Multimodal Face Anti-Counterfeiting Recognition Method, Device and Medium

By introducing prompt learning technology in the multimodal face anti-counterfeiting recognition task, it processes the visual and prompt token sequences of multimodal face images, and combined with the pre-trained neural network model, the problem of low recognition accuracy in traditional models in multimodal face recognition is solved, and higher recognition accuracy is achieved.

CN119360425BActive Publication Date: 2025-05-27GREATER BAY AREA UNIV (IN PREPARATION)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411291598.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-05-27
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

The traditional deep learning algorithm model performs poorly in multimodal face anti-counterfeiting recognition tasks, has low recognition accuracy, and is difficult to effectively deal with imbalances and feature deviations between data of different modalities.

Method used

By introducing prompt learning technology, modal prompt sets, task prompts and multimodal face images are obtained, and visual token sequences and prompt token sequences are obtained, and face anti-counterfeiting recognition is performed by combining pre-trained neural network models.

Benefits of technology

The neural network model has improved its understanding and processing ability of the modal characteristics of multimodal face data, enhanced its detection ability of subtle false face clues, and improved the recognition accuracy of face anti-counterfeiting recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360425B_ABST
    Figure CN119360425B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-modal face anti-counterfeiting recognition method, device and medium, which are applied to the field of face recognition technology. The method includes: obtaining a modal prompt set, a task prompt and a multi-modal face image, where the modal prompt set includes visual prompts related to the image modality, and the task prompt is a visual prompt associated with face anti-counterfeiting recognition; processing the modal prompt set, the task prompt and the multi-modal face image to obtain a visual token sequence and a prompt token sequence. The visual token sequence includes a classification token and multiple feature tokens. The classification token is the global token of the vision transformer, and the feature token is the vector representation of the multi-modal face image. The prompt token sequence includes multiple prompt tokens, and the prompt token is the global token associated with face anti-counterfeiting recognition; performing face anti-counterfeiting recognition according to the visual token sequence and the prompt token sequence in combination with a neural network model to obtain an anti-counterfeiting recognition result. The present application improves the accuracy of face anti-counterfeiting recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of face recognition technology, and in particular to a multimodal face anti-counterfeiting recognition method, device and medium. Background Art

[0002] Face anti-counterfeiting recognition technology mainly uses the difference in features between real face images and fake face images to judge real and fake faces, thereby preventing fake faces from attacking the face recognition system. In related technologies, deep learning algorithm models such as convolutional neural networks are used to achieve face anti-counterfeiting recognition. However, due to the imbalance and feature deviation between data of different modalities, traditional deep learning algorithm models perform poorly in multimodal face anti-counterfeiting recognition tasks, and their face anti-counterfeiting recognition accuracy needs to be improved. Summary of the invention

[0003] The embodiments of the present application provide a multimodal face anti-counterfeiting recognition method, device and medium for improving the recognition accuracy of face anti-counterfeiting.

[0004] On the one hand, the embodiment of the present application provides a multimodal face anti-counterfeiting recognition method, comprising the following steps:

[0005] Acquire a modality prompt set, a task prompt, and a multimodal face image; the modality prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, each of which has a different modality;

[0006] The modal prompt set, the task prompt and the multimodal face image are processed to obtain a visual token sequence and a prompt token sequence; the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of a visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition;

[0007] According to the visual token sequence and the prompt token sequence, combined with a pre-trained neural network model, face anti-counterfeiting recognition is performed to obtain an anti-counterfeiting recognition result.

[0008] On the other hand, an embodiment of the present application provides a multimodal face anti-counterfeiting recognition device, comprising:

[0009] An acquisition module, used to acquire a modality prompt set, a task prompt, and a multimodal face image; the modality prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, and the modalities of each face image are different;

[0010] The first processing module is used to process the modal prompt set, the task prompt and the multimodal face image to obtain a visual token sequence and a prompt token sequence; the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of the visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition;

[0011] The second processing module is used to perform face anti-counterfeiting recognition based on the visual token sequence and the prompt token sequence in combination with a pre-trained neural network model to obtain an anti-counterfeiting recognition result.

[0012] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement the above-mentioned multimodal face anti-counterfeiting recognition method.

[0013] The multimodal face anti-counterfeiting recognition method, device and medium of the present application first obtain a modal prompt set, a task prompt and a multimodal face image, wherein the modal prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, and the modalities of each face image are different; then, the modal prompt set, the task prompt and the multimodal face image are processed to obtain a visual token sequence and a prompt token sequence, wherein the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of the visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition; finally, according to the visual token sequence and the prompt token sequence, combined with a pre-trained neural network model, face anti-counterfeiting recognition is performed to obtain an anti-counterfeiting recognition result. It can be seen that by introducing prompt learning technology in the multimodal face anti-counterfeiting recognition task, the neural network model can be guided to pay attention to the modal characteristics and face forgery clues in the multimodal face data through the modal prompt set associated with the prompt learning technology, thereby realizing face anti-counterfeiting recognition. This can not only improve the neural network model's understanding and processing capabilities of the modal characteristics of multimodal face data, but also improve the neural network model's ability to detect subtle false face clues, thereby improving the recognition accuracy of face anti-counterfeiting.

[0014] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of the multimodal face anti-counterfeiting recognition method provided by this application;

[0016] Figure 2 It is a schematic diagram of the multimodal anti-counterfeiting identification method provided by this application;

[0017] Figure 3 is a schematic diagram of the gated adaptation module provided by this application;

[0018] Figure 4 It is a structural diagram of the multimodal face anti-counterfeiting recognition device provided in this application. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.

[0021] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0023] Face recognition technology refers to biometric technology that identifies people based on facial feature information. With the development and widespread application of face recognition technology, higher security requirements have been put forward for face recognition technology to cope with increasingly complex attack methods. These attack methods include not only traditional printing photos and video playback, but also more advanced three-dimensional masks and deep fake technology, which use artificial intelligence technology to generate fake faces that are almost indistinguishable to the naked eye in an attempt to deceive face recognition systems.

[0024] In order to improve the security of face recognition systems, face anti-counterfeiting recognition technology uses the difference in features between real face images and fake face images to judge real and fake faces, thereby preventing fake faces from attacking face recognition systems. Face anti-counterfeiting recognition technology significantly improves the robustness and accuracy of face recognition systems by integrating face images of multiple modalities. For example, attackers use forged three-primary-color face images to attack face recognition systems. However, face images of other modalities, such as depth face images and infrared face images, may also reveal the difference between fake faces and real faces, thereby helping face recognition systems identify fake faces. Among them, biometric data mainly include three-primary color images, depth images, infrared images and multispectral images. Each modal data provides a unique level of information. For example, three-primary color images provide color and texture information of the face, and three-primary color images are the most common data source for face recognition. For another example, depth images reveal the three-dimensional shape and contour of the face, which helps to distinguish between real faces and flat faces. For another example, infrared images can provide thermal imaging information of the face in low-light or completely dark environments, which helps to identify hidden facial features.

[0025] In related technologies, deep learning algorithm models such as convolutional neural networks are used to achieve face anti-counterfeiting recognition. However, data in different modalities often have different resolutions, noise levels, and acquisition conditions, which leads to imbalance and feature deviation between data in different modalities. Traditional deep learning algorithm models perform poorly in multimodal face anti-counterfeiting recognition tasks, and their face anti-counterfeiting recognition accuracy needs to be improved.

[0026] In view of this, the embodiments of the present application provide a multimodal face anti-counterfeiting recognition method, device and medium, aiming to improve the recognition accuracy of face anti-counterfeiting, especially the recognition accuracy in a multimodal face anti-counterfeiting recognition scenario.

[0027] The implementation steps of the multimodal face anti-counterfeiting recognition method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0028] It should be noted that in each implementation of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics such as user facial images, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when each implementation of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of each implementation of the present application will be obtained.

[0029] The multimodal face anti-counterfeiting recognition method provided in the embodiment of the present application can be applied to a terminal, a server, or software running in a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited to this. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. In addition, the server can also be a node server in a blockchain network, but is not limited to this. Among them, blockchain is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm.

[0030] Reference Figure 1 , Figure 1: is a flowchart of a multimodal face anti-counterfeiting recognition method provided by the present application. The multimodal face anti-counterfeiting recognition method may include the following steps S101-S103:

[0031] S101, obtaining a modal prompt set, a task prompt, and a multimodal face image; the modal prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, and the modalities of each face image are different;

[0032] S102, processing the modal prompt set, the task prompt and the multimodal face image to obtain a visual token sequence and a prompt token sequence; the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of the visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition;

[0033] S103, performing face anti-counterfeiting recognition according to the visual token sequence and the prompt token sequence in combination with a pre-trained neural network model to obtain an anti-counterfeiting recognition result.

[0034] In an embodiment of the present application, first, a modal prompt set, a task prompt and a multimodal face image are obtained, the modal prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, and the modality of each face image is different; then, the modal prompt set, the task prompt and the multimodal face image are processed to obtain a visual token sequence and a prompt token sequence, the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global label of the visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global label associated with face anti-counterfeiting recognition; finally, according to the visual token sequence and the prompt token sequence, combined with a pre-trained neural network model, face anti-counterfeiting recognition is performed to obtain an anti-counterfeiting recognition result. It can be seen that by introducing prompt learning technology in the multimodal face anti-counterfeiting recognition task, the neural network model can be guided to pay attention to the modal characteristics and face forgery clues in the multimodal face data through the modal prompt set associated with the prompt learning technology, thereby realizing face anti-counterfeiting recognition. This can not only improve the neural network model's understanding and processing capabilities of the modal characteristics of multimodal face data, but also improve the neural network model's ability to detect subtle false face clues, thereby improving the recognition accuracy of face anti-counterfeiting.

[0035] In the above step S101, a modal prompt set and a task prompt are obtained. The modal prompt set may include visual prompts related to the image modality, and the task prompt refers to a visual prompt associated with face anti-counterfeiting recognition. At the same time, a multimodal face image is obtained. The multimodal face image may include face images of multiple different modalities, and these face images are all face images of the same user.

[0036] The above-mentioned modality prompt set may include visual prompts related to the image modality, and the above-mentioned modality prompt set mainly focuses on the unique properties of each modality of the image, such as the three-dimensional information of the depth image or the thermal characteristics of the infrared image.

[0037] The above-mentioned task prompts can be visual prompts associated with face anti-counterfeiting recognition, that is, guiding text information directly related to specific multimodal face anti-counterfeiting recognition tasks, such as known deceptive behavior patterns or local areas of the face that require special attention from the neural network model.

[0038] The above-mentioned visual cues refer to learnable parameters or feature vectors, which can be pre-set according to actual conditions, and the embodiments of the present application do not make specific limitations on this.

[0039] For example, the above visual prompt may be specifically expressed in the form of a vector having the same size as that of the input image.

[0040] It should be noted that visual cues are mainly used to improve model adaptability and performance, and are applicable to the field of computer vision, such as the application of large pre-trained models. Specifically, by introducing visual cues in the input space of the model, the model can adapt to new downstream tasks without comprehensive fine-tuning. These visual cues can be pixel-level perturbations, specific image regions, or task-related feature vectors, which can be updated during model fine-tuning, while the main part of the model remains unchanged.

[0041] Optionally, both the above modal prompt set and the above task prompt adopt input-level prompts, that is, visual prompts will only accompany the input into the network reasoning, and new visual prompts will not be inserted in each layer.

[0042] The above-mentioned task prompts and the above-mentioned modal prompt sets can be pre-set according to actual conditions, and the embodiments of the present application do not make specific limitations on this.

[0043] The multimodal facial image may include multiple facial images belonging to the same user, and the modalities of the facial images are different.

[0044] The modalities of the above-mentioned facial images can be set according to actual conditions, and the embodiments of the present application do not make any specific limitations on this.

[0045] For example, the multimodal facial images may include facial images in near-infrared modality, facial images in thermal imaging modality, facial images in multispectral modality, and facial images in depth modality, but are not limited thereto.

[0046] In the above step S102, after obtaining the modal prompt set, task prompt and multimodal face image, the modal prompt set, task prompt and multimodal face image are processed. The processing may include but is not limited to data preprocessing and preliminary feature extraction, so as to obtain a visual token sequence and a prompt token sequence, so as to realize face anti-counterfeiting recognition based on the visual token sequence and the prompt token sequence.

[0047] The above visual token sequence may include a classification token and multiple feature tokens.

[0048] The above classification token is a global tag used in the Vision-Transformer (ViT) structure for downstream classification tasks. As a special tag, the classification token can aggregate the information of all other tokens to achieve global feature aggregation. Since it does not contain image content, it can reduce the bias towards a specific token in the sequence, thereby enabling the neural network model to obtain better and more stable classification performance.

[0049] The above-mentioned feature token is a vector representation. Specifically, by segmenting the multimodal face image, a number of fixed-size pixel blocks can be obtained, and the feature token refers to the vector representation obtained by mapping a single fixed-size pixel block in the multimodal face image, which mainly represents the local area in the multimodal face image and is input into the visual transformer structure as part of the visual token sequence.

[0050] It should be noted that the difference between classification tokens and feature tokens is that classification tokens aggregate the overall information of multimodal face images and occupy a fixed position in the input sequence and output sequence of the neural network model; while feature tokens represent local areas of multimodal face images and their positions in the input sequence are variable. In addition, classification tokens are mainly used to make final classification decisions in the output sequence of the neural network model, while feature tokens participate in feature extraction and information interaction in the middle layer of the neural network model, so that the neural network model learns the representation of local features.

[0051] The above prompt token sequence may include multiple prompt tokens.

[0052] The above-mentioned prompt token refers to a global tag associated with face anti-counterfeiting recognition. The prompt token is also a special tag. Through some auxiliary face anti-counterfeiting recognition prompt information, the prompt token can help the neural network model better perceive the current multimodal face anti-counterfeiting recognition task and obtain more global task information, so as to better assist the neural network model to complete the multimodal face anti-counterfeiting recognition task.

[0053] The number of the above hint tokens is equal to the number of the above feature tokens.

[0054] The above-mentioned processing of multimodal facial images and modal prompt sets to obtain visual token sequences and prompt token sequences may include preprocessing and feature extraction of multimodal facial images to obtain visual token sequences, and preprocessing and feature extraction of modal prompt sets to obtain prompt token sequences, but is not limited to this.

[0055] In the above step S103, after obtaining the visual token sequence and the prompt token sequence, the visual token sequence and the prompt token sequence are used as inputs of a pre-trained neural network model, and face anti-counterfeiting recognition is performed through the neural network model to obtain an anti-counterfeiting recognition result.

[0056] The above anti-counterfeiting identification results can be set according to actual conditions, and the embodiments of the present application do not make specific limitations on this.

[0057] For example, if the multimodal face anti-counterfeiting recognition task is a binary classification task, that is, detecting whether the face is alive or not, the anti-counterfeiting recognition result may include either a live type or a deception type, where the live type is used to indicate that the multimodal face image is a real face image, and the deception type is used to indicate that the multimodal face image is a false face image.

[0058] For another example, if the multimodal face anti-counterfeiting recognition task is a multi-class classification task, that is, specifically detecting which type of face forgery behavior the multimodal face image belongs to, the anti-counterfeiting recognition result is the specific face forgery behavior.

[0059] The above neural network model can be set according to actual conditions, and this embodiment does not make any specific limitation to this.

[0060] The above steps will be further described below.

[0061] In some implementations, the above processing of the modal prompt set, the task prompt, and the multimodal face image to obtain the visual token sequence and the prompt token sequence may include:

[0062] Performing facial region cropping processing on each face image to obtain each cropped face image;

[0063] Performing enhancement processing on each cropped face image to obtain each enhanced face image;

[0064] Normalizing each enhanced face image to obtain each normalized face image;

[0065] Each normalized face image is spliced ​​and vector mapped to obtain a visual token sequence.

[0066] In this embodiment, first, face region cropping is performed on each face image, aiming to crop the face region in each modal face image so as to accurately locate the face region; then, each cropped face image is enhanced, aiming to improve the generalization ability of the neural network model for face images under different environmental conditions, and to enhance the recognition ability of the neural network model for face forgery by simulating various possible attack methods; thereafter, each enhanced face image is normalized, so that the pixel value of each modal face image is scaled within a preset range, so as to reduce the negative impact of extreme changes in image background pixel values ​​on the performance of the neural network model, such as background noise interference, extreme changes in illumination, extreme changes in contrast, etc., and at the same time provide a consistent data format for subsequent processing; finally, each normalized face image is spliced ​​to obtain a visual token sequence for subsequent face anti-counterfeiting recognition.

[0067] The aforementioned enhancement processing of each cropped facial image may include simultaneously or sequentially performing data enhancement processing of color, illumination, noise and geometric features on each cropped facial image, but is not limited thereto.

[0068] The above-mentioned stitching and vector mapping processing of each normalized face image to obtain a visual token sequence may include stitching and fusing each normalized face image to obtain a stitched face image, performing pixel block segmentation processing on the stitched face image to obtain a plurality of pixel blocks, and then mapping processing on each pixel block to obtain a vector representation corresponding to the pixel block as a feature token; constructing a visual token sequence through multiple feature tokens and one classification token.

[0069] Alternatively, the above-mentioned stitching and vector mapping processing of each normalized face image to obtain a visual token sequence may include stitching each normalized face image into a stitched face image in the channel dimension, performing pixel block segmentation processing on the stitched face image to obtain a plurality of pixel blocks, and then mapping processing on each pixel block to obtain a vector representation corresponding to the pixel block as a feature token; constructing a visual token sequence through multiple feature tokens and one classification token, but not limited to this.

[0070] In an example, the normalized face image can be represented as , Represents the original pixel matrix of the face image, represents the normalized face image, Represents the maximum pixel value in the face image, Represents the minimum pixel value in the face image, , Indicates the pixel height of the normalized face image, Represents the pixel width of the normalized face image.

[0071] The above pixel height and the above pixel width can be set according to actual conditions, and this embodiment does not specifically limit this.

[0072] For example, the pixel height and the pixel width may both be 224, but are not limited thereto.

[0073] In some implementations, the above processing of the modal prompt set, the task prompt, and the multimodal face image to obtain the visual token sequence and the prompt token sequence may also include:

[0074] The features of each normalized face image are extracted through a shallow convolutional neural network to obtain content prompts;

[0075] Extracting visual cues corresponding to the modality of each face image from the modality cue set as modality visual cues;

[0076] The content prompts, modal visual prompts and task prompts are concatenated and fused to obtain a prompt token sequence.

[0077] In this implementation, first, a shallow convolutional neural network is used to perform feature extraction on each normalized facial image to capture the preliminary visual features of the facial image of each modality, such as edges, textures, and color distribution, so as to obtain content prompts. The content prompts contain basic feature information of the facial image of each modality, which provides a basis for subsequent feature supplementation of multimodal facial images and ensures effective connection and information supplementation between data of different modalities.

[0078] Then, modal prompt sets and task prompts are introduced. The modal prompt set focuses on the unique properties of each modality of the image, such as the three-dimensional information of depth images or the thermal characteristics of infrared images, while the task prompts focus on guidance information directly related to specific multimodal face anti-counterfeiting recognition tasks, such as known deceptive behavior patterns or local areas of the face that require special attention from the neural network model. Based on this, while obtaining the content prompt, the visual prompt corresponding to the modality of each facial image is extracted from the modal prompt set as the modal visual prompt. For example, if the multimodal facial image includes a facial image in the visible light modality, a facial image in the infrared modality and a facial image in the depth modality, the visual prompt related to the visible light modality, the visual prompt related to the infrared modality and the visual prompt related to the depth modality are obtained as the modal visual prompt. After that, the content prompt, the modal visual prompt and the task prompt are spliced ​​and fused to obtain a prompt token sequence. As a powerful information carrier, the prompt token sequence not only includes the basic information of the facial image of each modality, but also integrates the guidance information for the specific multimodal facial anti-counterfeiting recognition task and the unique attributes of the facial image of each modality. This not only enriches the input information of the neural network model, but also provides the neural network model with a more comprehensive contextual understanding, so that the neural network model can more accurately capture subtle false facial clues when processing multimodal facial images, thereby improving the neural network model's recognition ability for complex facial forgery behaviors.

[0079] It should be noted that the essence of content prompts is also visual prompts, which are also input-level prompts. Their specific form can be a vector form with the same size as the input image, but it is not limited to this. In the embodiment of the present application, all visual prompts can be divided into two categories according to the response relationship. One category is the visual prompts that directly respond to the image input content, that is, content prompts, and their generation mainly relies on the parameter training of the shallow convolutional differential network that processes the input data; the other category is to directly use learnable parameters or feature vectors as visual prompts, that is, modal prompt sets (modal visual prompts) and task prompts, and their generation mainly relies on preset and learnable parameters or feature vectors directly as corresponding visual prompts, and relies on the optimization update of back propagation.

[0080] The above shallow convolutional neural network can be any neural network used to implement differential convolution.

[0081] The above shallow convolutional neural network can be set according to actual conditions, and this embodiment does not make any specific limitation to this.

[0082] For example, the shallow convolutional neural network may include a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence, wherein the first convolutional layer and the third convolutional layer are both ordinary convolutional layers, and the convolutional kernel sizes of the first convolutional layer and the third convolutional layer are both 1 1. The first and third convolutional layers are used to change the number of feature channels so that feature extraction can be achieved through differential convolution. The second convolutional layer is a differential convolutional layer, and the convolution kernel size of the second convolutional layer is 3. 3. The second convolutional layer is used to perform differential convolution on the features; a preset activation function is set behind the third convolutional layer.

[0083] The above activation function can be set according to actual conditions, and this embodiment does not specifically limit this.

[0084] For example, the activation function may be a SiLU activation function; or, the activation function may be a ReLU activation function, but is not limited thereto.

[0085] The above-mentioned splicing and fusing processing of the content prompt, the modal visual prompt and the task prompt to obtain the prompt token sequence may include fusing the content prompt, the modal visual prompt and the task prompt into the prompt token sequence.

[0086] Alternatively, the above-mentioned concatenation and fusion processing of the content prompt, the modal visual prompt and the task prompt to obtain the prompt token sequence may include concatenating the content prompt, the modal visual prompt and the task prompt into a prompt token sequence in the channel dimension, as shown in the following formula (1):

[0087] (1);

[0088] In formula (1), is the prompt token sequence, which is the input of the first encoding structure; For modal visual cues, For task prompts, Provides content tips; It means that content prompts, modal visual prompts and task prompts are spliced ​​in the channel dimension to ensure that all information related to face anti-counterfeiting recognition can be considered by the neural network model at the same time; is the number of channels, which can be set according to actual conditions, for example, the number of channels is 64, but not limited thereto.

[0089] In some embodiments, reference Figure 2 The schematic diagram of the multimodal face anti-counterfeiting recognition method shown in FIG. Figure 2 middle, Indicates input to A sequence of visual tokens encoding a structure, Indicates input to A sequence of hint tokens encoding the structure, Indicates content hint. Indicates a modal visual cue, Indicates task prompts. Indicates The feature sequence of prompt tokens output by the encoding structure, Indicates The visual token feature sequence output by the encoding structure is Indicates The visual feature sequence output by the visual transformer of the encoding structure is Indicates The mixed visual features output by the gated adaptation module of the encoding structure, Indicates The hybrid concatenation features output by the gated adaptation module of the encoding structure.

[0090] The neural network model may include a plurality of sequentially connected encoding structures and classifiers, wherein, for the last encoding structure, the visual token feature sequence output by the encoding structure is used as the input of the classifier; for the encoding structures other than the last encoding structure, the visual token feature sequence output by the encoding structure is used as the visual token sequence of the next encoding structure, and the prompt token feature sequence output by the encoding structure is used as the prompt token sequence of the next encoding structure;

[0091] Each encoding structure may include a visual transformer, a gated adaptation module and a prompt convolution module; the visual transformer is used to extract features from a visual token sequence to obtain a visual feature sequence; the gated adaptation module is used to extract features from a visual feature sequence and a prompt token sequence through a preset hybrid expert model to obtain hybrid splicing features and hybrid visual features; the visual transformer is also used to obtain a visual token feature sequence based on the hybrid visual features and the visual feature sequence; the prompt convolution module is used to obtain a prompt token feature sequence based on the hybrid splicing features and the prompt token sequence; the classifier is used to perform face anti-counterfeiting recognition based on the visual token feature sequence to obtain an anti-counterfeiting recognition result.

[0092] In this embodiment, an improved neural network model is provided, and multiple encoding structures of the neural network model are used to capture more and more comprehensive subtle false face clues. Each encoding structure introduces a visual transformer, a gated adaptation module and a prompt convolution module, wherein the gated adaptation module adopts a hybrid expert model structure and an isolated gating mechanism to achieve feature capture. This can improve the neural network model's ability to detect subtle false face clues and enhance the accuracy of face anti-counterfeiting recognition.

[0093] Specifically, in a single encoding structure, there are:

[0094] First, the visual token sequence is input into the visual transformer, which captures the global dependency and position information of the visual token sequence through multi-head attention mechanism and position encoding, and understands the complex relationships within the visual token sequence, such as spatial features and semantic information, to obtain the visual feature sequence, thus achieving deep feature extraction.

[0095] Then, the visual feature sequence and the hint token sequence are used as the input of the gated adaptation module. The feature extraction of the visual feature sequence and the hint token sequence is performed through the hybrid expert model structure and the isolated gating mechanism of the gated adaptation module to obtain hybrid splicing features and hybrid visual features, which aim to capture the visual feature representation of the face in the visual feature sequence and the hint feature representation in the hint token sequence, and to mine the subtle false face clues hidden in the high-frequency details of the face image, so as to obtain richer and more comprehensive image features and effective face forgery features.

[0096] Afterwards, the mixed visual features and the visual feature sequence are used as the intermediate input of the visual transformer, and the mixed visual features and the visual feature sequence are processed by the visual transformer to achieve cross-module feature fusion, thereby obtaining a visual token feature sequence, which will be used as one of the outputs of the encoding structure.

[0097] At the same time, the hybrid splicing features and the prompt token sequence are used as the input of the prompt convolution module. The hybrid splicing features and the prompt token sequence are extracted and fused by the prompt convolution module to obtain the prompt token feature sequence, which will be used as another output of the encoding structure. In this way, cross-module feature fusion is achieved, so that the prompt token feature sequence not only contains deep data features, but also integrates key facial visual feature representations and prompt feature representations associated with the face forgery recognition task; in addition, it can effectively capture and enhance subtle fake face clues, which are crucial to revealing potential face forgery behaviors, thereby enhancing the neural network model's perception of local fake face clues.

[0098] For other coding structures except the last one, the visual token feature sequence output by the coding structure is used as the visual token sequence of the next coding structure, and the prompt token feature sequence output by the coding structure is used as the prompt token sequence of the next coding structure, thus realizing cyclic feature extraction. For the last coding structure, the visual token feature sequence output by the coding structure is used as the input of the classifier, and the classifier uses the visual token feature sequence output by the last coding structure to perform face anti-counterfeiting recognition and obtain the anti-counterfeiting recognition result, thereby realizing face anti-counterfeiting recognition.

[0099] The number of the above coding structures can be set according to actual conditions, and this embodiment does not specifically limit this.

[0100] The above-mentioned vision transformer can be a module based on Vision-Transformer.

[0101] The above-mentioned gating adaptation module can be any module configured with a hybrid expert model.

[0102] The above-mentioned hint convolution module can be any module configured with a neural network for implementing feature extraction.

[0103] The above classifier can be any module used to implement a binary classification task or a multi-class classification task.

[0104] For example, the above classifier may be a classifier based on a Softmax function, but is not limited thereto.

[0105] The above-mentioned performing anti-counterfeiting recognition of a face based on a visual token feature sequence to obtain an anti-counterfeiting recognition result may include using a Softmax function to process the visual token feature sequence to obtain an anti-counterfeiting recognition result.

[0106] In some embodiments, reference Figure 2 In the visual transformer, the feature extraction of the visual token sequence to obtain the visual feature sequence may include:

[0107] Performing standardization on the visual token sequence to obtain a standardized visual token sequence;

[0108] The standardized visual token sequence is processed through a multi-head attention mechanism to obtain a sub-visual feature sequence;

[0109] The sub-visual feature sequence and the visual token sequence are added to obtain a visual feature sequence.

[0110] In this implementation, first, the visual token sequence is standardized to obtain a standardized visual token sequence; then, the standardized visual token sequence is processed through a multi-head attention mechanism to obtain a sub-visual feature sequence, so that the global dependency and position information of the standardized visual token sequence are captured through multiple sub-attention layers in the multi-head attention mechanism, and the complex relationships within the standardized visual token sequence, such as spatial features and semantic information, are understood, thereby achieving deep feature extraction; finally, the sub-visual feature sequence and the visual token sequence are added to transfer the low-level visual token sequence to the high-level sub-visual feature sequence, thereby obtaining a visual feature sequence, thereby achieving feature fusion and transmission.

[0111] In some embodiments, the hybrid expert model combines multiple expert networks together to obtain better prediction performance. The expert network can be understood as a specific neural network. The advantage of the hybrid expert model is that it can improve computing efficiency through a sparse activation mechanism. Specifically, in practical applications, only a portion of the outputs of the expert networks may be valid for the current task. By activating only these expert networks, the hybrid expert model can reduce unnecessary calculations, thereby improving computing speed and efficiency. In addition, the adaptability of the hybrid expert model enables it to flexibly handle different data distributions and task requirements. In the multimodal face anti-counterfeiting recognition task, the hybrid expert model can dynamically adjust the weights and activation modes of each expert network according to different types of attacks and different data environments to achieve optimal recognition performance.

[0112] Although hybrid expert models have many advantages in theory, in practical applications, how to design and train these expert networks and how to effectively integrate their outputs remain a challenge. In addition, with the growth of data scale and the increase of modality types, hybrid expert models may require more computing resources and storage space. Therefore, more research and innovation are needed in model design, training strategies, resource management, and system integration to achieve the effective application of hybrid expert models in multimodal face anti-counterfeiting recognition tasks and ensure the scalability and practicality of the system.

[0113] Specifically, the existing hybrid expert model still faces the following challenges in the multimodal face anti-counterfeiting recognition task:

[0114] On the one hand, for the hybrid expert model, the decision-making process of its gating network is crucial. In the multimodal face anti-counterfeiting recognition task, the input data not only contains useful facial feature information, but also may contain noise introduced by environmental factors, such as illumination changes, facial occlusion, expression changes, etc. The gating network of the hybrid expert model is often more sensitive to these noises, which leads to the risk of erroneous activation of the gating network or inhibition of the expert network, causing the hybrid expert model to misjudge when facing normal facial changes or environmental noise, affecting the accuracy and stability of face anti-counterfeiting recognition.

[0115] On the other hand, in the multimodal face anti-counterfeiting recognition task, subtle fake face clues are often hidden in high-frequency details, which may be information missed by advanced counterfeiting techniques. For example, in high-resolution face images, details such as the microstructure of the skin, the distribution of pores, and tiny wrinkles caused by facial muscle movement are all key clues to judging the authenticity of the face. In addition, under high dynamic range lighting, fake faces may show unnatural reflections and shadows, which are also key clues to judging the authenticity of the face. In the multimodal face anti-counterfeiting recognition task, it is crucial to capture the above key clues. The hybrid expert model usually activates the corresponding expert network through a gating network, and captures the key clues through the activated expert network. However, the gating network may mistakenly activate some inappropriate expert networks, resulting in certain limitations of the hybrid expert model in capturing these high-frequency details, unnatural reflections and shadows.

[0116] On the other hand, cue learning plays an important role in guiding the neural network model to focus on key information in the input data. However, in the hybrid expert model, the sensitivity of its gating network to the cue token may cause the hybrid expert model to rely too much on specific cues, thereby ignoring other equally important feature information in the face image. For example, if the cue token focuses mainly on a specific area of ​​the face, such as the eyes or mouth, and ignores subtle changes in other areas of the face, this may limit the hybrid expert model's comprehensive analysis of the overall facial features.

[0117] In this regard, this embodiment provides a gated adaptation module to enable the hybrid expert model to be directly applied to the multimodal face anti-counterfeiting recognition task, thereby giving full play to the performance advantages of the hybrid expert model in the multimodal face anti-counterfeiting recognition task. Figure 3 In the above-mentioned gated adaptation module, the hybrid expert model includes several expert networks, each of which corresponds to a specific subtask, and they independently process input data and extract features; the above-mentioned feature extraction of the visual feature sequence and the prompt token sequence by the preset hybrid expert model to obtain the hybrid splicing feature and the hybrid visual feature may include:

[0118] Performing convolution processing on multiple feature tokens in the visual feature sequence to obtain multiple convolution-processed feature tokens;

[0119] Concatenate multiple convolution-processed feature tokens and prompt token sequences to obtain mixed concatenated features;

[0120] The mixed concatenation features are gated based on the product key-value indexing mechanism to obtain a probability distribution, which is used to characterize the contribution of each expert network.

[0121] Activating at least one expert network according to the probability distribution to obtain at least one activated expert network;

[0122] For each feature token after convolution processing, extract features of the feature token after convolution processing by at least one of the activated expert networks to obtain an initial token feature representation corresponding to the feature token after convolution processing;

[0123] The classification tokens in the visual feature sequence and the token feature representations corresponding to the feature tokens after multiple convolution processes are concatenated to obtain a mixed visual feature.

[0124] In this embodiment, first, multiple feature tokens in the visual feature sequence are convolved to obtain multiple feature tokens after convolution, in order to achieve preliminary feature extraction; then, multiple feature tokens after convolution and multiple prompt tokens in the prompt token sequence are spliced ​​to obtain mixed splicing features, in order to achieve feature fusion. In this way, the mixed splicing features are obtained through convolution and splicing, and the mixed splicing features, as the intermediate output of the gated adaptation module, will be fused with the rich prompt information provided by the prompt convolution module in the subsequent steps, so as to achieve cross-module feature fusion and capture and enhance subtle false face clues.

[0125] After obtaining the hybrid concatenated features, the hybrid concatenated features are gated based on the product key-value indexing mechanism, aiming to quickly retrieve the most relevant expert network based on the features of the input data and the low-dimensional subkeys, and generate a probability distribution based on the features of the input data and the retrieved expert subset, which can indicate the contribution of each expert network to the feature extraction process. After that, at least one expert network is activated according to the probability distribution, aiming to use the probability distribution to activate at least one expert network, and these activated expert networks will be used to process each convolution-processed feature token. After activating at least one expert network, for each convolution-processed feature token, at least one activated expert network is used to extract features from the convolution-processed feature token to obtain the token feature representation corresponding to the convolution-processed feature token. Finally, the classification tokens in the visual feature sequence and the token feature representations corresponding to the multiple convolution-processed feature tokens are concatenated to obtain a hybrid visual feature to achieve feature fusion.

[0126] It can be seen that the hybrid splicing features obtained through convolution processing and splicing processing as the object of gating processing can balance the differences between different prompt tokens and reduce the possibility of conflicts between different prompt tokens. In addition, the core of the gating adaptation module lies in the isolation gating mechanism and the hybrid expert model, where:

[0127] The isolated gating mechanism is mainly reflected in two aspects. On the one hand, the gated network processing and the hybrid expert model are isolated, which can effectively reduce the adverse effects of noise introduced by environmental factors on the gating decision, reduce the risk of erroneous activation or inhibition of the gated network, and make the expert networks of the hybrid expert model not make misjudgments when facing normal facial changes or environmental noise. On the other hand, it is based on the gating processing of the product key value indexing mechanism. Through this gating processing, the expert network suitable for feature extraction of each feature token can be more accurately selected and activated, and the feature capture ability of the hybrid expert model can be improved. In addition, the product key value indexing mechanism can provide efficient indexing operations, which increases the number of available expert networks by reducing the computational overhead of the indexing operation of the gating decision. More expert networks are conducive to extracting fine-grained false face clues and coarse-grained false face clues. For example, this embodiment can provide 1,600 available expert networks by combining product key value indexing and parameter efficient experts, while the previous ordinary indexing and expert settings often limit the number of expert networks to about 64.

[0128] The hybrid expert model uses the expert networks activated in the hybrid expert model to realize feature extraction. Through the expert networks activated in the hybrid expert model, it is possible to capture subtle false face clues hidden in the high-frequency details of the face image, and to explore the unnatural reflections and shadow textures presented by the false faces, thereby improving the neural network model's ability to detect subtle false face clues and ensuring the accuracy of face anti-counterfeiting recognition.

[0129] The hybrid expert model may include several expert networks, and the expert network may be a parameter-efficient expert network, which is composed of a single-neuron multi-layer perceptron, that is, there is only one hidden layer, and the layer has only one neuron.

[0130] The number of the above expert networks can be set according to actual conditions, and this embodiment does not specifically limit this.

[0131] For example, the number of the expert networks is 1600, but is not limited thereto.

[0132] The above-mentioned convolution processing on multiple feature tokens in the visual feature sequence may include convolution processing on multiple feature tokens in the visual feature sequence through a preset convolution layer, but is not limited to this.

[0133] The convolution kernel size of the above-mentioned preset convolution layer can be set according to actual conditions, and this embodiment does not make any specific limitation on this.

[0134] For example, the above preset convolution layer can be a convolution kernel with a size of 1 1 convolutional layer, but not limited to this.

[0135] The above-mentioned concatenating the feature tokens and prompt token sequences after multiple convolution processes to obtain the mixed concatenated features may include concatenating the feature tokens and prompt token sequences after multiple convolution processes into the mixed concatenated features in the channel dimension.

[0136] The above-mentioned splicing processing of feature tokens and prompt token sequences after multiple convolution processes to obtain a mixed splicing feature can include splicing feature tokens and prompt tokens belonging to the same position according to the position of each feature token in the visual feature sequence and the position of each prompt token in the prompt token sequence, so as to obtain a mixed splicing feature, but is not limited to this.

[0137] The above probability distribution refers to the contribution of each expert network to the feature extraction process. The final output can be a mixed visual feature or a token feature representation, but is not limited to this.

[0138] The above probability distribution may include a feature gating score and a prompt gating score, wherein the feature gating score refers to the contribution of each expert network to the feature token, and the prompt gating score refers to the contribution of each expert network to the prompt token.

[0139] The above-mentioned gating processing of the mixed splicing features based on the product key value indexing mechanism to obtain the probability distribution may include performing product key value indexing processing on the mixed splicing features, and then processing the mixed splicing features after the product key value indexing processing through a gating network to obtain the probability distribution, but is not limited to this.

[0140] The above-mentioned activating at least one expert network according to the probability distribution to obtain at least one activated expert network may include sorting a number of expert networks according to the contribution of each expert network, wherein the sorting sequence number of the expert network is positively correlated with the contribution of the expert network, that is, the larger the sorting sequence number of the expert network, the greater the contribution of the expert network; activating the last N expert networks to obtain N activated expert networks, N≥1.

[0141] Alternatively, the above-mentioned activating at least one expert network according to the probability distribution to obtain at least one activated expert network may include sorting a number of expert networks according to the contribution of each expert network, wherein the sorting number of the expert network is negatively correlated with the contribution of the expert network, that is, the smaller the sorting number of the expert network, the greater the contribution of the expert network; activating the first N expert networks to obtain N activated expert networks, N≥1, but is not limited to this.

[0142] The above-mentioned feature extraction of the feature tokens after convolution processing by at least one activated expert network to obtain the token feature representation corresponding to the feature tokens after convolution processing may include extracting the feature tokens after convolution processing by each activated expert network to obtain the token features output by each activated expert network; and concatenating the token features output by each activated expert network to obtain the token feature representation corresponding to the feature tokens after convolution processing, but is not limited to this.

[0143] The above-mentioned splicing processing of the classification token in the visual feature sequence and the token feature representations corresponding to multiple convolution-processed feature tokens to obtain a mixed visual feature may include splicing the classification token in the visual feature sequence in front of the token feature representation corresponding to the first convolution-processed feature token or behind the token feature representation corresponding to the last convolution-processed feature token, thereby obtaining a mixed visual feature, but is not limited to this.

[0144] In some implementations, in the gated adaptation module, the gated processing based on the product key value indexing mechanism on the mixed splicing features to obtain the probability distribution may include:

[0145] Perform product key-value retrieval processing on the mixed concatenation features to obtain a retrieval score;

[0146] The retrieval score is processed by a preset noise term to obtain a retrieval score fused with the noise term;

[0147] Isolate the retrieval scores fused with noisy items to obtain gated scores;

[0148] The gated scores are processed by the Top-K function and the Softmax function to obtain the probability distribution.

[0149] This embodiment is mainly implemented by a gating network based on a product key-value indexing mechanism. Specifically, in the gating network, first, the mixed splicing features are processed by the product key-value retrieval technology to obtain a preliminary retrieval score; then, the retrieval score is processed by the noise term to obtain a retrieval score fused with noise, and then the retrieval score fused with noise is isolated to obtain a preliminary gating score; then, the gating score is processed by the Top-K function to select the Top-K expert networks; finally, the Top-K expert networks are converted into the final gating score result, i.e., the probability distribution, by the Softmax function. In this way, the expert network suitable for feature extraction of each feature token can be selected and activated more accurately, thereby improving the feature capture capability of the hybrid expert model.

[0150] The above gating process based on the product key value index mechanism satisfies the following formula (2):

[0151] (2);

[0152] In formula (2), Represents the probability distribution The contribution of the expert network; Represents the Top-K function; For fusion of noisy input Conduct isolation treatment; Indicates that the input is retrieved by product key value retrieval technique and preset model parameters The result of the operation; represents the preset noise term; Represents the number of selected expert networks; the input of this formula Refers to mixed splicing features.

[0153] In some embodiments, in the gated adaptation module, the feature extraction of the feature token after convolution processing by at least one activated expert network to obtain a token feature representation corresponding to the feature token after convolution processing may include:

[0154] The feature tokens after convolution processing are input into each activated expert network, and features are extracted through each activated expert network to obtain the initial feature representation corresponding to each activated expert network;

[0155] The initial feature representations corresponding to each activated expert network are weightedly combined to obtain the token feature representation corresponding to the feature token after convolution processing.

[0156] In this embodiment, after activating the expert network, first, the feature tokens after convolution processing are input into each activated expert network, and feature extraction is performed through each activated expert network to obtain the initial feature representation corresponding to each activated expert network, aiming to capture the subtle false face clues hidden in the high-frequency details of the face image, and to mine the unnatural reflection and shadow texture presented by the false face; finally, the initial feature representations corresponding to each activated expert network are weightedly combined to obtain the token feature representation corresponding to the feature tokens after convolution processing, so as to achieve feature fusion. In this way, the ability of the neural network model to detect subtle false face clues can be improved, and the accuracy of face anti-counterfeiting recognition can be improved.

[0157] The output of the above activated expert network satisfies the following formula (3):

[0158] (3);

[0159] In formula (3), represents the output of a single activated expert network, represents the operation of a single activated expert network, represents the parameter weight of a single activated expert network, For input Isolate the express The weight parameter of this formula is Refers to the feature token after convolution processing.

[0160] Optionally, the hybrid expert model supports some form of interaction or combination with the output of the gating network, that is, the output of the hybrid expert model interacts or combines with the output of the gating network to obtain the final feature representation, and its specific implementation can be set according to actual conditions. For example, a nonlinear transformation is performed on the token feature representation corresponding to the feature token after convolution processing and the probability distribution of the gating network output, so as to obtain the token feature representation corresponding to the feature token after convolution processing.

[0161] The output of the above-mentioned gate adaptation module satisfies the following formula (4):

[0162] (4);

[0163] In formula (4), Represents the token feature representation corresponding to the feature token after a single convolution process, Indicates The output of the activated expert network, For the The parameter weights of the activated expert network, represents the parameter weight set of all activated expert networks, represents the total number of activated expert networks, It refers to The contribution of an activated expert network, which is also the output of the gating network.

[0164] In some embodiments, reference Figure 2 In the visual transformer, the visual token feature sequence is obtained according to the mixed visual feature and the visual feature sequence, which may include:

[0165] Performing standardization processing on the visual feature sequence to obtain a standardized visual feature sequence;

[0166] Performing forward propagation processing on the standardized visual feature sequence to obtain a visual feature sequence after forward propagation;

[0167] The visual feature sequence after forward propagation and the mixed visual feature are added to obtain a visual token feature sequence.

[0168] In this embodiment, in the visual transformer, first, the visual feature sequence is standardized to obtain a standardized visual feature sequence; then, the standardized visual feature sequence is forward propagated to obtain a forward-propagated visual feature sequence; finally, the forward-propagated visual feature sequence and the mixed visual feature are added to obtain a visual token feature sequence, thereby realizing cross-module feature fusion.

[0169] In some embodiments, reference Figure 2 In the above-mentioned prompt convolution module, the prompt token feature sequence is obtained according to the mixed concatenation feature and the prompt token sequence, which may include:

[0170] Perform differential convolution processing on the prompt token sequence to obtain a prompt token sequence after convolution processing;

[0171] Perform differential convolution processing on the mixed splicing features to obtain the mixed splicing features after convolution processing;

[0172] Concatenate the prompt token sequence after the convolution process and the mixed concatenation feature after the convolution process to obtain the prompt concatenation feature;

[0173] The prompt concatenated features are processed through the attention mechanism to obtain the prompt token feature sequence.

[0174] In this embodiment, in the prompt convolution module, first, the prompt token sequence is differentially convolved to obtain the prompt token sequence after convolution, and the mixed splicing feature is differentially convolved to obtain the mixed splicing feature after convolution, so as to realize feature extraction; then, the prompt token sequence after convolution and the mixed splicing feature after convolution are spliced ​​to obtain the prompt splicing feature to realize multimodal feature fusion; finally, the prompt splicing feature is processed by the attention mechanism to obtain the prompt token feature sequence, which not only contains deep data features, but also integrates key facial visual feature representations and prompt feature representations associated with the facial forgery recognition task. It can be seen that the prompt convolution module can make the neural network model more focused on key facial features, while avoiding the neural network model from being disturbed or misled by noise information, so as to effectively capture, perceive and enhance subtle false face clues, improve the quality and depth of feature extraction, and enhance the neural network model's perception of local false face clues.

[0175] The prompt token feature sequence output by the above prompt convolution module satisfies the following formula (5):

[0176] (5);

[0177] In formula (5), Indicates The hint token feature sequence output by the hint convolution module of each encoding structure, for all encoding structures except the last one, It can be understood as A sequence of hint tokens encoding a structure; No. A sequence of hint tokens encoding a structure; Indicates The hybrid concatenation features output by the gated adaptation module of the encoding structure; Indicates splicing processing; Represents the attention mechanism.

[0178] In some embodiments, performing face anti-counterfeiting recognition based on the visual token feature sequence to obtain the anti-counterfeiting recognition result may include:

[0179] The Softmax function is used to process the classification tokens in the visual token feature sequence to obtain the anti-counterfeiting recognition result.

[0180] In this implementation, the Softmax function is used to process the classification tokens in the visual token feature sequence, thereby achieving face anti-counterfeiting recognition.

[0181] In some embodiments, the above method may further include:

[0182] According to the anti-counterfeiting identification results, the parameters of the neural network model are updated.

[0183] In this implementation, in order to improve the performance and generalization ability of the neural network model, the parameters of the neural network model are updated based on the anti-counterfeiting identification results.

[0184] The above-mentioned updating of the parameters of the neural network model according to the anti-counterfeiting identification result may include, when the anti-counterfeiting identification result may include either a living type or a deception type, calculating the difference value between the anti-counterfeiting identification result and the expected identification result through a preset loss function, and updating the parameters of the visual transformer of the neural network model based on the difference value in combination with back propagation, but is not limited to this.

[0185] The above loss function can be set according to actual conditions, and this embodiment does not make any specific limitation to this.

[0186] For example, the above loss function can be a cross entropy loss function, which satisfies the following formula (6):

[0187] (6);

[0188] In formula (6), Indicates the difference between the anti-counterfeiting recognition result and the expected recognition result; Indicates the expected recognition result, that is, the true label of the anti-counterfeiting recognition result, and the true label is 1 or 0; is the probability that the anti-counterfeiting identification result is category 1.

[0189] To facilitate the understanding of the above-mentioned multi-modal face anti-counterfeiting recognition method of the present application, the principle of the above-mentioned multi-modal face anti-counterfeiting recognition method of the present application will be explained below with an example. Figure 2 and Figure 3 The specific process of implementing face anti-counterfeiting recognition in this example can be shown in the following steps S201-S205.

[0190] S201, obtain modal prompt set and task prompt , visible light face images, depth face images and infrared face images. The modal prompt set includes visual cues related to the image modality. The task prompts are visual cues associated with face anti-counterfeiting recognition. The face images of the above modalities are all face images of the same user.

[0191] S202, firstly, the visible light face image, the depth face image and the infrared face image are cropped, enhanced and normalized in turn to obtain normalized visible light face image, normalized depth face image and normalized infrared face image, wherein the enhancement processing refers to the data enhancement processing of color, illumination, noise and geometric features, and then the normalized visible light face image, the normalized depth face image and the normalized infrared face image are spliced ​​into a spliced ​​face image in the channel dimension, and the spliced ​​face image is segmented into pixel blocks to obtain a plurality of pixel blocks, and then for each pixel block, the pixel block is mapped to obtain a vector representation corresponding to the pixel block as a feature token, and finally a visual token sequence is constructed through a plurality of feature tokens and a classification token. .

[0192] In addition, after obtaining the normalized visible light face image, the normalized depth face image and the normalized infrared face image, the features of the normalized visible light face image, the normalized depth face image and the normalized infrared face image are extracted through a shallow convolutional neural network to obtain content prompts. , and then extracting the visual prompts corresponding to the modalities of each of the face images from the modal prompt set as modal visual prompts , and then the content prompts are , modal visual cues And task tips Concatenate into a sequence of prompt tokens .

[0193] S203, the visual token sequence and the prompt token sequence The first input to the neural network model The coding structure is The coding structure is processed to obtain the The feature sequence of prompt tokens output by the encoding structure and visual token feature sequences .

[0194] S204, judgment Is it true? If not, the prompt token feature sequence output by the current encoding structure is used as the prompt token sequence input to the next encoder, and the visual token feature sequence output by the current encoding structure is used as the visual token sequence input to the next encoder. , return to the above step S203 to implement the cyclic feature extraction; if so, enter step S205.

[0195] S205, using the Softmax function to process the classification tokens in the visual token feature sequence output by the last encoding structure to obtain an anti-counterfeiting recognition result.

[0196] The above steps start from start.

[0197] Among them, for a single coding structure, its implementation principle is as follows:

[0198] Visual Token Sequence is input to the visual transformer. In the visual transformer, first, the visual token sequence The standardized visual token sequence is processed to obtain a standardized visual token sequence; then, the standardized visual token sequence is processed by a multi-head attention mechanism to obtain a sub-visual feature sequence; finally, the sub-visual feature sequence and the visual token sequence are added to obtain a visual feature sequence .

[0199] Visual feature sequence and the prompt token sequence Together they serve as the input of the gated adaptation module. In the gated adaptation module, first, a convolution kernel size of 1 is used. 1 convolutional layer for visual feature sequence Convolution is performed on multiple feature tokens in the sequence to obtain multiple feature tokens after convolution. Secondly, multiple feature tokens after convolution and prompt token sequences are convolutionally processed. Perform splicing processing to obtain mixed splicing features ; Then, a gating network based on the product key-value indexing mechanism is used to concatenate the mixed features Processing is performed to obtain a probability distribution. The implementation of the gating network is shown in the above formula (2). After that, the feature tokens after convolution processing are input into each activated expert network, and feature extraction is performed through each activated expert network to obtain the initial feature representation corresponding to each activated expert network. Then, the initial feature representation corresponding to each activated expert network is weightedly combined to obtain the token feature representation corresponding to the feature token after convolution processing. The implementation of a single activated expert network is shown in the above formula (3). Finally, the visual feature sequence The classification token in is concatenated in front of the token feature representation corresponding to the feature token after the first convolution to obtain a mixed visual feature. .

[0200] Mixed visual features and visual feature sequences As the intermediate input of the visual transformer. In the visual transformer, first, the visual feature sequence The standardized visual feature sequence is obtained by performing a forward propagation process on the standardized visual feature sequence to obtain a forward propagated visual feature sequence; finally, the forward propagated visual feature sequence and the mixed visual feature sequence are processed. Perform addition processing to obtain the visual token feature sequence .

[0201] At the same time, the mixed splicing features and the prompt token sequence Together they serve as the input of the prompt convolution module. In the prompt convolution module, first, the prompt token sequence Perform differential convolution to obtain the prompt token sequence after convolution, and the mixed splicing features Perform differential convolution to obtain the mixed splicing features after convolution; then, concatenate the prompt token sequence after convolution and the mixed splicing features after convolution to obtain the prompt splicing features; finally, process the prompt splicing features through the attention mechanism to obtain the prompt token feature sequence .

[0202] In addition, refer to Figure 4 , the embodiment of the present application also provides a multi-modal face anti-counterfeiting recognition device, which may include:

[0203] The acquisition module 301 is used to acquire a modality prompt set, a task prompt, and a multimodal face image; the modality prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, and the modalities of each face image are different;

[0204] The first processing module 302 is used to process the modal prompt set, the task prompt and the multimodal face image to obtain a visual token sequence and a prompt token sequence; the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of the visual transformer, and the feature token is a vector representation obtained by mapping the pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition;

[0205] The second processing module 302 is used to perform face anti-counterfeiting recognition based on the visual token sequence and the prompt token sequence in combination with a pre-trained neural network model to obtain an anti-counterfeiting recognition result.

[0206] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0207] Finally, an embodiment of the present application also provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement the above-mentioned multimodal face anti-counterfeiting recognition method.

[0208] The contents of the above method embodiments are all applicable to the present storage medium embodiments. The functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0209] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the application is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Optional embodiments are expected, wherein the order of various operations is changed and the sub-operation described as a part of a larger operation is performed independently.

[0210] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present application. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional techniques of the engineer. Therefore, those skilled in the art can implement the present application set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the attached claims and their equivalents.

[0211] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several programs to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0212] The logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable programs for implementing logical functions, and may be embodied in any computer-readable medium for use by a program execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch and execute a program from a program execution system, device or apparatus), or in conjunction with such program execution system, device or apparatus. For purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by a program execution system, device or apparatus, or in conjunction with such program execution system, device or apparatus.

[0213] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0214] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0215] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0216] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.

[0217] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the described embodiments. Technical personnel familiar with the field can also make various equivalent modifications or substitutions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A multimodal face anti-counterfeiting recognition method, characterized in that: The following steps are involved: Acquire a modality prompt set, a task prompt, and a multimodal face image; the modality prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, each of which has a different modality; The modal prompt set, the task prompt and the multimodal face image are processed to obtain a visual token sequence and a prompt token sequence; the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of a visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition; According to the visual token sequence and the prompt token sequence, combined with a pre-trained neural network model, face anti-counterfeiting recognition is performed to obtain an anti-counterfeiting recognition result; Wherein, the neural network model includes several encoding structures and classifiers connected in sequence; for the last encoding structure, the visual token feature sequence output by the encoding structure is used as the input of the classifier; for other encoding structures except the last encoding structure, the visual token feature sequence output by the encoding structure is used as the visual token sequence of the next encoding structure, and the prompt token feature sequence output by the encoding structure is used as the prompt token sequence of the next encoding structure; the encoding structure includes a visual transformer, a gated adaptation module and a prompt convolution module, wherein the visual transformer is used to extract features from the visual token sequence to obtain a visual feature sequence; the gated adaptation module is used to extract features from the visual feature sequence and the prompt token sequence through a preset hybrid expert model to obtain a mixed splicing feature and a mixed visual feature; the visual transformer is also used to obtain a visual token feature sequence based on the mixed visual feature and the visual feature sequence; the prompt convolution module is used to obtain a prompt token feature sequence based on the mixed splicing feature and the prompt token sequence; the classifier is used to perform face anti-counterfeiting recognition based on the visual token feature sequence to obtain the anti-counterfeiting recognition result; The hybrid expert model includes a plurality of expert networks; the feature extraction of the visual feature sequence and the prompt token sequence by the preset hybrid expert model to obtain hybrid splicing features and hybrid visual features includes: Performing convolution processing on a plurality of feature tokens in the visual feature sequence to obtain a plurality of convolution-processed feature tokens; Performing concatenation processing on the feature tokens after the convolution processing and the prompt token sequence to obtain the mixed concatenation feature; Performing gating processing on the mixed concatenation features based on a product key value indexing mechanism to obtain a probability distribution, wherein the probability distribution is used to characterize the contribution of each of the expert networks; Activating at least one of the expert networks according to the probability distribution to obtain at least one activated expert network; For each of the feature tokens after the convolution processing, feature extraction is performed on the feature token after the convolution processing through at least one of the activated expert networks to obtain a token feature representation corresponding to the feature token after the convolution processing; The classification tokens in the visual feature sequence and the token feature representations corresponding to the plurality of feature tokens after the convolution processing are concatenated to obtain the mixed visual feature.

2. The multimodal face anti-counterfeiting recognition method according to claim 1, characterized in that: The processing of the modal prompt set, the task prompt and the multimodal face image to obtain a visual token sequence and a prompt token sequence includes: Performing facial region cropping processing on each of the facial images to obtain cropped facial images; Performing enhancement processing on each of the cropped face images to obtain each enhanced face image; Normalizing each of the enhanced face images to obtain each normalized face image; Each of the normalized face images is subjected to splicing processing and vector mapping processing to obtain the visual token sequence.

3. The multimodal face anti-counterfeiting recognition method according to claim 2, characterized in that: The processing of the modal prompt set, the task prompt and the multimodal face image to obtain a visual token sequence and a prompt token sequence also includes: Extracting features from each of the normalized face images through a shallow convolutional neural network to obtain content prompts; Extracting, from the modality prompt set, visual prompts corresponding to the modality of each of the face images as modality visual prompts; The content prompt, the modal visual prompt and the task prompt are spliced ​​and fused to obtain the prompt token sequence.

4. The multimodal face anti-counterfeiting recognition method according to claim 1, characterized in that: The step of extracting features from the visual token sequence to obtain a visual feature sequence comprises: Performing standardization processing on the visual token sequence to obtain a standardized visual token sequence; Processing the standardized visual token sequence through a multi-head attention mechanism to obtain a sub-visual feature sequence; The sub-visual feature sequence and the visual token sequence are added together to obtain the visual feature sequence.

5. The multimodal face anti-counterfeiting recognition method according to claim 1, characterized in that: The gate processing based on the product key value index mechanism is performed on the mixed splicing feature to obtain a probability distribution, including: Performing product key value retrieval processing on the mixed concatenation features to obtain a retrieval score; Processing the search score by a preset noise term to obtain a search score integrated with the noise term; Isolating the retrieval score of the fused noisy item to obtain a gated score; The gating score is processed by a Top-K function and a Softmax function to obtain the probability distribution.

6. The multimodal face anti-counterfeiting recognition method according to claim 1, characterized in that: The extracting features of the feature tokens after the convolution processing by at least one of the activated expert networks to obtain token feature representations corresponding to the feature tokens after the convolution processing includes: Inputting the feature tokens after the convolution processing into each of the activated expert networks, performing feature extraction through each of the activated expert networks, and obtaining initial feature representations corresponding to each of the activated expert networks; The initial feature representations corresponding to the activated expert networks are weightedly combined to obtain token feature representations corresponding to the feature tokens after the convolution processing.

7. The multimodal face anti-counterfeiting recognition method according to claim 1, characterized in that: The step of obtaining a visual token feature sequence according to the mixed visual feature and the visual feature sequence comprises: Performing standardization processing on the visual feature sequence to obtain a standardized visual feature sequence; Performing forward propagation processing on the standardized visual feature sequence to obtain a visual feature sequence after forward propagation; The visual feature sequence after the forward propagation and the mixed visual feature are added together to obtain the visual token feature sequence.

8. The multimodal face anti-counterfeiting recognition method according to claim 1, characterized in that: The step of obtaining a prompt token feature sequence according to the mixed concatenation feature and the prompt token sequence includes: Performing differential convolution processing on the prompt token sequence to obtain a prompt token sequence after convolution processing; Performing differential convolution processing on the mixed splicing features to obtain mixed splicing features after convolution processing; Performing concatenation processing on the prompt token sequence after the convolution processing and the mixed concatenation feature after the convolution processing to obtain the prompt concatenation feature; The prompt concatenated features are processed through an attention mechanism to obtain the prompt token feature sequence.

9. A multi-modal face anti-counterfeiting recognition device, characterized in that: Applied to the multimodal face anti-counterfeiting recognition method according to any one of claims 1 to 8, the device comprises: An acquisition module, used to acquire a modality prompt set, a task prompt, and a multimodal face image; the modality prompt set includes visual prompts related to the image modality, the task prompt is a visual prompt associated with face anti-counterfeiting recognition, and the multimodal face image includes multiple face images belonging to the same user, and the modalities of each face image are different; The first processing module is used to process the modal prompt set, the task prompt and the multimodal face image to obtain a visual token sequence and a prompt token sequence; the visual token sequence includes a classification token and multiple feature tokens, the classification token is a global tag of the visual transformer, and the feature token is a vector representation obtained by mapping pixel blocks in the multimodal face image; the prompt token sequence includes multiple prompt tokens, and the prompt token is a global tag associated with face anti-counterfeiting recognition; The second processing module is used to perform face anti-counterfeiting recognition based on the visual token sequence and the prompt token sequence in combination with a pre-trained neural network model to obtain an anti-counterfeiting recognition result.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement the multimodal face anti-counterfeiting recognition method as described in any one of claims 1-8 when executed by the processor.