Multi-modal face anti-counterfeiting detection method and device based on cross-modal fusion, equipment and medium

CN117437677BActive Publication Date: 2026-09-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311398873.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2026-09-25
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

但是,由于人脸识别系统容易受到各种类型的呈现攻击,如通过打印照片、视频重播和3D面具等方式呈现攻击,使得人脸识别技术的安全问题日益凸显

Benefits of technology

[0050]本发明的有益效果在于:本发明通过获取包括RGB人脸图像、红外IR人脸图像和深度人脸图像的多模态人脸图像,并对多模态人脸图像进行多尺度特征提取得到多尺度特征图,从而将不同模态的多尺度特征图进行跨模态特征融合处理得到跨模态特征图,再将跨模态特征图进行拼接处理,可以得到更全面、更多尺度、更具互补性的目标跨模态特征图,进而可以根据目标跨模态特征图,判断多模态人脸图像是否为真实人脸图像,有利于提升人脸防伪技术的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117437677B_ABST
    Figure CN117437677B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-modal face anti-counterfeiting detection method, device, equipment and medium based on cross-modal fusion, belong to face anti-counterfeiting technical field.The detection method is processed by including RGB, IR and depth in multi-modal face image, to obtain the cross-modal feature map of fusion feature, specifically, multi-scale feature extraction is carried out to RGB, IR, depth face image to obtain multi-scale feature map, then the cross-modal feature fusion processing of different modal multi-scale feature map is obtained cross-modal feature map, based on cross-modal feature map is spliced and is obtained target cross-modal feature map, whether multi-modal face image is real face image according to target cross-modal feature map can be judged.The present application can obtain comprehensive, multi-scale, complementary target cross-modal feature map by the cross-modal feature fusion processing of face image between different modalities, is conducive to improving the accuracy of face anti-counterfeiting technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of face anti-counterfeiting technology, and relates to a multimodal face anti-counterfeiting detection method, device, equipment and medium based on cross-modal fusion. Background Technology

[0002] With the development of technology, the application scope of facial recognition technology is constantly expanding, and it is now widely used in fields such as intelligent security, e-commerce, medical care, and education. However, because facial recognition systems are vulnerable to various types of presentation attacks, such as those using printed photos, video replays, and 3D masks, the security issues of facial recognition technology are becoming increasingly prominent. To improve the ability of facial recognition systems to resist spoofing attacks, research on facial anti-spoofing technology has gradually increased in recent years. Therefore, how to effectively improve the accuracy of facial anti-spoofing technology has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a face anti-spoofing recognition technology that combines multiple modal face images to effectively improve the accuracy of face anti-spoofing technology.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] Option 1: A multimodal face anti-spoofing detection method based on cross-modal fusion, the method comprising the following steps:

[0006] S10. Acquire multimodal face images, including RGB face images, infrared (IR) face images, and depth face images;

[0007] S20. Perform multi-scale feature extraction on RGB face images, infrared IR face images, and depth face images respectively to obtain RGB multi-scale feature maps, IR multi-scale feature maps, and depth multi-scale feature maps;

[0008] S30. Perform cross-modal feature fusion processing on the RGB multi-scale feature map and the IR multi-scale feature map to obtain the first cross-modal feature map; perform cross-modal feature fusion processing on the RGB multi-scale feature map and the depth multi-scale map features to obtain the second cross-modal feature map; perform cross-modal feature fusion processing on the depth multi-scale feature map and the IR multi-scale feature map to obtain the third cross-modal feature map.

[0009] S40. The first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map are concatenated to obtain the target cross-modal feature map;

[0010] S50. Based on the target cross-modal feature map, determine whether the multimodal face image is a real face image.

[0011] Further, in step S20, for any modality face image, the same multi-scale feature extraction method is used to obtain the multi-scale feature map of that modality, including the following steps:

[0012] S21. Perform residual processing on a certain modality face image to obtain the first residual feature map;

[0013] S22. Perform multi-scale feature extraction on the first residual feature map to obtain a set of reference multi-scale feature maps;

[0014] S23. Calculate the channel attention weights for each reference multi-scale feature map in the reference multi-scale feature map set to obtain the first channel attention weight set.

[0015] S24. Normalize each first-channel attention weight in the first-channel attention weight set to obtain the second-channel attention weight set.

[0016] S25. Perform channel multiplication on the reference multi-scale feature map set and the second channel attention weight set to obtain the multi-scale feature map.

[0017] Specifically, step S22 is as follows:

[0018] S221. Extract the first local scale features from the first residual feature map to obtain the first scale feature map;

[0019] S222. Extract the second local scale features from the first residual feature map to obtain the second scale feature map;

[0020] S223. Extract the first global scale features from the first residual feature map to obtain the third scale feature map;

[0021] S224. The first-scale feature map, the second-scale feature map, and the third-scale feature map are spliced ​​together to obtain a reference multi-scale feature map set.

[0022] Step S25 specifically includes:

[0023] S251. Randomly obtain a reference scale feature map from the set of reference multi-scale feature maps, and denote it as the first reference scale feature map;

[0024] S252. Obtain the first target channel attention weights that match the first reference scale feature map from the second channel attention weight set;

[0025] S253. Multiply the first reference scale feature map with the first target channel attention weight to obtain the first feature map;

[0026] S254. From the set of reference multi-scale feature maps, arbitrarily select a reference scale feature map other than the first reference scale feature map, and denote it as the second reference scale feature map;

[0027] S255. Obtain the second target channel attention weights that match the second reference scale feature map from the second channel attention weight set;

[0028] S256. Multiply the second reference scale feature map with the second target channel attention weight to obtain the second feature map;

[0029] S257. The first feature map and the second feature map are concatenated to obtain a multi-scale feature map.

[0030] Furthermore, in step S30, the process of performing cross-modal feature fusion processing on the multi-scale feature maps of any two modalities to obtain a cross-modal feature map includes:

[0031] S31. For any two selected modes, denoted as mode I and mode II, the multi-scale feature maps under mode I and mode II are spliced ​​together to obtain the first spliced ​​feature map.

[0032] S32. Calculate the channel attention weights corresponding to the two selected modalities based on the first spliced ​​feature map, and denot them as the channel attention weight of modality I and the channel attention weight of modality II, respectively.

[0033] S33. The first modality I enhanced feature map is obtained by weighting the multi-scale feature map of modality I through the attention weight of modality I channel;

[0034] S34. Perform cross-modal fusion of the enhanced feature map of mode I and the multi-scale feature map of mode II to obtain the cross-modal feature map of mode II;

[0035] S35. The first modality II enhanced feature map is obtained by weighting the multi-scale feature map of modality II through the attention weight of modality II channel;

[0036] S36. Perform cross-modal fusion of the enhanced feature map of mode II and the multi-scale feature map of mode I to obtain the cross-modal feature map of mode I;

[0037] S37. Aggregate the cross-modal feature maps of Mode I and Mode II to obtain cross-modal feature maps of the two modes.

[0038] Specifically, step S37 is as follows:

[0039] S371. The modality I cross-modal feature map and the modality II cross-modal feature map are concatenated to obtain the second concatenated feature map;

[0040] S372. Calculate the spatial attention weights of modality I and modality II based on the second spliced ​​feature map;

[0041] S373. Perform a dot product operation between the spatial attention weights of modality I and the multi-scale feature map of modality I to obtain the enhanced feature map of modality I; perform a dot product operation between the spatial attention weights of modality II and the multi-scale feature map of modality II to obtain the enhanced feature map of modality II;

[0042] S374. The cross-modal feature maps of mode I and mode II are obtained by fusing the enhanced feature maps of mode I and mode II.

[0043] Further, step S50 specifically involves: determining a first feature value of the multimodal face image based on the target cross-modal feature map; determining a target probability based on the first feature value; if the target probability is greater than a preset probability threshold, then the multimodal face image is a real face image.

[0044] Option 2: A multimodal face anti-spoofing detection device based on cross-modal fusion for use with the method described in Option 1. The device includes an acquisition module, a processing module, and a judgment module.

[0045] The acquisition module is used to acquire multimodal face images;

[0046] The processing module is used to extract multi-scale features from multimodal face images to obtain multi-scale feature maps, and to perform cross-modal feature fusion processing on the multi-scale feature maps of any two modalities to obtain a first cross-modal feature map, a second cross-modal feature map, and a third cross-modal feature map. Then, the target cross-modal feature map is obtained by concatenating the first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map.

[0047] The judgment module is used to determine whether a multimodal face image is a real face image based on the target cross-modal feature map.

[0048] Option 3: A computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor. Wherein, when the processor executes the computer program, it implements the multimodal face anti-spoofing detection method based on cross-modal fusion as described in Option 1.

[0049] Option 4: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal face anti-spoofing detection method based on cross-modal fusion as described in Option 1.

[0050] The beneficial effects of this invention are as follows: This invention acquires multimodal face images including RGB face images, infrared IR face images, and depth face images, and performs multi-scale feature extraction on the multimodal face images to obtain multi-scale feature maps. Then, it performs cross-modal feature fusion processing on the multi-scale feature maps of different modalities to obtain cross-modal feature maps. After stitching the cross-modal feature maps, a more comprehensive, multi-scale, and more complementary target cross-modal feature map can be obtained. In this way, based on the target cross-modal feature map, it can be determined whether the multimodal face image is a real face image, which is beneficial to improving the accuracy of face anti-spoofing technology.

[0051] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0053] Figure 1 This is a flowchart illustrating the multimodal face anti-spoofing detection method based on cross-modal fusion provided in an embodiment of the present invention.

[0054] Figure 2 This is another flowchart illustrating the multimodal face anti-spoofing detection method based on cross-modal fusion provided in this embodiment of the invention;

[0055] Figure 3 This is a schematic diagram illustrating the processing flow of each module in the multi-scale attention feature extraction process;

[0056] Figure 4 This is a schematic diagram of the processing flow of each module in the cross-modal feature fusion process. Detailed Implementation

[0057] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0058] like Figure 1As shown, this is a multimodal face anti-spoofing detection method based on cross-modal fusion provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0059] S10. Acquire multimodal face images, such as RGB face images, infrared (IR) face images, depth face images, etc.

[0060] Multimodal face images can be used to indicate face images in multiple modalities. Each pixel in an RGB face image can contain three components (also known as three channels): red, green, and blue. An RGB face image can be understood as an image with color information. Infrared (IR) face images are obtained by detecting the heat radiated outward from the object being photographed. IR face images have characteristics such as high recognition accuracy, low-light night vision, and all-weather adaptability. Depth face images, also known as distance face images, can be understood as images where the distance (depth) from the image acquisition device to various points in the scene is used as pixel values. They can reflect the geometry of visible surfaces in the image.

[0061] S20. Perform multi-scale feature extraction on the RGB face image to obtain an RGB multi-scale feature map, perform multi-scale feature extraction on the IR face image to obtain an IR multi-scale feature map, and perform multi-scale feature extraction on the depth face image to obtain a depth multi-scale feature map.

[0062] The multi-scale feature map can be used to indicate the feature maps obtained after sampling an image at different granularities. Generally speaking, sampling an image at a smaller and denser granularity reveals more image details, while sampling an image at a larger and sparser granularity reveals the overall trend of the image. Therefore, a multi-scale feature map can include global scale feature maps with multiple different overall scale features and local scale feature maps with multiple different local features.

[0063] The methods for multi-scale feature extraction of RGB, IR, and depth face images are the same. Therefore, this embodiment takes the multi-scale feature extraction of RGB face images to obtain RGB multi-scale feature maps as an example to illustrate the specific method of multi-scale feature extraction:

[0064] S21. Perform residual processing on the RGB face image to obtain a first residual feature map; the first residual feature map can be used to indicate the feature map obtained after residual processing, which is obtained through conventional feature extraction methods.

[0065] S22. Perform multi-scale feature extraction on the first residual feature map to obtain a set of reference multi-scale feature maps;

[0066] S221. Extract the first local scale features from the first residual feature map to obtain the first scale feature map;

[0067] S222. Extract the second local scale features from the first residual feature map to obtain the second scale feature map;

[0068] S223. Extract the first global scale feature from the first residual feature map to obtain the third scale feature map;

[0069] S224. Based on the first-scale feature map, the second-scale feature map, and the third-scale feature map, a set of reference multi-scale feature maps is obtained.

[0070] In the above process, local scale features can be extracted from the first residual feature map using a smaller convolutional kernel, while global scale features can be extracted from the first residual feature map using a larger convolutional kernel. After scale feature extraction, the feature maps of each scale can be concatenated in channel order to obtain a reference multi-scale feature map. For example, the feature maps can be concatenated sequentially in the order of the first scale feature map, the second scale feature map, and the third scale feature map to obtain the reference multi-scale feature map.

[0071] During concatenation, feature maps at each scale can be truncated according to the number of convolutional kernels (i.e., the number of channels), so that the size of the resulting reference multi-scale feature map is the same as the size of the first residual feature map. For example, if the first residual feature map is convolved with four kernels of different sizes, then one-quarter of each of the four scale feature maps can be truncated, and these four truncated feature maps can be concatenated to obtain the reference multi-scale feature map.

[0072] S23. Calculate channel attention weights for each reference multi-scale feature map in the reference multi-scale feature map set to obtain a first channel attention weight set; wherein, each first channel attention weight in the first channel attention weight set is related to each reference multi-scale feature map in the reference multi-scale feature map set. Figure 1 One-to-one correspondence.

[0073] S24. Normalize each first-channel attention weight in the first-channel attention weight set to obtain the second-channel attention weight set; wherein each second-channel attention weight in the second-channel attention weight set corresponds one-to-one with each first-channel attention weight in the first-channel attention weight set.

[0074] S25. Perform channel multiplication on the reference multi-scale feature map set and the second channel attention weight set to obtain the RGB multi-scale feature map.

[0075] S251. Randomly obtain a reference scale feature map from the set of reference multi-scale feature maps, and denote it as the first reference scale feature map;

[0076] S252. Obtain the first target channel attention weights that match the first reference scale feature map from the second channel attention weight set;

[0077] S253. Multiply the first reference scale feature map with the first target channel attention weight to obtain the first feature map;

[0078] The first feature map has channel attention features corresponding to the first reference scale feature map. The channel attention features can be obtained by calculating the first target channel attention weight. Thus, during the multiplication operation of the first reference scale feature map and the first target channel attention weight, the first reference scale feature map and the channel attention features can be combined to obtain the first feature map with channel attention features.

[0079] S254. From the set of reference multi-scale feature maps, arbitrarily select a reference scale feature map other than the first reference scale feature map, and denote it as the second reference scale feature map;

[0080] S255. Obtain the second target channel attention weights that match the second reference scale feature map from the second channel attention weight set;

[0081] S256. Multiply the second reference scale feature map with the second target channel attention weight to obtain the second feature map;

[0082] S257. The first feature map and the second feature map are concatenated in channel order to obtain an RGB multi-scale feature map. During the concatenation, each multi-scale feature map is truncated according to the number of convolution kernels (i.e., the number of channels) so that the size of the concatenated RGB multi-scale feature map is the same as the size of the first residual feature map.

[0083] In the above process, the channel attention weight calculation method can be implemented through the SE (where S stands for Squeeze compression and E stands for Excitation) attention module. Specifically, each reference multi-scale feature map in the reference multi-scale feature map set is used as the input feature map of the SE attention module, so that the SE attention module can compress the spatial information of the input feature map and combine the learned channel attention information with the input feature map to obtain a feature map with channel attention, namely the RGB multi-scale feature map mentioned above.

[0084] S30. Perform cross-modal feature fusion processing on the RGB multi-scale feature map and the IR multi-scale feature map to obtain the first cross-modal feature map.

[0085] S40. Perform cross-modal feature fusion processing on the RGB multi-scale feature map and the depth multi-scale feature map to obtain the second cross-modal feature map.

[0086] S50. Perform cross-modal feature fusion processing on the depth multi-scale feature map and the IR multi-scale feature map to obtain the third cross-modal feature map.

[0087] In steps S30-S50, the cross-modal feature map can be used to indicate a feature map containing features of face images from different modalities. Face images from different modalities can have strong complementarity, but simply assembling heterogeneous information from face images from different modalities may result in the features between face images not being well fused. Therefore, this embodiment adopts a cross-modal feature fusion strategy so that the final obtained cross-modal feature map can include more complementary fused features. The cross-modal feature fusion strategy is based on the feature separation and aggregation system (SA-Gate). For example, by utilizing the advantages of SA-Gate in terms of channel and spatial correlation, face images from multiple modalities are fused pairwise to obtain multiple cross-modal fused features. These cross-modal fused features are then fed into subsequent modules for learning, thereby obtaining more complementary features, which is beneficial to improving the accuracy of face anti-spoofing technology.

[0088] This embodiment takes step S30 as an example to illustrate the specific process of obtaining cross-modal feature maps through a cross-modal feature fusion strategy:

[0089] S31. The RGB multi-scale feature map and the IR multi-scale feature map are concatenated to obtain the first concatenated feature map.

[0090] S32. Perform global average pooling on the first stitched feature map. The RGB multi-scale feature map and the IR multi-scale feature map can both be understood as features composed of multiple feature values. Therefore, global average pooling is to perform an average operation on multiple feature values ​​to obtain the fused feature value.

[0091] Then, the attention weights of the RGB channels and the IR channels are calculated based on the fused feature values ​​by training the multi-layer perceptron (MLP) artificial neural network model; alternatively, the attention weights of the RGB channels and the IR channels can also be determined by the aforementioned SE attention module.

[0092] S33. Determine the first RGB enhanced feature map based on the RGB channel attention weights and the RGB multi-scale feature map;

[0093] The first RGB enhanced feature map is used to indicate the feature map obtained after weighting the RGB multi-scale feature map by RGB channel attention weights;

[0094] S34. Based on the first RGB enhanced feature map and the IR multi-scale feature map, obtain the IR cross-modal feature map;

[0095] Among them, the IR cross-modal feature map (which can be denoted as IR cross-modal feature) Figure 1 This is used to indicate the feature map obtained after cross-modal fusion of the IR multi-scale feature map and the first RGB enhanced feature map;

[0096] S35. Determine the first IR enhancement feature map based on the IR channel attention weights and the IR multi-scale feature map;

[0097] The first IR enhanced feature map is used to indicate the feature map obtained after weighting the IR multi-scale feature map by the IR channel attention weight;

[0098] S36. Based on the first IR enhanced feature map and the RGB multi-scale feature map, obtain the RGB cross-modal feature map;

[0099] Among them, the RGB cross-modal feature map (which can be denoted as RGB cross-modal feature) Figure 1 This is used to indicate the feature map obtained after cross-modal fusion of the RGB multi-scale feature map and the first IR enhanced feature map;

[0100] S37. Aggregate the IR cross-modal feature map and the RGB cross-modal feature map to obtain the first cross-modal feature map:

[0101] S371. The IR cross-modal feature map and the RGB cross-modal feature map are concatenated to obtain the second concatenated feature map;

[0102] S372. Based on the second stitched feature map, determine the IR spatial attention weight and RGB spatial attention weight. For example, the spatial attention weight of the face image can be obtained through a 1×1 convolutional layer, and then the spatial attention weight can be normalized by the softmax function to obtain the IR spatial attention weight and RGB spatial attention weight.

[0103] S373. Perform a dot product operation between the IR spatial attention weights and the IR multi-scale feature map to obtain the second IR enhanced feature map;

[0104] S374. Perform a dot product operation between the RGB spatial attention weights and the RGB multi-scale feature map to obtain the second RGB enhanced feature map.

[0105] S375. The second IR enhanced feature map and the second RGB enhanced feature map are fused to obtain the first cross-modal feature map.

[0106] In steps S40 and S50, the steps of obtaining cross-modal feature maps through cross-modal feature fusion processing can both be performed in the manner described in step S30, and will not be repeated in this embodiment.

[0107] S60. The first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map are concatenated to obtain the target cross-modal feature map.

[0108] The target cross-modal feature map can be used to indicate the feature map obtained by concatenating the first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map. The concatenation of the first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map can be achieved through a fully connected layer.

[0109] S70. Based on the target cross-modal feature map, determine whether the multimodal face image is a real face image:

[0110] S71. Determine the first feature value of the multimodal face image based on the target cross-modal feature map;

[0111] The target cross-modal feature map is processed through multiple residual blocks, then global average pooling (GAP) and fully connected (FC) to obtain the first feature value. This first feature value is the feature obtained after image processing of the multimodal face image, which includes more comprehensive, multi-scale and more complementary features of the multimodal face image.

[0112] S72. The first feature value is processed by the softmax function to determine the target probability. The target probability is used to indicate the probability that the multimodal face image is a real face image.

[0113] S73. When the target probability is greater than a preset probability threshold, determine the multimodal face image as a real face image.

[0114] like Figure 2 The diagram shown is a flowchart of a multimodal face anti-spoofing detection method based on cross-modal fusion provided by another embodiment of the present invention, including the following steps:

[0115] 1) Obtain a face image;

[0116] 2) Extract face image features using residual units to obtain the first residual feature map, the second residual feature map, and the third residual feature map;

[0117] 3) Multi-scale feature extraction is performed on the first residual feature map, the second residual feature map, and the third residual feature map through the multi-scale feature module to obtain RGB multi-scale feature map, IR multi-scale feature map, and Depth multi-scale feature map;

[0118] 4) Cross-modal feature maps are obtained by fusing RGB multi-scale feature maps, IR multi-scale feature maps, and Depth multi-scale feature maps through the cross-modal feature fusion module;

[0119] 5) The target probability is obtained by processing the cross-modal feature map through residual units, global average pooling layers, fully connected layers, and the Softmax function;

[0120] 6) Determine the authenticity of a face image based on the target probability.

[0121] For details of steps 1) to 6) above, please refer to [link / reference]. Figure 1 The detailed descriptions in the corresponding embodiments will not be repeated in this embodiment.

[0122] In the two embodiments described above, the multi-scale feature extraction method can be implemented using a multi-scale attention feature extraction module, such as... Figure 3 As shown, firstly, the input features (such as the first, second, or third residual feature map mentioned above) are passed through four different convolutional layers to obtain a multi-scale feature map. Secondly, this multi-scale feature map is passed through a Concat function, a fully connected layer, a ReLU function, a fully connected layer, and a sigmoid function to obtain the first channel attention weight. Then, the first channel attention weight is multiplied with the multi-scale feature map to obtain the multi-scale features corresponding to the input features.

[0123] Multiple convolutional layers can be understood as a multi-scale convolutional structure. This multi-scale convolutional structure can be used to extract multi-scale features from multimodal face images (such as RGB face images, IR face images, and depth face images mentioned above) to obtain multi-scale features of the multimodal face image. The multi-scale convolutional structure can include multiple convolutional layers with different kernel sizes, such as... Figure 3 The multi-scale convolutional structure shown includes four convolutional layers with different kernel sizes.

[0124] Taking multi-scale feature extraction of RGB face images as an example, the process of obtaining a reference multi-scale feature map through multi-scale feature extraction using a multi-scale convolutional structure is explained in detail:

[0125] The reference multi-scale feature map obtained by multi-scale feature extraction through multi-scale convolutional structures can be found in formula (1):

[0126] F=Conv(k,×k,G,)(X)(1)

[0127] In the formula, F i Let k represent the i-th reference multi-scale feature map. Taking a multi-scale convolutional structure containing four convolutional layers with different kernel sizes as an example, then i = 0, 1, 2, 3; i G represents the size of the i-th convolutional kernel in a multi-scale convolutional structure. i Let X represent the size of the group corresponding to the i-th convolutional kernel in the multi-scale convolutional structure, and let X represent the first residual feature map. The reference multi-scale feature map can be obtained by concatenating the channels in order, that is, by concatenating them in the order of F0, F1, F2, F3.

[0128] The process of extracting channel attention weights based on reference multi-scale feature maps to obtain the first channel attention weights corresponding to different scales can be seen in the following formula:

[0129] Z i =σ(W1δ(W0(F) i (2)

[0130] In the formula, Z i σ represents the attention weight of the i-th first channel; σ represents the activation function, such as the sigmoid activation function; W0 and W1 represent operations performed through a fully connected layer; δ represents the neural activation function, such as the ReLU (Rectified Linear Unit) activation function.

[0131] The attention weights for the second channel are obtained by calibrating the attention weights for the first channel using the Softmax function.

[0132] att i =Softmax(Z) i (3)

[0133] In the formula, att i This represents the attention weight of the i-th second channel.

[0134] The second-channel attention weights of the corresponding channels are multiplied with the corresponding reference multi-scale feature maps to obtain the various multi-scale feature maps:

[0135]

[0136] In the formula, Y i This represents the i-th multi-scale feature map, or the feature map of the multi-scale channel attention weights; This represents channel-wise multiplication; then, based on i multi-scale feature maps, an RGB multi-scale feature map is obtained:

[0137] Y'=Cat([Y0,Y,Y2,Y3]) (5)

[0138] In the formula, Y′ represents the RGB multi-scale feature map, Cat represents the operation through the fully connected layer, Y0 represents the multi-scale feature map obtained by convolution operation through convolution kernel 0, Y1 represents the multi-scale feature map obtained by convolution operation through convolution kernel 1, Y2 represents the multi-scale feature map obtained by convolution operation through convolution kernel 2, and Y3 represents the multi-scale feature map obtained by convolution operation through convolution kernel 3.

[0139] The multi-scale attention feature extraction module can integrate multi-scale spatial information and cross-channel attention into each feature map, thereby enabling better interaction between local and global channel attention information. This allows for the focus and extraction of more detailed important features from multimodal face images, enabling the face anti-spoofing process to utilize the features included in the aforementioned feature maps to identify the authenticity of face images, thus improving the accuracy of face anti-spoofing technology.

[0140] The method for extracting multi-scale features from IR face images to obtain IR multi-scale feature maps, and the method for extracting multi-scale features from depth face images to obtain depth multi-scale feature maps, can be the same as the method for extracting multi-scale features from RGB face images to obtain RGB multi-scale feature maps.

[0141] Figure 4 The process flow of each module in the cross-modal feature fusion is illustrated. Specifically: First, the RGB multi-scale features and IR multi-scale features are processed by the Concat function, MLP, and sigmoid function to obtain channel attention weights. Second, the RGB multi-scale features and IR multi-scale features are multiplied by their corresponding channel attention weights to obtain RGB enhanced features and IR enhanced features. The IR multi-scale features are then corrected based on the RGB enhanced features to obtain IR cross-modal features, and the RGB multi-scale features are also corrected based on the IR enhanced features to obtain RGB cross-modal features. The IR enhanced features and RGB enhanced features are then concatenated using the Concat function, and the concatenated features are processed by a 1x1 convolution and a sigmoid function to obtain spatial attention weights. Finally, the RGB multi-scale features and IR multi-scale features are multiplied by their corresponding spatial attention weights, and the processed features are fused to finally obtain the cross-modal features. The same approach can be used to perform cross-modal feature fusion processing on RGB multi-scale feature maps and depth multi-scale feature maps, as well as cross-modal feature fusion processing on depth multi-scale feature maps and IR multi-scale feature maps.

[0142] This example illustrates cross-modal feature fusion using RGB multi-scale feature maps and IR multi-scale feature maps; other multi-scale feature maps can be implemented in the same way.

[0143] First, the attention weights for the RGB channels and the IR channels are determined using RGB multi-scale feature maps and IR multi-scale feature maps:

[0144] W rgb =σ(MLP(GAP(F) rgb ||F ir (6)

[0145] W ir =σ(MLP(GAP(F) rgb ||F ir ))) (7)

[0146] In the formula, W rgb W represents the attention weights of the RGB channels. ir This represents the IR channel attention weights. MLP indicates the operation of calculating channel attention weights using MLP. GAP indicates the operation of global average pooling. F rgb Represents RGB multi-scale feature maps, F ir This represents the IR multi-scale feature map, and "||" indicates the splicing operation.

[0147] Secondly, the first RGB enhanced feature map is obtained based on the RGB channel attention weights and the RGB multi-scale feature map, and the first IR enhanced feature map is obtained based on the IR channel attention weights and the IR multi-scale feature map.

[0148]

[0149]

[0150] In the formula, F′ rgb F′ represents the first RGB enhanced feature map. ir This represents the first IR enhancement feature map.

[0151] Next, based on the first RGB enhanced feature map and the IR multi-scale feature map, the IR cross-modal feature map is obtained; based on the first IR enhanced feature map and the RGB multi-scale feature map, the RGB cross-modal feature map is obtained.

[0152]

[0153]

[0154] in, Represents the IR cross-modal feature map. This represents the RGB cross-modal feature map.

[0155] Then, the IR spatial attention weights are multiplied by the IR multi-scale feature map to obtain the second IR enhanced feature map, and the RGB spatial attention weights are multiplied by the RGB multi-scale feature map to obtain the second RGB enhanced feature map.

[0156]

[0157]

[0158] In the formula, F″ ir This represents the second IR enhancement feature map. C represents the attention weights in the IR space. ir This represents the convolution operation for the IR mode. Indicates the second splicing feature map; F″ rgb This represents the second RGB enhanced feature map. C represents the RGB spatial attention weights. rgb This represents a convolution operation for the RGB modality. Here, ⊙ represents element-wise multiplication, which can be understood as multiplying corresponding elements in a matrix (such as the multi-scale feature map mentioned above).

[0159] Finally, the second IR enhanced feature map and the second RGB enhanced feature map are fused to obtain the first cross-modal feature map:

[0160] M r-i =F″ rgb +F″ ir (14)

[0161] In the formula, M r-i This represents the first transmodal feature map. Similarly, the second transmodal feature map M can be obtained using the method described above. r-d and the third cross-modal feature map M d-i .

[0162] In another embodiment of the present invention, a multimodal face anti-spoofing detection device based on cross-modal fusion is provided. The device includes an acquisition module, a processing module, a judgment module, and a determination module.

[0163] The acquisition module is used to acquire multimodal face images, including RGB face images, IR face images, and depth face images.

[0164] The processing module is used to extract multi-scale features from face images of various modalities to obtain multi-scale feature maps of each modality, to perform cross-modal feature fusion processing on the multi-scale feature maps of various modalities to obtain cross-modal feature maps, to perform splicing processing on multiple cross-modal feature maps to obtain target cross-modal feature maps, and to determine whether a multi-modal face image is a real face image.

[0165] The determination module is used to determine the first feature value of the multimodal face image based on the target cross-modal feature map, determine the target probability based on the first feature value, and determine the authenticity of the multimodal face image by comparing the target probability with a preset probability threshold.

[0166] Each module in the aforementioned multimodal face anti-spoofing detection device based on cross-modal fusion can be implemented entirely or partially through software, hardware, or a combination thereof. That is, each module can be embedded in the processor of the computer device in hardware form or independent of the processor, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0167] In another embodiment of the present invention, a computer device is provided, which can be a server or a client. If it is a server, the computer device may include a processor, memory, network interface, and database connected via a system bus; if it is a client, the computer device may include a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory, wherein the non-volatile storage media stores an operating system, computer programs, and a database, and the internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for network connection communication between the server and the client. The computer program, when executed by the processor, implements... Figure 1 The method shown is a multimodal face anti-spoofing detection method based on cross-modal fusion.

[0168] In another embodiment of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements as follows: Figure 1 The method shown is a multimodal face anti-spoofing detection method based on cross-modal fusion.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multimodal face anti-spoofing detection method based on cross-modal fusion, characterized in that: The method includes the following steps: S10. Acquire multimodal face images, including RGB face images, infrared (IR) face images, and depth face images; S20. Perform multi-scale feature extraction on RGB face images, infrared IR face images, and depth face images respectively to obtain RGB multi-scale feature maps, IR multi-scale feature maps, and depth multi-scale feature maps; S30. Perform cross-modal feature fusion processing on the RGB multi-scale feature map and the IR multi-scale feature map to obtain the first cross-modal feature map; perform cross-modal feature fusion processing on the RGB multi-scale feature map and the depth multi-scale map features to obtain the second cross-modal feature map; perform cross-modal feature fusion processing on the depth multi-scale feature map and the IR multi-scale feature map to obtain the third cross-modal feature map. The process of performing cross-modal feature fusion on multi-scale feature maps of any two modalities to obtain cross-modal feature maps includes: S31. For any two selected modes, denoted as mode I and mode II, the multi-scale feature maps under mode I and mode II are spliced ​​together to obtain the first spliced ​​feature map. S32. Calculate the channel attention weights corresponding to the two selected modes based on the first spliced ​​feature map, and denot them as the channel attention weight of mode I and the channel attention weight of mode II, respectively. S33. The first modality I enhanced feature map is obtained by weighting the multi-scale feature map of modality I using the attention weight of the modality I channel; S34. Perform cross-modal fusion of the first modality I enhanced feature map and the modality II multi-scale feature map to obtain the modality II cross-modal feature map; S35. The first modality II enhanced feature map is obtained by weighting the multi-scale feature map of modality II using the modality II channel attention weights. S36. Perform cross-modal fusion of the first modality II enhanced feature map and the modality I multi-scale feature map to obtain the modality I cross-modal feature map; S37. Aggregate the cross-modal feature map of mode I and the cross-modal feature map of mode II to obtain cross-modal feature maps of the two modes; S371. The modality I cross-modal feature map and the modality II cross-modal feature map are spliced ​​together to obtain a second spliced ​​feature map; S372. Calculate the spatial attention weights of modality I and modality II based on the second spliced ​​feature map; S373. Perform a dot product operation between the spatial attention weights of modality I and the multi-scale feature map of modality I to obtain the enhanced feature map of modality I; perform a dot product operation between the spatial attention weights of modality II and the multi-scale feature map of modality II to obtain the enhanced feature map of modality II; S374. The enhanced feature map of the second mode I and the enhanced feature map of the second mode II are fused to obtain the cross-modal feature map of mode I and mode II; S40. The first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map are concatenated to obtain the target cross-modal feature map; S50. Based on the target cross-modal feature map, determine whether the multimodal face image is a real face image.

2. The method according to claim 1, characterized in that: In step S20, for any modality face image, the same multi-scale feature extraction method is used to obtain the multi-scale feature map of that modality, including the following steps: S21. Perform residual processing on a certain modality face image to obtain the first residual feature map; S22. Perform multi-scale feature extraction on the first residual feature map to obtain a set of reference multi-scale feature maps; S23. Calculate the channel attention weights for each reference multi-scale feature map in the reference multi-scale feature map set to obtain the first channel attention weight set. S24. Normalize each first-channel attention weight in the first-channel attention weight set to obtain the second-channel attention weight set. S25. Perform channel multiplication on the reference multi-scale feature map set and the second channel attention weight set to obtain the multi-scale feature map.

3. The method according to claim 2, characterized in that: Step S22 includes: S221. Extract the first local scale features from the first residual feature map to obtain the first scale feature map; S222. Extract the second local scale features from the first residual feature map to obtain the second scale feature map; S223. Extract the first global scale features from the first residual feature map to obtain the third scale feature map; S224. The first-scale feature map, the second-scale feature map, and the third-scale feature map are spliced ​​together to obtain a reference multi-scale feature map set.

4. The method according to claim 2, characterized in that: Step S25 includes: S251. Randomly obtain a reference scale feature map from the set of reference multi-scale feature maps, and denot it as the first reference scale feature map; S252. Obtain the first target channel attention weight that matches the first reference scale feature map from the second channel attention weight set; S253. Multiply the first reference scale feature map with the first target channel attention weight to obtain the first feature map; S254. Randomly select a reference scale feature map other than the first reference scale feature map from the set of reference multi-scale feature maps, and denote it as the second reference scale feature map. S255. Obtain the second target channel attention weights that match the second reference scale feature map from the second channel attention weight set; S256. Multiply the second reference scale feature map with the second target channel attention weight to obtain the second feature map; S257. The first feature map and the second feature map are concatenated to obtain a multi-scale feature map.

5. The method according to claim 1, characterized in that: Step S50 specifically involves: determining a first feature value of the multimodal face image based on the target cross-modal feature map; determining a target probability based on the first feature value; if the target probability is greater than a preset probability threshold, then the multimodal face image is a real face image.

6. A multimodal face anti-spoofing detection device based on cross-modal fusion according to any one of claims 1 to 5, characterized in that: The device includes an acquisition module, a processing module, and a judgment module; The acquisition module is used to acquire multimodal face images, which include RGB face images, infrared (IR) face images, and depth face images; The processing module is used to extract multi-scale features from multimodal face images to obtain multi-scale feature maps, and to perform cross-modal feature fusion processing on the multi-scale feature maps of any two modalities to obtain a first cross-modal feature map, a second cross-modal feature map, and a third cross-modal feature map. Then, the target cross-modal feature map is obtained by concatenating the first cross-modal feature map, the second cross-modal feature map, and the third cross-modal feature map. The judgment module is used to determine whether a multimodal face image is a real face image based on the target cross-modal feature map.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, it implements the multimodal face anti-spoofing detection method based on cross-modal fusion as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the multimodal face anti-spoofing detection method based on cross-modal fusion as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • RGB-D multi-mode fusion person detection method based on asymmetric double-flow network

    CN110956094A

  • Living body detection method and device, electronic equipment, storage medium and program product

    CN114743277A

  • Remote sensing image fusion method based on large kernel attention mechanism for multi-scale feature enhancement

    CN114936995A

  • Multi-modal sentiment classification method and device, equipment and storage medium

    CN115114408A