Face fusion attack detection method based on global features and channel shuffling attention vision transformer
By using the method of using the whole domain feature and channel shuffle attention visual transformer in face fusion attack detection, integrating spatial domain and frequency domain information, and combining channel shuffle and SE attention mechanisms, the problem of insufficient flexibility and limited accuracy of feature extraction in the existing technology is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510178508.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The existing face fusion attack detection technology has the problem of insufficient feature extraction and limited accuracy under complex scenarios and advanced attacks.
The face fusion attack detection method based on the whole domain feature and channel shuffle attention visual transformer is adopted. By integrating the spatial domain and frequency domain information of the image, combining the channel shuffle and SE attention mechanism, the whole domain features are extracted and classified.
It significantly improves the ability to identify complex features, enhances the detection accuracy and reliability of the system when handling fusion attacks, avoids information redundancy, and improves feature extraction effect.
Smart Images

Figure CN120108018A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of face recognition, and in particular to a face fusion attack detection method based on global features and channel shuffled attention visual transformer. Background Art
[0002] Face Recognition System (FRS) is a biometric identification system that authenticates a person by analyzing facial features. Due to its contactless nature, convenience, and high accuracy, the system has been widely used in many fields, such as security monitoring, electronic payment, smart device unlocking, and identity authentication;
[0003] However, an emerging identity theft method, the face fusion attack, poses a significant threat to the security of face recognition systems by fusing two or more facial images to generate a single image that contains the features of the original images. This fused facial image is similar to the original image in appearance and biometrics, which can deceive face recognition systems. This attack is particularly severe in the issuance and verification of electronic travel documents. In many countries today, passport applications allow the submission of facial images in analog or digital form. An attacker (such as a blacklisted criminal) can fuse his or her own facial features with those of an accomplice to generate a fused image that is difficult to detect. Because the fused image is highly similar to the attacker, the attacker can use a forged electronic machine-readable travel document (eMRTD) to pass border control, which will seriously threaten national and public security.
[0004] There are currently two different techniques for fusion attack detection (i) No-reference fusion attack detection technique (ii) Differential fusion attack detection technique. In the no-reference fusion detection technique, the image is analyzed individually without any reference and then classified as a real image or a fused image. In differential fusion attack detection, the analysis is performed by comparing the obtained reference image with a stored-based reference image. In addition, depending on the type of processing of the image data, no-reference fusion attack detection can be of two types (a) Print-scan attack detection, where the captured digital photo is then printed and handed over to the passport issuing center, where it is digitized again using a scanning device and subsequently stored in the eMRTD. (b) Digital attack detection, where the digitally captured face can be directly used to detect fusion attacks. Because digital passport photos are used in many countries to renew passport applications, the detection of digital face fusion attacks is focused on.
[0005] The most popular fusion attack detection techniques in digital attack detection can be roughly divided into three categories: texture-based methods, image quality-based methods, and deep learning-based methods. Texture-based methods perform detection by analyzing texture features in images, and are simple and intuitive to operate. Image quality-based methods use the noise or abnormal signals introduced in the quantized fusion process for detection, and have a certain degree of versatility. Deep learning-based methods automatically extract complex features in images by training deep neural networks, thereby achieving efficient deformation detection. Although existing methods have made significant progress in face fusion attack detection, they still have certain limitations. Texture-based methods rely too much on manually designed features, resulting in inflexible feature extraction and low accuracy. Image quality-based methods perform poorly in complex scenarios and advanced attacks, and have limited accuracy. In recent years, deep learning-based methods have received widespread attention, and have shown obvious advantages in feature extraction, but these methods usually focus on the global information of the image and ignore the tiny details in the image, resulting in limitations in identifying complex and subtle features in fused images.
[0006] To this end, technicians in this field have proposed a face fusion attack detection method based on global features and channel shuffled attention visual transformer. Summary of the invention
[0007] In order to solve the above technical problems, the present invention provides a face fusion attack detection method based on global features and channel shuffled attention visual transformer to solve the problems raised in the background technology.
[0008] The face fusion attack detection method based on global features and channel shuffled attention visual transformer includes the following steps:
[0009] S1. Preprocess the selected face, segment and normalize the face of the image according to the eye coordinates detected by the dlib landmark detector;
[0010] S2, perform color space conversion on the input RGB image X, convert it into YCbCr color space, and perform frequency domain feature extraction;
[0011] S3, performing contrast enhancement operation in color space on the input RGB image X and performing spatial domain feature extraction;
[0012] S4, after obtaining the spatial domain features and frequency domain features, fuse them to construct global features and obtain a global feature image;
[0013] S5. The global feature image is input into the designed channel shuffle attention visual transformer ShufAttViT module for feature extraction and classification. The ShufAttViT module introduces channel shuffling and SE attention mechanism based on MobileViT to capture subtle changes in facial features and perform face fusion attack detection.
[0014] Preferably, in step S1, the face of the image is segmented and normalized using the eye coordinates detected by the dlib landmark detector, and then the cropped face area is adjusted to a size of 256×256 pixels.
[0015] Preferably, the step S2 performs color space conversion on the input RGB image X, converting it into a YCbCr color space, which is expressed as:
[0016] X Ycbcr =T(X)
[0017] Among them, X Ycbc represents the converted YCbCr image, and T represents the color space conversion operation;
[0018] The YCbCr image is divided into a group of 8×8 image blocks through a sliding window, and each image block is processed using discrete cosine transform to obtain the corresponding 8×8 spectrum matrix. Each spectrum value corresponds to the intensity of different frequency components in the image block, which is expressed as:
[0019] d (i,j) =DCT(P (i,j) )
[0020] Among them, P (i,j) represents the image block in row i and column j, d (i,j) is its corresponding DCT spectrum;
[0021] The DCT spectrum d (i,j) Perform logarithmic transformation, expressed as:
[0022] d log (i,j)=log(1+|d(i,j)|)
[0023] And reassemble all processed image blocks into the enhanced frequency domain image X frq , expressed as:
[0024] X frq =Reassemble{d log (i,j)}
[0025] Reassemble means rearranging all enhanced image blocks into a complete image.
[0026] Preferably, the step S3 extracts the spatial domain features by processing the RGB image using an adaptive histogram equalization technique to enhance the local contrast, retain the brightness information and improve the image details, which can be expressed as:
[0027] X enhanced =CLAHE(X)
[0028] Among them, X enhanced Represents the image after contrast enhancement, and performs contrast enhancement operation in the color space on the input RGB image X.
[0029] Preferably, in step S4, after obtaining the spatial feature X enhanced and frequency domain feature X frq After that, the two feature maps are concatenated in the channel dimension to obtain the final global feature map Xcombined, which is expressed as:
[0030] Xcombined=[X frq ||X enhanced ]
[0031] The symbol || represents the concatenation operation in the channel dimension.
[0032] Preferably, in step S5, the global feature image is input into the designed channel shuffle attention visual transformer ShufAttViT module for feature extraction and classification, and the ShufAttViT module performs local feature extraction and performs local feature extraction on the input feature map X∈R H×W×C Apply a 3×3 convolutional layer to extract local spatial information and obtain the feature map X 1 ∈R H×W×d , expressed as:
[0033] X 1 =Conv 3×3 (X)
[0034] Among them, Conv 3×3 Represents a 3×3 convolution operation, and then the feature map X 1 Divide into N non-overlapping flat patches, and get X U ∈R P×N×d , expressed as:
[0035] X U =Unfold(X 1 )
[0036] Perform global feature extraction, apply the Transformer encoder to each patch, capture global features through the multi-head self-attention mechanism, and generate global features X G ∈R P×N×d , expressed as:
[0037] X G (p) = Transformer(X U (p)),1≤p≤P
[0038] The obtained global features XG are reassembled into a feature map X of the original image size F ∈R H×W×C , expressed as:
[0039] X F =fold(X G )
[0040] Perform feature fusion and channel shuffling, and extract the local spatial information X through the 3×3 convolution layer 1 and reassembled into the global feature map X of the original size F Concatenate and get the feature map Xconcat1∈R H×W×2C , and perform channel shuffling operation on the concatenated feature map Xconcat1, expressed as:
[0041] X concat1 =Concat(X 1 ,X F )
[0042] X shuffle =ChannelShuffle(X concat1 )
[0043] Among them, Concat represents the feature fusion and splicing operation, and ChannelShuffle represents the channel shuffling operation.
[0044] Preferably, after the channel shuffling operation, the SE attention mechanism is applied to the original feature map X to obtain an enhanced feature map X SE ∈R H×W×C , the SE attention mechanism includes global average pooling, two-layer fully connected network and Sigmoid activation function, which is expressed as:
[0045]
[0046] Among them, H and W represent the length and width of the feature map respectively; the attention weight s is generated through a two-layer fully connected network, expressed as:
[0047] s=σ(W 2 δ(W 1 z))
[0048] Among them, σ is the Sigmoid activation function, δ is the ReLU activation function, and W 1 and W2 is the weight of the fully connected layer;
[0049] The original feature map is weighted and expressed as:
[0050] X SE =X·s
[0051] The feature map X after channel shuffle shuffle The original feature map X after SE enhancement SE Perform splicing and fusion to obtain the final feature map X out ∈R H×W×(C+d) , expressed as:
[0052] X concat2 =Concat(X shuffle ,X SE )
[0053] A 3×3 convolutional layer is used to process the fused feature map to further fuse local and global information to obtain the final output feature map, which is expressed as:
[0054] X out =Conv 3×3 (X concat2 )
[0055] Among them, X out Represents the feature map of the final output.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention comprehensively integrates the spatial domain and frequency domain information of the image to fully capture the details and global features in the image. This global feature extraction strategy significantly improves the recognition ability of complex features and enhances the detection accuracy and reliability of the system when dealing with fusion attacks; and proposes a lightweight ShufAttViT module, which integrates the advanced visual transformer architecture of channel shuffling and SE attention mechanism. Based on the MobileViT module, this network avoids information redundancy by optimizing the feature fusion mechanism and strengthening the information interaction ability between channels, and improves the extraction effect of local and global features. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a block diagram of face detection based on global features and channel shuffled attention visual transformer fusion of the present invention;
[0059] Figure 2 It is a schematic diagram of the frequency domain feature extraction process of the present invention;
[0060] Figure 3 This is the ShufAttViT module diagram of the present invention. DETAILED DESCRIPTION
[0061] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0062] Embodiment: The present invention provides a face fusion attack detection method based on global features and channel shuffle attention visual transformer, comprising the following steps:
[0063] like Figure 1 The block diagram of face detection based on global features and channel shuffled attention visual transformer is shown:
[0064] S1. Preprocess the selected face, segment and normalize the face of the image using the eye coordinates detected by the dlib landmark detector, and then adjust the cropped face area to 256×256 pixels for subsequent feature extraction;
[0065] S2, extract frequency domain features, the extraction process is as follows Figure 2 As shown, it is used to reveal subtle forgery clues hidden in the frequency space. The input RGB image X is converted into the YCbCr color space, which is expressed as:
[0066] X Ycbcr= T(X)
[0067] Among them, X Ycbc represents the converted YCbCr image, and T represents the color space conversion operation;
[0068] The YCbCr image is divided into a group of 8×8 image blocks through a sliding window, and each image block is processed using discrete cosine transform to obtain the corresponding 8×8 spectrum matrix. Each spectrum value corresponds to the intensity of different frequency components in the image block, which is expressed as:
[0069] d (i,j)= DCT(P (i,j) )
[0070] Among them, P (i,j) represents the image block in row i and column j, d (i,j) is its corresponding DCT spectrum;
[0071] The DCT spectrum d (i,j) Perform logarithmic transformation, expressed as:
[0072] d log (i,j)=log(1+|d(i,j)|)
[0073] This reveals the frequency characteristics more clearly to enhance the performance of high-frequency components while retaining low-frequency information;
[0074] And reassemble all processed image blocks into the enhanced frequency domain image X frq , expressed as:
[0075] X frq =Reassemble{d log (i,j)}
[0076] Reassemble means rearranging all enhanced image blocks into a complete image. This frequency feature extraction method can more clearly reveal subtle fusion clues hidden in the image by effectively separating and strengthening frequency domain information, thereby significantly improving the accuracy and reliability of detection.
[0077] S3. Extract spatial features. The spatial features come directly from the RGB image, which contains rich color and brightness information and can reflect the visual characteristics of the image. The RGB image is processed by using the adaptive histogram equalization technology to enhance the local contrast, retain the brightness information and improve the image details. It is expressed as:
[0078] X enhanced =CLAHE(X)
[0079] Among them, X enhanced Represents the image after contrast enhancement. This processing method can better highlight the subtle structure and color differences in the image, making the spatial features more significant. The contrast enhancement operation in the color space is performed on the input RGB image X to better highlight the important features in the image.
[0080] S4, construct global features, obtain global feature images, and obtain spatial features X enhanced and frequency domain feature X frq Finally, the two feature maps are concatenated in the channel dimension to obtain the final global feature map Xcombined, which is expressed as:
[0081] Xcombined=[X frq ||X enhanced ]
[0082] The symbol || represents the concatenation operation in the channel dimension. This fusion method can fully utilize the complementarity of spatial and frequency domain features, allowing the model to consider both spatial and frequency information when conducting further feature learning. This comprehensive method can more comprehensively capture the fusion clues in the image, thereby significantly improving the accuracy and reliability of detection.
[0083] S5. Input the global feature image into the designed channel shuffle attention visual transformer ShufAttViT module for feature extraction and classification. The ShufAttViT module introduces channel shuffling and SE attention mechanism on the basis of MobileViT to further enhance the feature fusion and information interaction capabilities between channels. ShufAttViT increases the interaction between channels through channel shuffling and avoids information redundancy, while the SE attention mechanism enables the network to adaptively adjust the importance of channels, thereby significantly improving the expressiveness of the feature map, enabling the model to more accurately capture subtle facial feature changes, and improving the accuracy and robustness of face fusion attack detection;
[0084] The ShufAttViT module first extracts local features and then extracts the input feature map X∈R H×W×C Apply a 3×3 convolutional layer to extract local spatial information and obtain the feature map X 1 ∈R H×W×d , expressed as:
[0085] X 1 =Conv 3×3 (X)
[0086] Among them, Conv 3×3 Represents a 3×3 convolution operation, and then the feature map X 1 Divide into N non-overlapping flat patches, and get X U ∈R P×N×d , expressed as:
[0087] X U =Unfold(X 1 )
[0088] Perform global feature extraction, apply the Transformer encoder to each patch, capture global features through the multi-head self-attention mechanism, and generate global features X G ∈R P×N×d , expressed as:
[0089] X G (p) = Transformer(X U (p)),1≤p≤P
[0090] In order to preserve the spatial information, the obtained global features XG are reassembled into a feature map X of the original image size F ∈R H×W×C , expressed as:
[0091] X F =fold(X G )
[0092] Feature fusion and channel shuffling are performed to achieve effective fusion of local and global features. The local spatial information X extracted by the 3×3 convolutional layer 1 and reassembled into the global feature map X of the original size F Concatenate and get the feature map Xconcat1∈R H×W×2C , this fusion method makes each feature map contain both detail information and global context information, and performs channel shuffling operation on the concatenated feature map Xconcat1, which is expressed as:
[0093] X concat1 =Concat(X 1 ,X F )
[0094] X shuffle =ChannelShuffle(X concat1 )
[0095] Among them, Concat represents the feature fusion and splicing operation, and ChannelShuffle represents the channel shuffling operation, which enhances the interaction between channels by rearranging the channel order. Channel shuffling not only effectively promotes the information mixing between channels and avoids information redundancy, but also significantly improves the feature extraction capability of the network.
[0096] After the channel shuffle operation, the SE attention mechanism is applied to the original feature map X to obtain the enhanced feature map X SE ∈R H ×W×C , the SE attention mechanism includes global average pooling, two-layer fully connected network and Sigmoid activation function, which is expressed as:
[0097]
[0098] Among them, H and W represent the length and width of the feature map respectively; the attention weight s is generated through a two-layer fully connected network, expressed as:
[0099] s=σ(W 2 δ(W 1 z))
[0100] Among them, σ is the Sigmoid activation function, δ is the ReLU activation function, and W 1 and W 2 is the weight of the fully connected layer;
[0101] The original feature map is weighted and expressed as:
[0102] X SE =X·s
[0103] The feature map X after channel shuffle shuffle The original feature map X after SE enhancement SE Perform splicing and fusion to obtain the final feature map X out ∈R H×W×(C+d) , expressed as:
[0104] X concat2 =Concat(X shuffle ,X SE )
[0105] A 3×3 convolutional layer is used to process the fused feature map to further fuse local and global information to obtain the final output feature map, which is expressed as:
[0106] X out =Conv 3×3 (X concat2 )
[0107] Among them, X out Represents the feature map of the final output.
[0108] The ShufAttViT network continues to adopt the lightweight design concept and provides three different scale network configurations for mobile vision tasks - small (S), extra small (XS) and extra extra small (XXS) to meet different computing needs and application scenarios, as shown in Table 1 below. This configuration enables the ShufAttViT network to run efficiently under different hardware conditions.
[0109] Table 1: ShufAttViT network structure
[0110]
[0111]
[0112] The initial layer of the network is a standard convolutional layer with a stride of 3×3. This convolutional layer is responsible for preliminary feature extraction. After that, the network stacks MobileNetV2 (MV2) blocks and ShufAttViT blocks in sequence to gradually extract more advanced features. In the ShufAttViT module, the MobileViT module has been significantly improved. The ShufAttViT module adopts an innovative design, especially the introduction of Channel Shuffle and SE (Squeeze-and-Excitation) attention mechanisms. The purpose of these designs is to enhance the feature fusion effect and the information interaction ability between channels. The channel shuffle mechanism helps to rearrange the feature channels to improve the effect of feature fusion, while the SE attention mechanism adjusts the feature weights according to the importance of the channel, thereby improving the ability of feature expression.
[0113] Despite these innovations, the spatial dimensions of the feature maps remain multiples of 2, ensuring compatibility with the structure of the ShufAttViT module, where both the height and width of the feature maps are kept as multiples of 2 (i.e., h=w=2).
[0114] In the ShufAttViT network, the MV2 block continues to be responsible for the downsampling operation of the image, which is consistent with the MobileViT network. By introducing the ShufAttViT module, the feature representation ability of the network has been significantly improved. This means that the network can better capture local details and global information in the image, thereby performing better in visual tasks. Overall, the ShufAttViT network not only retains the high efficiency of the MobileViT network, but also further optimizes the quality of feature extraction.
[0115] From the above, it can be seen that: for fused face detection, the invention proposes a fusion face detection method of global features and channel shuffle attention visual transformer. This method can capture the macro structure and micro details of the image at the same time by combining spatial domain and frequency domain information. This comprehensive feature extraction strategy enables the model to better identify and distinguish subtle changes in complex scenes, thereby more effectively separating real facial features from fused facial features. Based on this, a lightweight visual transformer model is further designed. Through deep learning and extraction of global features, the detection ability of subtle and complex features is significantly improved, and fused faces can be effectively detected. Thereby, the person verification technology is effectively applied to airports, stations, subways, border inspection gates, customs gates, scenic spots, examination halls, communities, financial and business outlets, business units, parks, communities and other scenes with personnel identity verification needs. Integrating face detection as a key step and core link in the verification work is of great significance for identifying and verifying the real and effective identities of people entering and leaving the country, cracking down on illegal people entering and leaving the country, cracking down on forged documents and fake identities, screening out suspicious persons, and timely and accurately discovering illegal people entering and leaving the country.
[0116] Importantly, it should be noted that the construction and arrangement of the present application shown in a plurality of different exemplary embodiments are only exemplary. Although only a few embodiments are described in detail in this disclosure, it should be readily understood by those who refer to this disclosure that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in the application. Without departing from the scope of the present invention, other replacements, modifications, changes and omissions may be made in the design, operating conditions and arrangement of the exemplary embodiments. Therefore, the present invention is not limited to specific embodiments, but extends to a variety of modifications still falling within the scope of the appended claims.
[0117] Additionally, in order to provide a concise description of exemplary embodiments, all features of an actual embodiment (ie, those features that are not relevant to the best mode presently contemplated for carrying out the invention or those that are not relevant to implementing the invention) may not be described.
[0118] It will be appreciated that in the development of any actual implementation, as in any engineering or design project, numerous implementation-specific decisions may be made. Such a development effort may be complex and time-consuming, but will be a routine task of design, fabrication, and production for those of ordinary skill having the benefit of this disclosure without undue experimentation.
[0119] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A face fusion attack detection method based on global features and channel shuffled attention visual transformer, characterized in that: The following steps are involved: S1. Preprocess the selected face, segment and normalize the face of the image according to the eye coordinates detected by the dlib landmark detector; S2, perform color space conversion on the input RGB image X, convert it into YCbCr color space, and perform frequency domain feature extraction; S3, performing contrast enhancement operation in color space on the input RGB image X and performing spatial domain feature extraction; S4, after obtaining the spatial domain features and frequency domain features, fuse them to construct global features and obtain a global feature image; S5. The global feature image is input into the designed channel shuffle attention visual transformer ShufAttViT module for feature extraction and classification. The ShufAttViT module introduces channel shuffling and SE attention mechanism based on MobileViT to capture subtle changes in facial features and perform face fusion attack detection.
2. The face fusion attack detection method based on global features and channel shuffled attention visual transformer as claimed in claim 1 is characterized by: In the step S1, the face of the image is segmented and normalized using the eye coordinates detected by the dlib landmark detector, and then the cropped face area is adjusted to a size of 256×256 pixels.
3. The face fusion attack detection method based on global features and channel shuffled attention visual transformer as claimed in claim 1 is characterized by: The step S2 performs color space conversion on the input RGB image X, converting it into a YCbCr color space, which is expressed as: X Ycbcr =T(X) Among them, X Ycbc represents the converted YCbCr image, and T represents the color space conversion operation; The YCbCr image is divided into a group of 8×8 image blocks through a sliding window, and each image block is processed using discrete cosine transform to obtain the corresponding 8×8 spectrum matrix. Each spectrum value corresponds to the intensity of different frequency components in the image block, which is expressed as: d (i,j) =DCT(P (i,j) ) Among them, P (i,j) represents the image block in row i and column j, d (i,j) is its corresponding DCT spectrum; The DCT spectrum d (i,j) Perform logarithmic transformation, expressed as: d log (i,j)=log(1+|d(i,j)|) And reassemble all processed image blocks into the enhanced frequency domain image X frq , expressed as: X frq =Reassemble{d log (i,j)} Reassemble means rearranging all enhanced image blocks into a complete image.
4. The face fusion attack detection method based on global features and channel shuffle attention visual transformer as claimed in claim 1 is characterized by: The step S3 extracts spatial features by processing the RGB image using adaptive histogram equalization technology to enhance local contrast, retain brightness information and improve image details, which can be expressed as: X enhanced =CLAHE(X) Among them, X enhanced Represents the image after contrast enhancement, and performs contrast enhancement operation in the color space on the input RGB image X.
5. The face fusion attack detection method based on global features and channel shuffled attention visual transformer as claimed in claim 1 is characterized by: In step S4, the spatial domain feature X is obtained. enhanced and frequency domain feature X frq After that, the two feature maps are concatenated in the channel dimension to obtain the final global feature map Xcombined, which is expressed as: Xcombined=[X frq ||X enhanced ] The symbol || represents the concatenation operation in the channel dimension.
6. The face fusion attack detection method based on global features and channel shuffled attention visual transformer as claimed in claim 1 is characterized by: In step S5, the global feature image is input into the designed channel shuffle attention visual transformer ShufAttViT module for feature extraction and classification. The ShufAttViT module performs local feature extraction and performs local feature classification on the input feature map X∈R H×W×C Apply a 3×3 convolutional layer to extract local spatial information and obtain the feature map X1∈R H×W×d , expressed as: X1=Conv 3×3 (X) Among them, Conv 3×3 represents a 3×3 convolution operation, and then the feature map X1 is split into N non-overlapping flattened patches to obtain X U ∈R P×N×d , expressed as: X U =Unfold(X1) Perform global feature extraction, apply the Transformer encoder to each patch, capture global features through the multi-head self-attention mechanism, and generate global features X G ∈R P×N×d , expressed as: X G (p)=Transformer(X U (p)),1≤p≤P The obtained global features XG are reassembled into a feature map X of the original image size F ∈R H×W×C , expressed as: X F =fold(X G ) Perform feature fusion and channel shuffling, extract the local spatial information X1 through the 3×3 convolution layer and reassemble it into the global feature map X of the original size F Concatenate and get the feature map Xconcat1∈R H×W×2C , and perform channel shuffling operation on the concatenated feature map Xconcat1, expressed as: X concat1 =Concat(X1,X F ) X shuffle =ChannelShuffle(X concat1 ) Among them, Concat represents the feature fusion and splicing operation, and ChannelShuffle represents the channel shuffling operation.
7. The face fusion attack detection method based on global features and channel shuffled attention visual transformer as claimed in claim 6 is characterized by: After the channel shuffle operation, the SE attention mechanism is applied to the original feature map X to obtain the enhanced feature map X SE ∈R H×W×C , the SE attention mechanism includes global average pooling, two-layer fully connected network and Sigmoid activation function, which is expressed as: Among them, H and W represent the length and width of the feature map respectively; the attention weight s is generated through a two-layer fully connected network, expressed as: s=σ(W2δ(W1z)) Among them, σ is the Sigmoid activation function, δ is the ReLU activation function, W1 and W2 are the weights of the fully connected layer; The original feature map is weighted and expressed as: X SE =X·s The feature map X after channel shuffle shuffle The original feature map X after SE enhancement SE Perform splicing and fusion to obtain the final feature map X out ∈R H×W×(C+d) , expressed as: X concat2 =Concat(X shuffle ,X SE ) A 3×3 convolutional layer is used to process the fused feature map to further fuse local and global information to obtain the final output feature map, which is expressed as: X out =Conv 3×3 (X concat2 ) Among them, X out Represents the feature map of the final output.