Counterfeit face detection method and device, electronic equipment and storage medium

By performing image fusion and block feature extraction methods on the original image and frequency domain information images, the problem of insufficient face distribution and feature extraction in the prior art is solved, and the accuracy of fake face detection is improved.

CN119942657APending Publication Date: 2025-05-06DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411860010.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing fake face detection methods do not fully consider the distribution and feature extraction of faces in the image, resulting in low detection accuracy.

Method used

By obtaining the original image containing the target face, performing frequency domain transformation and frequency domain inverse transformation, a frequency domain information image is obtained, and image fused with the original image to obtain a fused feature image. Then, block feature extraction is performed on the fused feature image, and fake face detection is performed on the target face based on multiple block feature maps.

Benefits of technology

The detection accuracy of forged faces is improved, and the block-level space-frequency domain fusion feature map is supervised and predicted to more accurately determine whether the faces in the image are forged.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942657A_ABST
    Figure CN119942657A_ABST
Patent Text Reader

Abstract

The invention provides a forged face detection method and device, electronic equipment and a storage medium, and relates to the technical field of image processing. In the application, after an original image including a target face is obtained, frequency domain transformation and frequency domain inverse transformation can be performed on the original image to obtain a frequency domain information image; thirdly, performing image fusion on the original image and the frequency domain information image to obtain a fusion feature image in which the spatial domain feature and the frequency domain feature are fused; and finally, performing block feature extraction on the fused feature image to obtain a plurality of block feature maps, and performing forged face detection on the target face based on the plurality of block feature maps to determine whether the target face is a forged face or a real face. Therefore, the counterfeit detection result of the original image containing the target face is supervised and predicted on the block-level space-frequency domain fusion feature map, and the detection precision of the counterfeit face is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a forged face detection method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of deep learning technology, deep face fake technology, also known as artificial intelligence (AI) face-changing technology, has gradually matured. However, its widespread application has also brought serious security risks. Therefore, the identification of deep face fakes has become an important part of image authentication.

[0003] Currently, forged face detection is usually achieved by extracting spatial domain features and / or frequency domain features of an image containing the target object's face to capture forged traces of the face, thereby achieving forged face detection.

[0004] It can be seen that the existing forged face detection methods do not fully consider the distribution of faces in images (e.g., positions), and / or there is a problem of insufficient feature extraction of images. For example, only spatial domain and frequency domain feature extraction is performed on an image containing a target object face, and then the spatial domain features and frequency domain features obtained are used to determine whether the face in the image is forged. Therefore, the existing forged face detection methods have low detection accuracy for forged faces. Summary of the invention

[0005] The embodiments of the present application provide a forged face detection method, device, electronic device and storage medium to improve the detection accuracy of forged faces.

[0006] In a first aspect, an embodiment of the present application provides a method for detecting a fake face, the method comprising:

[0007] Acquire an original image containing a target face, and perform frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image; wherein the original image and the frequency domain information image are both spatial domain images, and the frequency domain information image includes multiple frequency domain features of the target face;

[0008] Perform image fusion on the original image and the frequency domain information image to obtain a fused feature image, and perform block feature extraction on the fused feature image to obtain multiple block feature maps; wherein different block feature maps correspond to different image regions in the fused feature image;

[0009] Based on multiple block feature maps, forged face detection is performed on the target face to determine the detection result of the target face; the detection result is a forged face or a real face.

[0010] In an optional embodiment, performing frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image includes:

[0011] Perform facial key point detection on the original image to obtain the position coordinates corresponding to multiple facial key points;

[0012] Based on multiple position coordinates and fixed coordinates respectively set for multiple facial key points, face alignment is performed on the target face to obtain a face alignment image;

[0013] Perform frequency domain transformation and inverse frequency domain transformation on the face alignment image to obtain a frequency domain information image.

[0014] In an optional embodiment, facial key point detection is performed on the original image to obtain position coordinates corresponding to a plurality of facial key points, including:

[0015] Determine the location area of ​​the target face in the original image, and crop the original image based on the location area to obtain a face area image;

[0016] Perform facial key point detection on the face area image to obtain the position coordinates corresponding to multiple facial key points.

[0017] In an optional embodiment, performing frequency domain transformation and frequency domain inverse transformation on the face alignment image to obtain a frequency domain information image includes:

[0018] Performing frequency domain transformation on the face alignment image to obtain the face alignment image after frequency domain transformation;

[0019] Based on a plurality of preset filtering frequencies, filtering is performed on the face alignment images after the frequency domain transformation respectively to obtain a plurality of face alignment images after the filtering;

[0020] Performing frequency domain inverse transformation on the multiple filtered face alignment images to obtain multiple frequency domain inverse transformed face alignment images, and obtaining a frequency domain information image based on the multiple frequency domain inverse transformed face alignment images.

[0021] In an optional embodiment, performing image fusion on the original image and the frequency domain information image to obtain a fused feature image includes:

[0022] The original image and the frequency domain information image are fused in the frequency channel dimension to obtain a fused feature image; wherein each frequency channel corresponds to a filtering frequency.

[0023] In an optional embodiment, block feature extraction is performed on the fused feature image to obtain multiple block feature maps, including:

[0024] Using spatial attention mechanism and / or channel attention mechanism, extract features from the fused feature image to obtain a global feature map of the target face;

[0025] Based on a preset number of block classifiers, the global feature map is divided into a plurality of block feature maps; wherein each block classifier is used to determine whether an image region associated with a corresponding block feature map is a forged region.

[0026] In an optional embodiment, performing forged face detection on a target face based on a plurality of block feature maps to determine a detection result of the target face includes:

[0027] Determine classification labels corresponding to the plurality of block feature maps respectively; wherein each classification label is used to indicate whether an image region associated with the corresponding block feature map is a forged region;

[0028] Based on multiple classification labels and their corresponding classification weights, the detection result of the target face is determined.

[0029] In a second aspect, an embodiment of the present application further provides a forged face detection device, the device comprising:

[0030] An image acquisition module is used to acquire an original image containing a target face, and perform frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image; wherein both the original image and the frequency domain information image are spatial domain images, and the frequency domain information image includes multiple frequency domain features of the target face;

[0031] The feature extraction module is used to perform image fusion on the original image and the frequency domain information image to obtain a fused feature image, and to perform block feature extraction on the fused feature image to obtain a plurality of block feature maps; wherein different block feature maps correspond to different image regions in the fused feature image;

[0032] The face detection module is used to perform forged face detection on the target face based on multiple block feature maps to determine the detection result of the target face; the detection result is a forged face or a real face.

[0033] In an optional embodiment, when performing frequency domain transformation and frequency domain inverse transformation on the original image to obtain the frequency domain information image, the image acquisition module is specifically used to:

[0034] Perform facial key point detection on the original image to obtain the position coordinates corresponding to multiple facial key points;

[0035] Based on multiple position coordinates and fixed coordinates respectively set for multiple facial key points, face alignment is performed on the target face to obtain a face alignment image;

[0036] Perform frequency domain transformation and inverse frequency domain transformation on the face alignment image to obtain a frequency domain information image.

[0037] In an optional embodiment, when performing facial key point detection on the original image to obtain position coordinates corresponding to a plurality of facial key points, the image acquisition module is specifically used to:

[0038] Determine the location area of ​​the target face in the original image, and crop the original image based on the location area to obtain a face area image;

[0039] Perform facial key point detection on the face area image to obtain the position coordinates corresponding to multiple facial key points.

[0040] In an optional embodiment, when performing frequency domain transformation and frequency domain inverse transformation on the face alignment image to obtain the frequency domain information image, the image acquisition module is specifically used to:

[0041] Performing frequency domain transformation on the face alignment image to obtain the face alignment image after frequency domain transformation;

[0042] Based on a plurality of preset filtering frequencies, filtering is performed on the face alignment images after the frequency domain transformation respectively to obtain a plurality of face alignment images after the filtering;

[0043] Performing frequency domain inverse transformation on the multiple filtered face alignment images to obtain multiple frequency domain inverse transformed face alignment images, and obtaining a frequency domain information image based on the multiple frequency domain inverse transformed face alignment images.

[0044] In an optional embodiment, when the original image and the frequency domain information image are fused to obtain a fused feature image, the feature extraction module is specifically used to:

[0045] The original image and the frequency domain information image are fused in the frequency channel dimension to obtain a fused feature image; wherein each frequency channel corresponds to a filtering frequency.

[0046] In an optional embodiment, when performing block feature extraction on the fused feature image to obtain a plurality of block feature maps, the feature extraction module is specifically used to:

[0047] Using spatial attention mechanism and / or channel attention mechanism, extract features from the fused feature image to obtain a global feature map of the target face;

[0048] Based on a preset number of block classifiers, the global feature map is divided into a plurality of block feature maps; wherein each block classifier is used to determine whether an image region associated with a corresponding block feature map is a forged region.

[0049] In an optional embodiment, when performing forged face detection on a target face based on a plurality of block feature maps to determine a detection result of the target face, the face detection module is specifically used to:

[0050] Determine classification labels corresponding to the plurality of block feature maps respectively; wherein each classification label is used to indicate whether an image region associated with the corresponding block feature map is a forged region;

[0051] Based on multiple classification labels and their corresponding classification weights, the detection result of the target face is determined.

[0052] In a third aspect, an embodiment of the present application further provides an electronic device, including:

[0053] Processor; and

[0054] Memory for storing programs,

[0055] The program includes instructions, and when the instructions are executed by the processor, the processor executes the forged face detection method as described in the first aspect.

[0056] In a fourth aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the forged face detection method as described in the first aspect.

[0057] In a fifth aspect, the present application provides a computer program product, which, when called by a computer, enables the computer to execute the steps of the forged face detection method as described in the first aspect.

[0058] The beneficial effects of this application are as follows:

[0059] In the forged face detection method provided in the embodiment of the present application, after obtaining the original image containing the target face, the original image can be subjected to frequency domain transformation and frequency domain inverse transformation to obtain a frequency domain information image; then, the original image and the frequency domain information image are subjected to image fusion to obtain a fused feature image that is a fusion of spatial domain features and frequency domain features; finally, block feature extraction is performed on the fused feature image to obtain a plurality of block feature maps, and forged face detection is performed on the target face based on the plurality of block feature maps to determine whether the target face is a forged face or a real face.

[0060] It can be seen that supervising and predicting the forged detection results of the original image containing the target face on the block-level spatial-frequency domain fusion feature map improves the detection accuracy of forged faces.

[0061] In addition, other features and advantages of the present application will be described in the subsequent description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described here are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0063] Figure 1 A schematic diagram of an optional system architecture applicable to the embodiments of the present application;

[0064] Figure 2 A schematic diagram of an implementation process of a forged face detection method provided in an embodiment of the present application;

[0065] Figure 3 A schematic diagram of an implementation flow of a method for acquiring a frequency domain information image provided in an embodiment of the present application;

[0066] Figure 4 A schematic diagram of the distribution of key points of a face provided in an embodiment of the present application;

[0067] Figure 5 A schematic diagram of a specific scenario for obtaining a face area image provided in an embodiment of the present application;

[0068] Figure 6 A logical schematic diagram of obtaining a frequency domain information image provided by an embodiment of the present application;

[0069] Figure 7 A schematic diagram of a specific scenario for predicting the detection result of a target face provided in an embodiment of the present application;

[0070] Figure 8 A method based on the embodiment of the present application is provided Figure 2 Schematic diagram of application scenarios;

[0071] Fig. 9 A schematic diagram of the structure of a forged face detection device provided in an embodiment of the present application;

[0072] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0073] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.

[0074] It should be understood that the various steps described in the method implementation of the present application can be performed in different orders and / or performed in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0075] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0076] It should be noted that the modifications of "one" and "plurality" mentioned in the present application are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0077] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0078] First, the design concept of the embodiment of the present application is briefly introduced below:

[0079] Deep fake face technology can make it extremely difficult to distinguish between false content (e.g., fake faces) and real content (e.g., real faces) by performing precise operations on multimedia data such as images containing faces. This also makes the widespread application of deep fake face technology bring serious security risks. Therefore, deep fake face detection has become a crucial issue in the current field of face recognition. Exemplarily, the current deep fake face detection methods can include but are not limited to the following 5 methods, namely A1 to A5:

[0080] A1: By extracting spatial domain features and frequency domain features, and using a gating mechanism (i.e., channel attention mechanism) to adaptively fuse shallow low-frequency features and deep frequency domain features, deep face fake detection can be achieved.

[0081] A2: For spatial domain images, a convolutional neural network (CNN) model and a vision transformer (VIT) model are combined to achieve deep fake face detection and improve generalization capabilities.

[0082] A3: Combine spatial domain features and frequency domain features, and use the spatial-channel attention mechanism to fuse features, so as to achieve deep fake face detection.

[0083] A4: Detect deep fake faces in patches in spatial domain images.

[0084] A5: Frequency domain deepfake detection on face images.

[0085] It is not difficult to find that although the above method A1 takes into account both spatial domain features and frequency domain features and classifies facial images, in fact, a fake facial image is usually only faked in some areas. For example, only the eyes, mouth or the entire face may be faked, while the hair, background and other areas are real images. Therefore, the above method A1 only outputs a true or false judgment probability for the overall feature, which violates the overall true or false feature distribution of the image.

[0086] Although the above method A2 adopts the CNN model and the VIT model to improve the generalization ability of the model, thereby improving the accuracy of deep fake face detection for spatial domain images, it ignores other domain features except the spatial domain features.

[0087] Similar to the above method A1, the above method A3 also fails to take into account that only a part of the area of ​​a fake face image is usually faked. Although the above method A4 performs block area fake detection on spatial domain images, it ignores the frequency domain features. In addition, although the above method A5 can perform frequency domain deep fake detection on fake face images, it does not combine spatial domain features.

[0088] In summary, existing deep fake face detection methods do not fully consider the distribution of faces in images (e.g., position, proportion, etc.), and / or have the problem of insufficient feature extraction of images. Therefore, existing deep fake face detection methods have low detection accuracy for fake faces.

[0089] In view of this, in order to solve or improve the above-mentioned problems, an embodiment of the present application proposes a method for detecting forged faces, which may specifically include: first, obtaining an original image containing a target face, and performing frequency domain transform and frequency domain inverse transform on the original image to obtain a frequency domain information image; wherein the original image and the frequency domain information image are both spatial domain images, and the frequency domain information image includes multiple frequency domain features of the target face; secondly, performing image fusion on the original image and the frequency domain information image to obtain a fused feature image, and performing block feature extraction on the fused feature image to obtain multiple block feature maps; wherein different block feature maps correspond to different image areas in the fused feature image; finally, performing forged face detection on the target face based on the multiple block feature maps to determine the detection result of the target face; the detection result is a forged face or a real face.

[0090] In this way, the original face image (i.e., spatial domain image) and the frequency domain information image are fused to obtain a fused feature image containing spatial domain features and frequency domain features. The fused feature image is then subjected to block-level feature extraction and regional authenticity discrimination to obtain the detection result of the target face, thereby improving the detection accuracy of forged faces.

[0091] In particular, the preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application, and the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.

[0092] See also Figure 1 As shown, it is a schematic diagram of a system architecture applicable to an embodiment of the present application, and the system architecture may include: a terminal device (101a, 101b) and a server 102. The terminal device (101a, 101b) and the server 102 may exchange information through a communication network, wherein the communication mode adopted by the communication network may include: a wireless communication mode and a wired communication mode. Exemplarily, the terminal device (101a, 101b) may access the network through cellular mobile communication technology and communicate with the server 102. Wherein, the cellular mobile communication technology, for example, includes the fifth generation mobile communication (5th generation mobile networks, 5G) technology or the next generation mobile communication technology. Optionally, the terminal device (101a, 101b) may access the network through a short-range wireless communication mode and communicate with the server 102. Wherein, the short-range wireless communication mode, for example, includes wireless fidelity (wireless fidelity, Wi-Fi) technology.

[0093] The embodiment of the present application does not impose any restriction on the number of communication devices involved in the above system architecture. For example, the above system architecture may include more terminal devices, or may include fewer terminal devices, or may also include other network devices. Figure 1 As shown, only the terminal devices (101a, 101b) and the server 102 are described as examples, and the above communication devices and their respective functions are briefly introduced below.

[0094] The terminal device (101a, 101b) is a device that can provide voice and / or data connectivity to users, and can be a device that supports wired and / or wireless connection.

[0095] Exemplarily, the terminal devices (101a, 101b) may include, but are not limited to: mobile phones, tablet computers, laptop computers, PDAs, mobile internet devices (MID), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminal devices in industrial control, wireless terminal devices in unmanned driving, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, or wireless terminal devices in smart homes, etc.

[0096] In addition, a related client may be installed on the terminal device (101a, 101b), and the client may be software, such as an application (APP), a browser, a short video software, etc., or a web page, a mini-program, etc.; it should be noted that the terminal device (101a, 101b) in the embodiment of the present application may enable the client related to forged face detection to send the original image containing the target face to the server 102, so as to perform subsequent forged face detection and other method steps.

[0097] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0098] It should be noted that in the embodiment of the present application, a pre-trained forged face detection model is deployed on the server 102. After the server 102 obtains the original image containing the target face, it can use the forged face detection model to detect the authenticity of the target face and determine whether the target face is a forged face or a real face. Optionally, the forged face detection model can be obtained by the server 102 through training and adjustment based on the training sample set and the test sample set.

[0099] It is worth mentioning that in an embodiment of the present application, the server 102 can be used to obtain an original image containing a target face, and perform frequency domain transform and frequency domain inverse transform on the original image to obtain a frequency domain information image; then, perform image fusion on the original image and the frequency domain information image to obtain a fused feature image, and perform block feature extraction on the fused feature image to obtain multiple block feature maps; finally, perform forged face detection on the target face based on the multiple block feature maps to determine the detection result of the target face.

[0100] The following describes the fake face detection method provided by the exemplary embodiment of the present application in combination with the above-mentioned system architecture and with reference to the accompanying drawings. It should be noted that the above-mentioned system architecture is only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect.

[0101] See also Figure 2 As shown, it is a schematic diagram of an implementation process of a forged face detection method provided in an embodiment of the present application. The execution subject takes a server as an example. The specific implementation process of the method is as follows:

[0102] S201: Acquire an original image containing a target face, and perform frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image.

[0103] Among them, the original image and the frequency domain information image are both spatial domain images, and the frequency domain information image can include multiple frequency domain features of the target face, that is, it can be a spatial domain image after filtering out background features and noise. For example, the aforementioned frequency domain information image can include frequency domain features corresponding to multiple facial key points (such as eyes) of the aforementioned target face.

[0104] Optionally, the above-mentioned original image can be a grayscale image or a color (red green blue, RGB) image. Of course, it can also be an image with any other number of color channels, which is not limited in the embodiment of the present application.

[0105] In addition, in order to facilitate the subsequent image fusion of the original image and the frequency domain information image, it can be ensured that the size of the generated frequency domain information image is the same as the size of the original image. Exemplarily, the image size of the aforementioned frequency domain information image and the image size of the aforementioned original image are both N×N, that is, the height (H) and width (W) are both N.

[0106] Furthermore, the frequency domain transform can be discrete cosine transform (DCT), or any other frequency domain transform, which is not limited in the present application. Correspondingly, the frequency domain inverse transform can be inverse DCT, or inverse transform of other frequency domain transforms.

[0107] It should also be noted that the kernel size (expressible as n×n) of the frequency domain feature extractor used in the frequency domain transformation in the embodiment of the present application is less than or equal to the image size of the aforementioned frequency domain information image and the image size of the aforementioned original image, that is, n≤N.

[0108] In an optional implementation, when executing step S201, after obtaining the original image containing the target face, the server can perform face alignment processing on the original image according to the fixed coordinates set for the preset multiple facial key points, and then perform frequency domain transformation and frequency domain inverse transformation on the original image after the face alignment processing to ensure that the target face can be subsequently detected as a forged face. Figure 3 As shown, it is a schematic diagram of an implementation flow of a method for acquiring a frequency domain information image provided in an embodiment of the present application. The execution subject is still taken as an example of a server, and the method steps are specifically as follows:

[0109] S301: Detect facial key points on the original image to obtain position coordinates corresponding to a plurality of facial key points.

[0110] Exemplarily, when executing step S301, refer to Figure 4 As shown in FIG. 1 , taking the key points at the center of the eyes, the key points at the center of the ears, and the key points at the center of the lips as examples, the server can obtain the position coordinates of the five facial key points in the reference coordinate system by detecting the five facial key points. For example, the position coordinates corresponding to the five facial key points are: (x 1 ,y 1 )、(x 2 ,y 2 )、(x 3 ,y 3 )、(x 4 ,y 4 ) and (x 5 ,y 5 ).

[0111] It should be noted that the above reference coordinate system can be a view coordinate system with the upper left corner vertex as the coordinate system origin O, the right as the positive direction of the X axis, and the downward as the positive direction of the Y axis. Of course, it can also be other coordinate reference systems.

[0112] S302: Based on the multiple position coordinates and the fixed coordinates respectively set for the multiple facial key points, perform face alignment on the target face to obtain a face alignment image.

[0113] Still taking the above five facial key points as an example, assuming that the fixed coordinates set for the above five facial key points are: (150, 150), (250, 150), (20, 170), (380, 170) and (200, 200). The server can adjust the above five facial key points from their respective position coordinates to the corresponding fixed coordinates through image transformation. Optionally, the above coordinate position adjustment method can adopt an affine transformation method.

[0114] Based on the above method, the face to be detected (i.e., the target face) is spatially normalized by the key points of the face, i.e., face alignment. In this way, it can be ensured that the features related to the shape and texture of the facial features, which are not related to the position of the facial features, can be extracted later, thereby improving the accuracy of forged face detection, as well as the performance and stability of the forged face detection method.

[0115] In order to improve the speed or efficiency of detecting facial key points, the server can first determine the position area of ​​the target face in the original image, and crop the original image based on the position area to obtain the corresponding face area image, and then perform facial key point detection on the face area image to obtain the position coordinates corresponding to multiple facial key points.

[0116] See also Figure 5 As shown in the figure, take the original image containing only one target object as an example. First, the face position detection can be performed on the original image containing the target object in combination with the face features to obtain Figure 5 Next, if the face detection frame completely contains the target face of the target object, the face detection frame can be expanded according to a preset expansion size to obtain Figure 5 The face detection area corresponding to the face detection frame shown in (B) is the face area. Figure 5 The face area shown in (B) is cropped to obtain Figure 5 The cropped image (i.e., the face area image) shown in (C) in FIG.

[0117] like Figure 5As shown in (D) in the figure, the aforementioned cropped image can also be adjusted (e.g., enlarged or reduced) to a fixed size. It should be noted that in the embodiment of the present application, there is no specific limitation on the expansion method of the above-mentioned face detection frame. For example, the above-mentioned face detection frame can also be expanded according to the intersection over union (IoU) between the detection area corresponding to the face detection frame and the image area corresponding to the target face, so as to obtain a face detection frame that can just completely cover the target face. In addition, in the embodiment of the present application, there is no specific limitation on the shape of the above-mentioned face detection frame. For example, the shape of the above-mentioned face detection frame can be circular, elliptical or rectangular, or it can be an irregular shape.

[0118] S303: Perform frequency domain transformation and frequency domain inverse transformation on the face alignment image to obtain a frequency domain information image.

[0119] Exemplarily, when executing step S303, after obtaining the face alignment image A, the server may use DCT to perform frequency domain transformation on the face alignment image A to obtain a frequency domain transformation image S, and then use DCT inverse transformation to perform frequency domain inverse transformation on the frequency domain transformation image S to obtain a frequency domain information image F. In this way, not only is the feature extraction of the face alignment image A in the spatial domain in the frequency domain achieved, but also because the frequency domain information image F obtained after the frequency domain transformation and the frequency domain inverse transformation is still an image in the spatial domain, the subsequent fusion speed of the spatial domain features and the frequency domain features is improved.

[0120] In an optional implementation, the server may perform a frequency domain transform on the face alignment image to obtain a face alignment image after the frequency domain transform (i.e., a frequency domain transform image); then, based on a preset plurality of filter frequencies (e.g., m filter frequencies), the face alignment image after the frequency domain transform is respectively filtered to obtain a plurality of filtered face alignment images (i.e., filtered frequency domain transform images); finally, perform a frequency domain inverse transform on the plurality of filtered face alignment images to obtain a plurality of frequency domain inverse transform face alignment images (i.e., filtered frequency domain transform images), and obtain a frequency domain information image based on the plurality of frequency domain inverse transform face alignment images.

[0121] The above-mentioned preset multiple filter frequencies can be different filter frequencies set for the fixed coordinates corresponding to each facial key point in the standard face image. In this way, by filtering the frequency domain transformation image with a filter of a specific filter frequency, the frequency signal of the corresponding position (i.e., the fixed coordinate) can be obtained. In other words, by filtering the frequency domain transformation image multiple times, frequency domain transformation images of different frequencies can be obtained. Since the local frequency domain feature extraction is performed on the face alignment image, the accuracy or precision of the subsequent forged face detection is also improved.

[0122] For example, see Figure 6 As shown in the figure, after inputting an original image containing a face, the server can first align the original image to obtain a face-aligned image. Assuming that the image size of the face-aligned image is N×N, if the face-aligned image is a grayscale image, the input size is 1×N×N, where 1 is the number of color channels C, and N is the height H and width W of the face-aligned image. Next, the face-aligned image can be transformed in the frequency domain, for example, using DCT frequency domain transformation, and the kernel size of the DCT frequency domain feature extractor is n×n, where n satisfies the condition n≤N. Thus, the frequency domain transformation graph S is obtained.

[0123] Furthermore, the server may filter the frequency domain transformation image S multiple times to obtain frequency domain transformation images of different frequencies. For example, the frequency domain transformation image S may be filtered m times, where m≤n. The i-th filtering retains the frequency signals at positions ((i-1)×n / m, (i-1)×n / m) to (i×n / m, i×n / m) in the n*n region, and obtains the frequency domain transformation image S after filtering. i . Thus, m frequency domain transformed images S after filtering are obtained 1 ~S m . Furthermore, for S 1 ~S m Perform inverse DCT frequency domain transform to obtain the frequency domain information image F in the spatial domain 1 ~F m , and combine them into a frequency domain information image F of m channels.

[0124] It should be noted that the above filtering method is not specifically limited in the embodiments of the present application. For example, it can be Gaussian filtering or other filtering methods. Figure 6 The filtering method shown is only an example, and its purpose is to extract the features of different frequency bands (or frequencies). Figure 6 The filtering method shown can extract the frequency domain features on the left diagonal.

[0125] S202: performing image fusion on the original image and the frequency domain information image to obtain a fused feature image, and performing block feature extraction on the fused feature image to obtain multiple block feature maps.

[0126] Among them, different block feature maps correspond to different image regions in the fused feature image. In other words, each block feature map contains: fused features of frequency domain features and spatial domain features of a part of the image region in the fused feature image.

[0127] Specifically, when executing step S202, after obtaining the frequency domain information image F, the server can fuse the frequency domain information image F already in the spatial domain with the original image Y. For example, the aforementioned image fusion method can be vector concatenation (Concat) along the channel direction. Optionally, the aforementioned along the channel direction can be along the color channel direction.

[0128] In an optional implementation, since both the frequency domain information image F and the original image Y include frequency domain features corresponding to multiple different frequencies, the server can perform image fusion processing on the original image Y and the frequency domain information image F in the frequency channel dimension to obtain a fused feature image, wherein each frequency channel corresponds to a filter frequency. In this way, the feature fusion of the frequency domain information image and the original image is enhanced, thereby improving the accuracy of subsequent forged face detection.

[0129] It should be noted that the above-mentioned fused feature image can also be called a space-frequency domain fused feature image or a space-frequency domain feature image. Of course, it can also have other names, which is not limited in the embodiments of the present application.

[0130] Furthermore, after obtaining the fused feature image obtained by fusing the spatial domain features and the frequency domain features, the server can use various classic convolutional neural networks to perform deep image feature extraction on the fused feature image. In order to fully consider that the forged face image is usually only forged in part of the area, the server can perform block feature extraction on the fused feature image to obtain multiple block feature maps, and then determine the detection result of the target face based on the multiple block feature maps.

[0131] Optionally, the above-mentioned convolutional neural network may include but is not limited to: VGG, GoogleNet, Xception, ResNet, ResNeSt, MobileNet, ShuffleNet and other feature extraction networks, which are not specifically limited in the embodiments of the present application.

[0132] In order to enhance the extraction of effective features, an attention mechanism, such as a spatial attention mechanism and / or a channel attention mechanism, can also be introduced in the process of extracting block features from the fused feature image. Specifically, assuming that the size of the input feature map (i.e., the fused feature image) is C×H×W, the channel attention vector size can be C×1×1, which is used to selectively focus on more important channel features at the channel level. Among them, C represents the number of channels of the input feature map, and H×W represents the image size of the input feature map. The spatial attention mechanism vector size can be 1×H×W, which is used to selectively focus on the features of each channel at the spatial level. Finally, a block feature map that fuses spatial domain features and frequency domain features is obtained.

[0133] Therefore, in an optional implementation, after obtaining the fused feature image, the server can use a spatial attention mechanism and / or a channel attention mechanism to extract features from the fused feature image to obtain a global feature map of the target face; and then divide the global feature map into multiple block feature maps based on a preset number of block classifiers. Each block classifier can be used to determine whether an image region associated with a corresponding block feature map is a forged region.

[0134] For example, assuming that the number of the preset block classifiers is K×K, the server may divide the global feature map into K×K block feature maps, each block feature map corresponding to a block classifier. Optionally, in order to improve the accuracy of true and false image detection on the block feature map, multiple block classifiers may be used to perform true and false image detection on the same block feature map.

[0135] Any classifier K among the above K×K block classifiers can be obtained by training based on the sample image set and the corresponding supervisory signal. The aforementioned supervisory signal is used to indicate whether the sample image area corresponding to the block classifier K is a forged image. Exemplarily, assuming that the sample image area input to the block classifier K is a forged image, a "fake" supervisory signal can be assigned to the sample image area; conversely, assuming that the sample image area input to the block classifier K is a real image, a "real" supervisory signal can be assigned to the sample image area. Thus, the block classifier K is iteratively trained for multiple times until the authenticity detection loss value of the block classifier K for the sample image area meets the set convergence condition.

[0136] Optionally, the manner in which the supervisory signal is assigned to each block region during the training phase can also be freely defined, and the present application embodiment does not limit this. For example, fake supervisory signals can be assigned to all face regions where there are partial fake face images.

[0137] In order to reduce the computational complexity of the convolutional neural network, multiple downsampling processes can be performed in the process of extracting features from the fused feature image and obtaining the global feature map of the target face, thereby reducing the dimension of the fused feature image, reducing the image size, and compressing the fused features obtained by feature extraction.

[0138] For example, downsampling processing can be implemented through maximum pooling or average pooling to reduce the number of features and parameters, thereby filtering out features that have little effect on true and false image detection and have redundant information, and retaining key information.

[0139] S203: Perform forged face detection on the target face based on multiple block feature maps to determine the detection result of the target face.

[0140] The detection result of the target face can be a fake face or a real face. In other words, after determining multiple block feature maps, the server can perform fake face detection on the local area of ​​the target face, thereby determining whether the target face in the original image is a fake face, that is, whether the original image is a deep fake image, thereby improving the accuracy of fake face detection.

[0141] In an optional implementation, when executing step S203, after the server extracts block features from the fused feature image and obtains multiple block feature maps, it can predict classification labels for the aforementioned multiple block feature maps based on multiple pre-trained block classifiers, thereby determining the detection result of the target face based on the obtained multiple classification labels and their corresponding classification weights. Each classification label can be used to indicate whether the image area associated with the corresponding block feature map is a forged area. The classification weight corresponding to each classification label is the classification proportion of each block classifier in determining whether the target face in the original image is a forged face. Optionally, the detection result of the aforementioned target face can be determined by the classification label value corresponding to the target face obtained in the following manner:

[0142]

[0143] Among them, Y represents the classification label value of the target face, K×K represents the number of block classifiers or the number of block feature maps, and α i represents the classification weight corresponding to the i-th block classifier (or block feature map), α 1 +α 2 +...+α K+K =1, the classification weight corresponding to each block classifier is determined by experiments, y i Indicates the classification label corresponding to the i-th block classifier (or block feature map). For example, if the block classifier determines that the image area associated with the corresponding block feature map is a real area, the output classification label is "1"; if the block classifier determines that the image area associated with the corresponding block feature map is a forged area, the output classification label is "0".

[0144] For example, see Figure 7As shown, assuming that the number of preset block classifiers is 3×3, that is, the global feature map containing the target face corresponds to 9 block feature maps, and the classification label output by each block classifier is "1" or "0", which is used to indicate whether the image area associated with the corresponding block feature map is a forged area. The server can use 9 block classifiers to determine the classification label value Y corresponding to the target face = 1×0.05+0×0.10+1×0.08+0×0.12+1×00.12+1×0.20+1×0.15+1×0.08+0×0.15+1×0.07=0.63. Assuming that the classification label threshold is 0.90, that is, when the classification label value Y is greater than or equal to the classification label threshold 0.90, it can be determined that the face detection result is a real face. Therefore, Figure 7 The detection result of the target face shown is known to be a forged face because the classification label value Y=0.63 is less than the classification label threshold 0.90.

[0145] Based on the above method, the K×K global feature map is passed through a block-level classifier to obtain K×K classification results (i.e., classification labels), and then the K×K classification results are weightedly summed to obtain the final classification result (i.e., the detection result of the target face); in this way, block-level supervision is more conducive to model learning and judging forgery features, thereby improving the accuracy of face forgery detection.

[0146] Based on the forged face detection method described in steps S201 to S203 above, refer to Figure 8 As shown in the figure, first, the original image is transformed in the frequency domain, filtered, and inversely transformed in the frequency domain to obtain a frequency domain information image containing multiple frequencies. The original image and the frequency domain information image are then fused and input into the convolutional neural network for feature extraction. Important information is adaptively mined through spatial attention and channel attention, thereby obtaining fine-grained facial features. Finally, the spatial-frequency domain fusion features of the block area are supervised and classified by a block-level classifier, which is more in line with the distribution of actual forged features on the image, and is more helpful for the model to learn and distinguish the differences between real and forged features, thereby improving the detection accuracy of forged faces.

[0147] In summary, in the forged face detection method provided in the embodiment of the present application, after obtaining the original image containing the target face, the original image can be transformed in the frequency domain and inversely transformed in the frequency domain to obtain a frequency domain information image; then, by performing image fusion on the original image and the frequency domain information image, a fused feature image that fuses the spatial domain features and the frequency domain features is obtained; finally, block feature extraction is performed on the fused feature image to obtain a plurality of block feature maps, and forged face detection is performed on the target face based on the plurality of block feature maps to determine whether the target face is a forged face or a real face. It can be seen that the forged detection result of the original image containing the target face is supervised and predicted on the block-level spatial-frequency domain fusion feature map, which improves the problem that the existing deep forgery detection method for faces does not fully consider the distribution of faces in the image, and / or that there is insufficient feature extraction of the image, thereby improving the detection accuracy of forged faces.

[0148] Further, based on the same technical concept, the embodiment of the present application provides a forged face detection device, which is used to implement the above method flow of the embodiment of the present application. Fig. 9 As shown, the forged face detection device 900 may include: an image acquisition module 901, a feature extraction module 902 and a face detection module 903, wherein:

[0149] The image acquisition module 901 is used to acquire an original image containing a target face, and perform frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image; wherein the original image and the frequency domain information image are both spatial domain images, and the frequency domain information image includes multiple frequency domain features of the target face;

[0150] The feature extraction module 902 is used to perform image fusion on the original image and the frequency domain information image to obtain a fused feature image, and perform block feature extraction on the fused feature image to obtain multiple block feature maps; wherein different block feature maps correspond to different image regions in the fused feature image;

[0151] The face detection module 903 is used to perform forged face detection on the target face based on multiple block feature maps to determine the detection result of the target face; the detection result is a forged face or a real face.

[0152] In an optional embodiment, when performing frequency domain transformation and frequency domain inverse transformation on the original image to obtain the frequency domain information image, the image acquisition module 901 is specifically used to:

[0153] Perform facial key point detection on the original image to obtain the position coordinates corresponding to multiple facial key points;

[0154] Based on multiple position coordinates and fixed coordinates respectively set for multiple facial key points, face alignment is performed on the target face to obtain a face alignment image;

[0155] Perform frequency domain transformation and inverse frequency domain transformation on the face alignment image to obtain a frequency domain information image.

[0156] In an optional embodiment, when performing facial key point detection on the original image to obtain position coordinates corresponding to a plurality of facial key points, the image acquisition module 901 is specifically used to:

[0157] Determine the location area of ​​the target face in the original image, and crop the original image based on the location area to obtain a face area image;

[0158] Perform facial key point detection on the face area image to obtain the position coordinates corresponding to multiple facial key points.

[0159] In an optional embodiment, when performing frequency domain transformation and frequency domain inverse transformation on the face alignment image to obtain the frequency domain information image, the image acquisition module 901 is specifically used to:

[0160] Performing frequency domain transformation on the face alignment image to obtain the face alignment image after frequency domain transformation;

[0161] Based on a plurality of preset filtering frequencies, filtering is performed on the face alignment images after the frequency domain transformation respectively to obtain a plurality of face alignment images after the filtering;

[0162] Performing frequency domain inverse transformation on the multiple filtered face alignment images to obtain multiple frequency domain inverse transformed face alignment images, and obtaining a frequency domain information image based on the multiple frequency domain inverse transformed face alignment images.

[0163] In an optional embodiment, when performing image fusion on the original image and the frequency domain information image to obtain a fused feature image, the feature extraction module 902 is specifically used to:

[0164] The original image and the frequency domain information image are fused in the frequency channel dimension to obtain a fused feature image; wherein each frequency channel corresponds to a filtering frequency.

[0165] In an optional embodiment, when performing block feature extraction on the fused feature image to obtain multiple block feature maps, the feature extraction module 902 is specifically used to:

[0166] Using spatial attention mechanism and / or channel attention mechanism, extract features from the fused feature image to obtain a global feature map of the target face;

[0167] Based on a preset number of block classifiers, the global feature map is divided into a plurality of block feature maps; wherein each block classifier is used to determine whether an image region associated with a corresponding block feature map is a forged region.

[0168] In an optional embodiment, when performing forged face detection on a target face based on a plurality of block feature maps to determine a detection result of the target face, the face detection module 903 is specifically used to:

[0169] Determine classification labels corresponding to the plurality of block feature maps respectively; wherein each classification label is used to indicate whether an image region associated with the corresponding block feature map is a forged region;

[0170] Based on multiple classification labels and their corresponding classification weights, the detection result of the target face is determined.

[0171] Based on the description of the above method embodiment and device embodiment, the exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform the method according to the embodiment of the present invention when executed by the at least one processor.

[0172] An embodiment of the present application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute a method according to an embodiment of the present application.

[0173] An embodiment of the present application also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute a method according to an embodiment of the present application.

[0174] See also Fig.10 As shown, the structured block diagram of the electronic device 1000 that can be used as the server or client of the present application will now be described, which is an example of the hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent the computer device of various forms of digital electronics, such as, laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only used as examples, and are not intended to limit the implementation of the present application described herein and / or required.

[0175] like Fig.10As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 to a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0176] A plurality of components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 may be any type of device capable of inputting information to the electronic device 1000, and the input unit 1006 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 1007 may be any type of device capable of presenting information, and may include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1008 may include but is not limited to a disk, an optical disk. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a worldwide interoperability for microwave access (WiMax) device, a cellular communication device, and / or the like.

[0177] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various AI computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above. For example, in some embodiments, the above-mentioned forged face detection method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. In some embodiments, the computing unit 1001 may be configured to perform the above-mentioned forged face detection method in any other appropriate manner (e.g., by means of firmware).

[0178] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0179] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0180] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tub (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0182] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0183] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

[0184] Furthermore, it should be understood that what is disclosed above is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope covered by the present application.

Claims

1. A method for detecting fake faces, characterized in that: include: Acquire an original image containing a target face, and perform frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image; wherein the original image and the frequency domain information image are both spatial domain images, and the frequency domain information image includes multiple frequency domain features of the target face; Performing image fusion on the original image and the frequency domain information image to obtain a fused feature image, and performing block feature extraction on the fused feature image to obtain a plurality of block feature maps; wherein different block feature maps correspond to different image regions in the fused feature image; Perform forged face detection on the target face based on the multiple block feature maps to determine a detection result of the target face; the detection result is a forged face or a real face.

2. The method according to claim 1, characterized in that The performing frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image includes: Performing facial key point detection on the original image to obtain position coordinates corresponding to a plurality of facial key points; Based on the multiple position coordinates and the fixed coordinates respectively set for the multiple facial key points, the target face is aligned to obtain a face alignment image; Perform frequency domain transformation and frequency domain inverse transformation on the face alignment image to obtain the frequency domain information image.

3. The method according to claim 2, characterized in that The detecting of facial key points on the original image to obtain position coordinates corresponding to a plurality of facial key points respectively includes: Determine a location area of ​​the target face in the original image, and crop the original image based on the location area to obtain a face area image; Perform facial key point detection on the face area image to obtain position coordinates corresponding to a plurality of facial key points.

4. The method according to claim 2, characterized in that The performing frequency domain transformation and frequency domain inverse transformation on the face alignment image to obtain the frequency domain information image includes: Performing frequency domain transformation on the face alignment image to obtain a face alignment image after frequency domain transformation; Based on a plurality of preset filtering frequencies, filtering is performed on the face alignment images after the frequency domain transformation respectively to obtain a plurality of face alignment images after filtering; Performing frequency domain inverse transformation on the multiple filtered face alignment images to obtain multiple frequency domain inverse transformed face alignment images, and obtaining the frequency domain information image based on the multiple frequency domain inverse transformed face alignment images.

5. The method according to claim 4, characterized in that The performing image fusion on the original image and the frequency domain information image to obtain a fused feature image includes: The original image and the frequency domain information image are subjected to image fusion processing in the frequency channel dimension to obtain the fused feature image; wherein each frequency channel corresponds to a filtering frequency.

6. The method according to any one of claims 1 to 5, characterized in that The step of extracting block features from the fused feature image to obtain a plurality of block feature maps comprises: Using a spatial attention mechanism and / or a channel attention mechanism to perform feature extraction on the fused feature image to obtain a global feature map of the target face; Based on a preset number of block classifiers, the global feature map is divided into the plurality of block feature maps; wherein each block classifier is used to determine whether an image region associated with a corresponding block feature map is a forged region.

7. The method according to any one of claims 1 to 5, characterized in that The performing forged face detection on the target face based on the multiple block feature maps to determine the detection result of the target face includes: Determining classification labels corresponding to the plurality of block feature maps respectively; wherein each classification label is used to indicate whether an image region associated with a corresponding block feature map is a forged region; Based on multiple classification labels and their corresponding classification weights, a detection result of the target face is determined.

8. A forged face detection device, characterized in that: include: An image acquisition module, used to acquire an original image containing a target face, and perform frequency domain transformation and frequency domain inverse transformation on the original image to obtain a frequency domain information image; wherein both the original image and the frequency domain information image are spatial domain images, and the frequency domain information image includes multiple frequency domain features of the target face; A feature extraction module, used to perform image fusion on the original image and the frequency domain information image to obtain a fused feature image, and perform block feature extraction on the fused feature image to obtain a plurality of block feature maps; wherein different block feature maps correspond to different image regions in the fused feature image; A face detection module is used to perform forged face detection on the target face based on the multiple block feature maps to determine the detection result of the target face; the detection result is a forged face or a real face.

9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.