Face forgery detection method and device, terminal and storage medium
By combining multi-layer convolutional neural networks and deep separable convolutional models with graph convolutional networks, and fusing facial action unit features and texture features, the problem of low accuracy in face forgery detection in existing technologies is solved, and higher detection accuracy and generalization ability are achieved.
Patent Information
- Application Number
- CN202210540707.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-05-17
AI Technical Summary
Existing face forgery detection techniques ignore the mutual exclusivity and co-occurrence of facial action units, resulting in low detection accuracy.
A multi-layer convolutional neural network model and a depthwise separable convolutional model are used to determine the global facial action unit features and global texture features of face images respectively, and these features are fused through a graph convolutional network model for detection.
The accuracy of face forgery detection is improved, and good generalization and detection results are provided.
Smart Images

Figure CN115050066B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of machine learning and computer vision technology, and more specifically, to a method, device, terminal, and storage medium for detecting face forgery. Background Art
[0002] Face forgery detection is to determine whether the face contained in a given image is forged.
[0003] Currently, there are two main approaches to face forgery detection. One is to use manually designed high-level semantic features for forgery detection, such as head posture consistency and abnormal blinking frequency. The other is to use data-driven facial defect features for forgery detection, such as inconsistent regional textures, abnormal generation artifacts, and abnormal spectral domain distribution.
[0004] However, the above methods ignore the mutual exclusivity and co-occurrence of facial action units, resulting in low accuracy in face forgery detection. Summary of the Invention
[0005] The main purpose of this application is to provide a face forgery detection method, device, terminal and storage medium to solve the problem of low accuracy of face forgery detection in related technologies.
[0006] To achieve the above objectives, in a first aspect, the present application provides a method for detecting forged faces, comprising:
[0007] receiving a face image;
[0008] Based on the face image and the multi-layer convolutional neural network model, the global facial action unit features corresponding to the face image are determined;
[0009] Based on the face image and the depth-wise separable convolutional model, the global texture features corresponding to the face image are determined;
[0010] The authenticity of face images is determined based on global facial action unit features and global texture features.
[0011] In one possible implementation, determining global facial action unit features corresponding to a facial image based on a facial image and a multi-layer convolutional neural network model includes:
[0012] Perform motion magnification on the face image to obtain a motion enhancement map corresponding to the face image;
[0013] The motion enhancement map is input into a multi-layer convolutional neural network model to obtain multiple feature maps;
[0014] Based on multiple feature maps, global facial action unit features corresponding to the face image are determined.
[0015] In one possible implementation, determining global facial action unit features corresponding to a face image based on multiple feature maps includes:
[0016] Determining, based on a plurality of facial key points and preset candidate frames set on the face image, a plurality of facial action unit regions corresponding to each feature map in a plurality of feature maps;
[0017] determining a facial action unit feature corresponding to each feature map based on the transformation coefficient and a plurality of facial action unit regions corresponding to each feature map;
[0018] Based on the facial action unit features corresponding to each feature map and the graph convolutional network model, the global facial action unit features are determined.
[0019] In one possible implementation, determining multiple facial action unit regions corresponding to each feature map in multiple feature maps based on multiple facial key points and preset candidate frames set on the face image includes:
[0020] Selecting, from the plurality of facial key points, a point with the smallest distance to each of the plurality of facial action units as the center of each facial action unit;
[0021] Match the preset candidate box to the center of each facial action unit to obtain the area of each facial action unit;
[0022] Each facial action unit region is aggregated to obtain multiple facial action unit regions corresponding to each feature map.
[0023] In one possible implementation, determining a facial action unit feature corresponding to each feature map based on the transformation coefficients and multiple facial action unit regions corresponding to each feature map includes:
[0024] determining transformation coefficients;
[0025] Using the transformation coefficient and the multiple facial action unit regions corresponding to each feature map, each facial action unit region in the multiple facial action unit regions is located, and features are extracted for each facial action unit region to obtain features corresponding to each facial action unit region;
[0026] Summarize the features corresponding to each facial action unit region to obtain features corresponding to multiple facial action unit regions;
[0027] The features corresponding to multiple facial action unit areas are sequentially convolved and pooled to obtain the facial action unit features corresponding to each feature map.
[0028] In one possible implementation, determining a global facial action unit feature based on the facial action unit feature corresponding to each feature graph and a graph convolutional network model includes:
[0029] Summarizing the facial action unit features corresponding to each feature map to obtain multiple facial action unit features corresponding to the multiple feature maps, wherein the multiple feature maps correspond to the multiple facial action unit features in a one-to-one manner;
[0030] Multiple facial action unit features corresponding to multiple feature maps are input into the graph convolutional network model and fused to obtain the global facial action unit feature.
[0031] In one possible implementation, determining global texture features corresponding to a facial image based on a facial image and a depthwise separable convolutional model includes:
[0032] Input the face image into the depthwise separable convolutional model to obtain the three-dimensional feature map corresponding to the face image;
[0033] A block-level loss function is used to supervise the authenticity of the sub-feature maps in the stereo feature map to obtain the shallow texture features corresponding to the stereo feature map;
[0034] Pooling is performed on shallow texture features to obtain global texture features.
[0035] In one possible implementation, determining the authenticity of a face image based on global facial action unit features and global texture features includes:
[0036] The global facial action unit features and global texture features are spliced together to obtain the target features;
[0037] Input the target feature into the classifier to obtain a first probability value and a second probability value corresponding to the target feature;
[0038] The first probability value and the second probability value are compared to obtain a comparison result, and the authenticity of the facial image is determined based on the comparison result.
[0039] In a second aspect, an embodiment of the present invention provides a face forgery detection device, comprising:
[0040] An image receiving module, configured to receive a face image;
[0041] An action unit feature determination module is used to determine the global facial action unit features corresponding to the face image based on the face image and the multi-layer convolutional neural network model;
[0042] A texture feature determination module is used to determine the global texture features corresponding to the face image based on the face image and the depth-separable convolution model;
[0043] The authenticity identification module is used to determine the authenticity of face images based on global facial action unit features and global texture features.
[0044] In a third aspect, an embodiment of the present invention provides a terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above methods for detecting forged faces are implemented.
[0045] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above face forgery detection methods.
[0046] Embodiments of the present invention provide a method, device, terminal, and storage medium for detecting facial forgery, comprising: receiving a facial image, determining global facial action unit features corresponding to the facial image based on the facial image and a multi-layer convolutional neural network model, then determining global texture features corresponding to the facial image based on the facial image and a depthwise separable convolutional model, and then determining the authenticity of the facial image based on the global facial action unit features and global texture features. The present invention uses a multi-layer convolutional neural network model to learn facial action unit features and model the co-occurrence dependencies of facial action units, so that facial motion features are further integrated with global dependencies to obtain global facial action unit features. This feature can help the model understand facial features more comprehensively for facial forgery detection. In addition, the global facial action unit features and global texture features are integrated to jointly perform facial forgery detection, providing good generalization for the facial forgery detection model and improving the accuracy of facial forgery detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings that constitute part of this application are used to provide a further understanding of this application and make other features, objects and advantages of this application more apparent. The illustrative embodiment drawings of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0048] Figure 1 This is a flowchart of a method for detecting forged faces provided by one embodiment of the present invention;
[0049] Figure 2 is a flowchart of a method for detecting forged faces provided by another embodiment of the present invention;
[0050] Figure 3 is a flowchart for implementing the determination of facial action unit features of each feature graph provided by an embodiment of the present invention;
[0051] Figure 4 is a schematic diagram of a multi-scale facial action unit dependency graph provided by an embodiment of the present invention;
[0052] Figure 51 is a schematic structural diagram of a face forgery detection device provided by an embodiment of the present invention;
[0053] Figure 6 is a schematic diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0055] The terms "first," "second," "third," "fourth," and so forth (if any) in the description and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present invention described herein can be practiced in sequences other than those illustrated or described herein.
[0056] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0057] It should be understood that in the present invention, "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0058] It should be understood that in the present invention, "multiple" refers to two or more. "And / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "Contains A, B and C", "Contains A, B, C" means that A, B, and C are all included, "Contains A, B or C" means that one of A, B, and C is included, and "Contains A, B and / or C" means that any one, any two, or any three of A, B, and C are included.
[0059] It should be understood that, in the present invention, "B corresponding to A," "B corresponding to A," "A corresponds to B," or "B corresponds to A" means that B is associated with A and B can be determined based on A. Determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information. A and B match when the similarity between A and B is greater than or equal to a preset threshold.
[0060] Depending on the context, "if" as used herein may be interpreted as "when" or "when" or "in response to determining" or "in response to detecting."
[0061] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0062] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below with reference to the accompanying drawings.
[0063] In one embodiment, Figure 1 As shown, a face forgery detection method is provided, comprising the following steps:
[0064] Step S101: receiving a face image;
[0065] Step S102: Based on the facial image and the multi-layer convolutional neural network model, determine the global facial action unit features corresponding to the facial image.
[0066] Among them, the global facial action unit feature refers to the features of all action units of the entire face.
[0067] After receiving a facial image, the present invention first performs motion amplification on the facial image to obtain a motion enhancement map corresponding to the facial image, then inputs the motion enhancement map into a multi-layer convolutional neural network model to obtain multiple feature maps, and then determines the global facial action unit features corresponding to the facial image based on the multiple feature maps.
[0068] Specific, combined Figure 2, when receiving a 299x299 RGB image of the face area (hereinafter referred to as the face image) and its corresponding facial key points, where the facial key points can be obtained in advance using corresponding tools. The face image is then sent to Magnet for motion amplification to enhance the movement performance of the facial muscles, and the motion enhancement map corresponding to the face image is output. The size of the motion enhancement map is also 299x299. Next, a multi-layer convolutional neural network model is used to extract features from the motion enhancement map, obtaining three feature maps, namely 76x76, 38x38, and 19x19 multi-layer feature maps. Finally, based on the three obtained feature maps, the global facial action unit features corresponding to the face image are determined.
[0069] Because the facial action units (FAUs) contained in shallow feature maps (76x76) have weak semantics but contain texture details, while the high-level feature maps (19x19) contain strong semantic features but lack texture details, adaptive FAU selection is performed on the feature maps at three different levels. The goal of adaptive FAU region selection is to automatically locate FAUs in the absence of annotation information to extract regional discriminative features. This module learns the locations of FAUs in a data-driven manner.
[0070] Therefore, determining the global facial action unit features corresponding to the face image based on multiple feature maps includes the following steps, specifically:
[0071] (1) Based on multiple facial key points and preset candidate frames set on the face image, multiple facial action unit regions corresponding to each feature map in the multiple feature maps are determined.
[0072] To determine the multiple facial action unit regions corresponding to each feature map in the multiple feature maps, it is necessary to first select the point with the smallest distance to each facial action unit in the multiple facial action units from the multiple facial key points as the center of each facial action unit, then match the preset candidate box for the center of each facial action unit to obtain each facial action unit region, and then summarize each facial action unit region to obtain the multiple facial action unit regions corresponding to each feature map.
[0073] Specifically, such as Figure 2The facial image shown is provided with multiple facial key points. Assume that the facial image corresponds to 17 facial action units, including AU1, AU2, AU4, AU5, AU6, AU7, AU9, AU12, AU14, AU15, AU16, AU17, AU20, AU23, AU25, AU26, and AU43. For each feature map, the point closest to each facial action unit among the multiple facial key points is first selected as the center (i.e., center coordinate) of the facial action unit. Then, a preset candidate box is matched to each facial action unit in each feature map. The preset candidate box is determined by matching the center of each facial action unit, and the preset candidate box size can be 9x9, 5x5, or 3x3. After matching the preset candidate box to each facial action unit in each feature map, each facial action unit area in each feature map can be obtained. Then, each facial action unit area is aggregated to obtain multiple facial action unit areas corresponding to each feature map.
[0074] (2) Based on the transformation coefficients and the multiple facial action unit regions corresponding to each feature map, the facial action unit features corresponding to each feature map are determined.
[0075] To determine the facial action unit features corresponding to each feature map, it is necessary to first determine the transformation coefficient, and then use the transformation coefficient and the multiple facial action unit areas corresponding to each feature map to locate each facial action unit area in the multiple facial action unit areas, and perform feature extraction on each facial action unit area to obtain the features corresponding to each facial action unit area, and then summarize the features corresponding to each facial action unit area to obtain the features corresponding to multiple facial action unit areas. Finally, the features corresponding to the multiple facial action unit areas are sequentially subjected to convolution calculation and pooling to obtain the facial action unit features corresponding to each feature map.
[0076] Specific, combined Figure 3 After obtaining multiple facial action unit areas corresponding to each feature map as mentioned above, a 1x1 convolution is used to compress each feature map channel to 1, and a feature vector with a length of 128 is obtained through global average pooling (GAP). Then, four transformation coefficients are predicted through a fully connected layer, namely, the length scaling coefficient, the width scaling coefficient, the up-down translation coefficient, and the left-right translation coefficient.
[0077] Subsequently, using the four resulting transformation coefficients and the multiple facial action unit regions corresponding to each feature map, each facial action unit region is adaptively located to extract features within the region and determine the features corresponding to each facial action unit region. The features corresponding to each facial action unit region are then extracted and subjected to three layers of convolution and pooling to obtain the motion features of each facial action unit (i.e., facial action unit features). For each feature map, 3x17 facial action unit motion features are extracted, each corresponding to a specific facial action unit, meaning that each feature map corresponds to 3x17 facial action unit features.
[0078] (3) Based on the facial action unit features corresponding to each feature map and the graph convolutional network model, the global facial action unit features are determined.
[0079] To determine the global facial action unit feature, the facial action unit features corresponding to each feature map must first be aggregated to obtain multiple facial action unit features corresponding to multiple feature maps, where the multiple feature maps correspond one-to-one to the multiple facial action unit features. These multiple facial action unit features corresponding to the multiple feature maps are then fed into a graph convolutional network model and fused to obtain the global facial action unit feature. The graph in a graph convolutional network (GCN) is a non-Euclidean data format that can be used to represent social networks, communication networks, protein molecular networks, and more. Graph convolutional networks model the node and structural features of a graph through an information propagation mechanism and are often used to mine co-occurrence relationships between nodes. This technique is used here to model the co-occurrence relationships of facial action units.
[0080] Specific, combined Figure 2 and Figure 4 , Multi-scale facial action unit dependency modeling requires learning the facial action unit dependency, including intra-layer unit modeling and inter-layer unit modeling. Among them, inter-layer unit modeling: for the same motion unit, the feature maps of different layers correspond to 3 nodes. These 3 nodes are connected in pairs to form the inter-layer unit model of this scheme. Intra-layer unit modeling: Each feature map contains 17 nodes, and the co-occurrence frequency between the 17 nodes is determined based on the co-occurrence frequency. If there is an edge between the two nodes, that is, if the co-occurrence frequency is greater than a certain threshold, there is a connecting edge. Otherwise, there is no connecting edge, which constitutes an intra-layer unit model. Then, the inter-layer unit model and the intra-layer unit model are combined to obtain 3x17 nodes and corresponding edges. Among them, the nodes are facial action unit features.
[0081] Next, a GCN graph convolutional network model is used to perform graph learning on the network graph composed of the obtained 3x17 facial action unit features (i.e., the multi-scale facial action unit dependency graph). In other words, the co-occurrence dependencies between facial action units are modeled to obtain a corresponding number of new features. These new features not only contain action unit motion characteristics but also topological dependency characteristics. Finally, these new features are fused to obtain the global facial action unit feature.
[0082] It should be noted that when learning the above graph, the activation state of the facial action unit is supervised, that is, whether the facial action unit is activated (real movement) is supervised to help the model better locate the facial action unit.
[0083] Step S103: Determine the global texture features corresponding to the face image based on the face image and the depthwise separable convolution model.
[0084] To determine the global texture features corresponding to the face image, the face image must first be input into the depthwise separable convolutional model to obtain the three-dimensional feature map corresponding to the face image. Then, a block-level loss function is used to supervise the authenticity of the sub-feature maps in the three-dimensional feature map to obtain the shallow texture features corresponding to the three-dimensional feature map. The shallow texture features are then pooled to obtain the global texture features.
[0085] Specifically, this embodiment is designed from the two perspectives of network design and supervision loss to drive the model to find shallow texture features, so as to improve the generalization ability of features in the face of unknown generation technology or unknown defects. Therefore, this solution inputs the face image into the depth separable convolution model, and finally obtains a three-dimensional feature map with a size of 38x38x256, wherein the depth separable convolution model is a 3-layer texture feature extraction model completely based on depth separable convolution. Then, a block-level loss function is used to supervise the authenticity of the 38x38 sub-feature map. Among them, the block-level supervision annotation is directly mapped from the global annotation. After the above-mentioned supervised learning, the shallow texture features are obtained, and then the shallow texture features are pooled to finally obtain the global texture features.
[0086] Furthermore, block-level supervised annotation primarily covers true and false, while false images can include: 1. Images with flaws and artifacts left over from the generation process. These flaws may come from traces of face-swapping, motion blur around facial features, incomplete tooth modeling, and more, and these flaws are more likely to occur in high-frequency space. 2. Images where the texture of the generated area is inconsistent with that of the surrounding area. Every person's facial texture is unique, and pasting the generated face back onto the target face will inevitably cause texture conflicts between the generated and original areas, which can serve as a basis for counterfeit detection. 3. Images containing inherent noise "fingerprints" from the GAN (Generation Network) generation tool and the camera's light sensor. The "fingerprint" of the GAN generation tool comes from a fixed convolution kernel, upsampling method, and other factors; the camera's light sensor has unique noise from the factory, and this feature is present in all generated and forged images.
[0087] Step S104: Determine the authenticity of the face image based on the global facial action unit features and the global texture features.
[0088] To determine the authenticity of a facial image, it is necessary to first concatenate the global facial action unit features and the global texture features to obtain the target features, then input the target features into a classifier to obtain the first probability value and the second probability value corresponding to the target features, and then compare the first probability value and the second probability value to obtain a comparison result, and determine the authenticity of the facial image based on the comparison result.
[0089] Specifically, the target feature is input into the classifier to obtain the first probability value and the second probability value corresponding to the target feature. The first probability value is assumed to represent true and the second probability value is assumed to represent false. When the first probability value is greater than the second probability value, the face image can be determined to be true; when the second probability value is greater than the first probability value, the face image is determined to be false.
[0090] An embodiment of the present invention provides a method for detecting face forgery, comprising: receiving a face image, determining global facial action unit features corresponding to the face image based on the face image and a multi-layer convolutional neural network model, then determining global texture features corresponding to the face image based on the face image and a depthwise separable convolutional model, and then determining the authenticity of the face image based on the global facial action unit features and global texture features. The present invention uses a multi-layer convolutional neural network model to learn facial action unit features and models the co-occurrence dependencies of facial action units, so that facial motion features are further integrated with global dependencies to obtain global facial action unit features. This feature can help the model understand facial features more comprehensively for face forgery detection. In addition, the global facial action unit features and global texture features are integrated to jointly perform face forgery detection, providing good generalization for the face forgery detection model and improving the accuracy of face forgery detection.
[0091] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0092] The following are device embodiments of the present invention. For details not fully described therein, reference may be made to the corresponding method embodiments described above.
[0093] Figure 5 A schematic diagram of the structure of a face forgery detection device provided by an embodiment of the present invention is shown. For ease of explanation, only the parts relevant to the embodiment of the present invention are shown. The face forgery detection device includes an image receiving module 51, an action unit feature determination module 52, a texture feature determination module 53, and an authenticity identification module 54. The details are as follows:
[0094] Image receiving module 51, used for receiving facial images;
[0095] an action unit feature determination module 52 for determining global facial action unit features corresponding to the face image based on the face image and the multi-layer convolutional neural network model;
[0096] A texture feature determination module 53 is used to determine the global texture features corresponding to the face image based on the face image and the depth-separable convolution model;
[0097] The authenticity identification module 54 is used to determine the authenticity of the face image based on the global facial action unit features and the global texture features.
[0098] In one possible implementation, the action unit feature determination module 52 includes:
[0099] The image magnification submodule is used to perform motion magnification on the face image to obtain a motion enhancement map corresponding to the face image;
[0100] The first model processing submodule is used to input the motion enhancement map into the multi-layer convolutional neural network model to obtain multiple feature maps;
[0101] The global feature determination submodule is used to determine the global facial action unit features corresponding to the face image based on multiple feature maps.
[0102] In one possible implementation, the global feature determination submodule includes:
[0103] A region determination unit, configured to determine, based on a plurality of facial key points and preset candidate frames set on the face image, a plurality of facial action unit regions corresponding to each of the plurality of feature maps;
[0104] a feature determination unit for determining a facial action unit feature corresponding to each feature map based on the transformation coefficient and a plurality of facial action unit regions corresponding to each feature map;
[0105] The global feature determination unit is used to determine the global facial action unit features based on the facial action unit features corresponding to each feature map and the graph convolutional network model.
[0106] In a possible implementation, the region determining unit includes:
[0107] a center selection subunit, configured to select, from a plurality of facial key points, a point with a minimum distance from each of a plurality of facial action units as the center of each facial action unit;
[0108] The region matching subunit is used to match the preset candidate box with the center of each facial action unit to obtain the region of each facial action unit;
[0109] The region determination subunit is used to aggregate each facial action unit region to obtain multiple facial action unit regions corresponding to each feature map.
[0110] In one possible implementation, the feature determination unit includes:
[0111] A coefficient determination subunit, used for determining a transformation coefficient;
[0112] a feature extraction subunit, configured to locate each facial action unit region among the plurality of facial action unit regions by using the transformation coefficients and the plurality of facial action unit regions corresponding to each feature map, and to perform feature extraction on each facial action unit region to obtain features corresponding to each facial action unit region;
[0113] A first feature aggregation subunit is used to aggregate the features corresponding to each facial action unit region to obtain features corresponding to multiple facial action unit regions;
[0114] The feature determination subunit is used to perform convolution calculation and pooling on the features corresponding to multiple facial action unit areas in sequence to obtain the facial action unit features corresponding to each feature map.
[0115] In a possible implementation, the global feature determination unit includes:
[0116] a second feature aggregation subunit, configured to aggregate the facial action unit features corresponding to each feature map to obtain a plurality of facial action unit features corresponding to the plurality of feature maps, wherein the plurality of feature maps correspond one-to-one to the plurality of facial action unit features;
[0117] The global feature determination subunit is used to input multiple facial action unit features corresponding to multiple feature maps into the graph convolutional network model and then fuse them to obtain the global facial action unit feature.
[0118] In a possible implementation, the texture feature determination module 53 includes:
[0119] The second model processing submodule is used to input the face image into the depthwise separable convolution model to obtain a three-dimensional feature map corresponding to the face image;
[0120] The supervised learning submodule is used to supervise the authenticity of the sub-feature maps in the stereo feature map using a block-level loss function to obtain the shallow texture features corresponding to the stereo feature map;
[0121] The feature pooling submodule is used to pool shallow texture features to obtain global texture features.
[0122] In one possible implementation, the authenticity identification module 54 includes:
[0123] The feature splicing submodule is used to splice the global facial action unit features and the global texture features to obtain the target features;
[0124] The classification calculation submodule is used to input the target feature into the classifier to obtain a first probability value and a second probability value corresponding to the target feature;
[0125] The authenticity identification submodule is used to compare the first probability value and the second probability value to obtain a comparison result, and determine the authenticity of the face image based on the comparison result.
[0126] Figure 6 Schematic diagram of a terminal provided by an embodiment of the present invention. Figure 6 As shown, the terminal 6 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the processor 61 executes the computer program 63, the steps in each of the above-mentioned face forgery detection method embodiments are implemented, such as Figure 1 Alternatively, when the processor 61 executes the computer program 63, the functions of the modules / units in the above-mentioned embodiments of the face forgery detection device are realized, for example Figure 5 Functions of the modules / units 51 to 54 are shown.
[0127] The present invention also provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the face forgery detection method provided by the various embodiments described above.
[0128] Among them, the readable storage medium can be a computer storage medium or a communication medium. Communication media include any medium that facilitates the transmission of computer programs from one place to another. Computer storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist in a communication device as discrete components. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0129] The present invention also provides a program product, comprising execution instructions stored in a readable storage medium. At least one processor of a device can read the execution instructions from the readable storage medium, and the at least one processor executes the execution instructions to cause the device to implement the face forgery detection methods provided in the various embodiments described above.
[0130] In the embodiments of the above-mentioned devices, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented by a hardware processor or implemented by a combination of hardware and software modules in the processor.
[0131] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for detecting forged faces, characterized in that: include: receiving a face image; Determining global facial action unit features corresponding to the facial image based on the facial image and a multi-layer convolutional neural network model; Determining global texture features corresponding to the facial image based on the facial image and a depthwise separable convolutional model; determining the authenticity of the face image based on the global facial action unit features and the global texture features; The determining, based on the facial image and the multi-layer convolutional neural network model, a global facial action unit feature corresponding to the facial image includes: Performing motion magnification on the facial image to obtain a motion enhancement map corresponding to the facial image; Inputting the motion enhancement map into the multi-layer convolutional neural network model to obtain a plurality of feature maps; determining, based on the plurality of feature maps, global facial action unit features corresponding to the face image; The determining, based on the multiple feature maps, global facial action unit features corresponding to the face image includes: Determining, based on a plurality of facial key points and preset candidate frames set on the face image, a plurality of facial action unit regions corresponding to each feature map in the plurality of feature maps; determining a facial action unit feature corresponding to each feature map based on the transformation coefficient and a plurality of facial action unit regions corresponding to each feature map; Determining the global facial action unit feature based on the facial action unit feature corresponding to each feature graph and the graph convolutional network model; The step of determining the facial action unit feature corresponding to each feature map based on the transformation coefficient and the plurality of facial action unit regions corresponding to each feature map comprises: determining the transform coefficients; Using the transformation coefficient and a plurality of facial action unit regions corresponding to each feature map, locating each facial action unit region among the plurality of facial action unit regions, and performing feature extraction on each facial action unit region to obtain features corresponding to each facial action unit region; Summarizing the features corresponding to each facial action unit region to obtain features corresponding to the multiple facial action unit regions; The features corresponding to the multiple facial action unit areas are sequentially subjected to convolution calculation and pooling to obtain facial action unit features corresponding to each feature map; The global facial action unit feature refers to the features of all action units in the entire face, where each facial action unit is obtained by matching the center of the facial action unit determined by facial key points with a preset candidate box. The facial action unit is used to characterize the muscle area in the face.
2. The face forgery detection method according to claim 1, wherein: The step of determining, based on a plurality of facial key points and preset candidate frames set on the face image, a plurality of facial action unit regions corresponding to each feature map in the plurality of feature maps comprises: Selecting, from the plurality of facial key points, a point with the shortest distance from each of the plurality of facial action units as the center of each facial action unit; Matching the center of each facial action unit with the preset candidate frame to obtain each facial action unit area; Each facial action unit region is aggregated to obtain multiple facial action unit regions corresponding to each feature map.
3. The face forgery detection method according to claim 1, wherein: The determining the global facial action unit feature based on the facial action unit feature corresponding to each feature graph and the graph convolutional network model includes: Summarizing the facial action unit features corresponding to each feature graph to obtain a plurality of facial action unit features corresponding to the plurality of feature graphs, wherein the plurality of feature graphs correspond to the plurality of facial action unit features in a one-to-one manner; Multiple facial action unit features corresponding to the multiple feature maps are input into the graph convolutional network model and then fused to obtain the global facial action unit feature.
4. The face forgery detection method according to any one of claims 1 to 3, wherein: The determining, based on the facial image and the depthwise separable convolution model, a global texture feature corresponding to the facial image includes: Inputting the face image into a depthwise separable convolutional model to obtain a three-dimensional feature map corresponding to the face image; Using a block-level loss function to perform supervised learning on the authenticity of the sub-feature map in the stereo feature map, to obtain shallow texture features corresponding to the stereo feature map; Pooling is performed on the shallow texture features to obtain the global texture features.
5. The face forgery detection method according to any one of claims 1 to 3, wherein: The determining the authenticity of the face image based on the global facial action unit feature and the global texture feature includes: Concatenating the global facial action unit feature and the global texture feature to obtain a target feature; Inputting the target feature into a classifier to obtain a first probability value and a second probability value corresponding to the target feature; The first probability value and the second probability value are compared to obtain a comparison result, and the authenticity of the facial image is determined based on the comparison result.
6. A face forgery detection device, characterized in that: include: An image receiving module, configured to receive a face image; an action unit feature determination module, configured to determine global facial action unit features corresponding to the facial image based on the facial image and a multi-layer convolutional neural network model; A texture feature determination module, configured to determine global texture features corresponding to the facial image based on the facial image and a depthwise separable convolution model; an authenticity identification module, configured to determine the authenticity of the face image based on the global facial action unit features and the global texture features; The determining, based on the facial image and the multi-layer convolutional neural network model, a global facial action unit feature corresponding to the facial image includes: Performing motion magnification on the facial image to obtain a motion enhancement map corresponding to the facial image; Inputting the motion enhancement map into the multi-layer convolutional neural network model to obtain a plurality of feature maps; determining, based on the plurality of feature maps, global facial action unit features corresponding to the face image; The determining, based on the facial image and the multi-layer convolutional neural network model, a global facial action unit feature corresponding to the facial image includes: Performing motion magnification on the facial image to obtain a motion enhancement map corresponding to the facial image; Inputting the motion enhancement map into the multi-layer convolutional neural network model to obtain a plurality of feature maps; determining, based on the plurality of feature maps, global facial action unit features corresponding to the face image; The determining, based on the multiple feature maps, global facial action unit features corresponding to the face image includes: Determining, based on a plurality of facial key points and preset candidate frames set on the face image, a plurality of facial action unit regions corresponding to each feature map in the plurality of feature maps; determining a facial action unit feature corresponding to each feature map based on the transformation coefficient and a plurality of facial action unit regions corresponding to each feature map; Determining the global facial action unit feature based on the facial action unit feature corresponding to each feature graph and the graph convolutional network model; The step of determining the facial action unit feature corresponding to each feature map based on the transformation coefficient and the plurality of facial action unit regions corresponding to each feature map comprises: determining the transform coefficients; Using the transformation coefficient and a plurality of facial action unit regions corresponding to each feature map, locating each facial action unit region among the plurality of facial action unit regions, and performing feature extraction on each facial action unit region to obtain features corresponding to each facial action unit region; Summarizing the features corresponding to each facial action unit region to obtain features corresponding to the multiple facial action unit regions; The features corresponding to the multiple facial action unit areas are sequentially subjected to convolution calculation and pooling to obtain facial action unit features corresponding to each feature map; The global facial action unit feature refers to the features of all action units in the entire face, where each facial action unit is obtained by matching the center of the facial action unit determined by facial key points with a preset candidate box. The facial action unit is used to characterize the muscle area in the face.
7. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the face forgery detection method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the face forgery detection method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Human face living body detection system applying double-branch three-dimensional convolution model, terminal and storage medium
CN111814574A