Facial key point-based counterfeit speaking face detection method and system
Through the detection method based on facial key points, combined with the graph attention mechanism and a bidirectional GRU encoder, a forged talking face detection network is constructed, which solves the problem of poor robustness in the video heavy compression and social communication scenarios in the prior art, and realizes high-precision forged video detection.
Patent Information
- Application Number
- CN202510022868.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
AI Technical Summary
The existing forged speaking face detection technology is poorly robust in video heavy compression and social communication scenarios, with low accuracy, making it difficult to effectively identify forged features in compressed forged videos.
Using a facial key point detection method, a facial key point adjacency network is constructed by obtaining face image sample video frames, and a facial key point adjacency network is constructed by combining the short-term forged trace encoder of the graph attention mechanism and a long-term forged trace encoder of the bidirectional GRU to construct a forged speaking face detection network, predict and network training are carried out, and the detection model is optimized.
It improves the robustness and accuracy of face detection forged speaking, and can accurately identify forged videos in complex compressed scenarios, significantly improving the practicality of detection.
Smart Images

Figure CN119964215A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of forged face detection, and in particular to a forged speaking face detection method and system based on facial key points. Background Art
[0002] Deepfake technology has brought many challenges and risks to all walks of life. Talking face fake technology not only poses a serious threat to personal privacy, but also brings new security challenges to commercial and government institutions. Therefore, it is urgent to take measures to prevent and respond to the risks that may be brought by talking face fake technology.
[0003] In order to resist the potential harm caused by deep fake technology, the detection technology for fake speaking faces has received widespread attention. However, the existing fake speaking face detection technology still has poor robustness, which restricts the practical application of fake detection technology. In online interactive scenarios, in order to achieve fast network transmission to ensure good real-time audio-visual interaction experience between the two communicating parties, video data usually needs to be compressed and encoded. The receiving party's video will have a certain degree of degradation. The compressed deep fake video has compression artifacts similar to tampering artifacts. This makes it more challenging to extract discriminative features from compressed fake videos. In terms of video feature extraction, most of the current deep fake detection schemes rely on pixel-level features. These schemes are very vulnerable to common video processing such as image compression, cropping, noise and other attacks. The detection scheme based on pixel-level features has poor robustness in video heavy compression scenarios, and the accuracy of fake video detection is greatly reduced by 10%.
[0004] Therefore, it is particularly important to propose a robust detection algorithm with stable and high accuracy. Summary of the invention
[0005] The present invention provides a method and system for detecting fake speaking faces based on facial key points, so as to solve the defects of poor robustness and low precision of the fake face technology in the prior art in video recompression and social communication scenarios.
[0006] In a first aspect, the present invention provides a method for detecting a fake speaking face based on facial key points, comprising: Obtain a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output facial key points of the face; A short-term forgery trace encoder based on the graph attention mechanism and a long-term forgery trace encoder based on the bidirectional GRU are cascaded to construct a forged speaking face detection network, the facial key points of the face are input into the forged speaking face detection network for prediction, and the network is trained in combination with the authenticity labels of the video frames to obtain an optimized forged speaking face detection model; The video to be detected is input into the optimized forged speaking face detection model to obtain a forged speaking face detection result.
[0007] According to a method for detecting fake speaking faces based on facial key points provided by the present invention, a sample video frame of a facial image is obtained, the sample video frame of the facial image is extracted to construct a facial key point adjacency network, and the facial key points of the face are output, including: The OpenFace open source tool is used to extract the location information of multiple facial key points of the face image sample video frame; Extracting forehead key points, eye key points, nose area and mouth area landmarks from the plurality of facial key point position information based on facial muscle movement of the human face to form a facial muscle movement network; Based on the forgery clues, the lip-jaw distance, lip opening and closing, nose-lip distance and facial contour are added to the facial muscle movement network of the human face to form the facial key point adjacency network.
[0008] According to a method for detecting forged speaking faces based on facial key points provided by the present invention, the short-term forged trace encoder of the graph attention mechanism comprises: Obtain the coordinate pairs corresponding to the key points of each face at any time, and integrate all the coordinate pairs to form a horizontal coordinate sequence and a vertical coordinate sequence; Constructing a graph structure from the abscissa sequence and the ordinate sequence, wherein the graph structure includes a node set and an edge set; Determine an adjacency matrix of the graph structure, wherein a non-zero element in the adjacency matrix is a connection strength between any two nodes, and a zero element is no connection between any two nodes; Representing the adjacency matrix based on the local sequence time length to obtain a spatiotemporal adjacency matrix, wherein the spatiotemporal adjacency matrix includes a connection relationship between a current timestamp node and neighboring nodes; Taking the feature set of each facial key point in the spatiotemporal adjacency matrix as input data to form a feature tensor, the feature tensor includes the feature dimension of each node, the number of frames and the number of nodes, and taking the time length of the local sequence as a sliding window to generate a local facial key point sequence; Introducing an attention mechanism into the local facial key point sequence to form a graph attention network, obtaining output features of any layer in the graph attention network, converting the output features of any layer into high-dimensional features, calculating the high-dimensional features using the attention mechanism, and obtaining attention scores of adjacent key points; The attention score is calculated by LeakyReLu activation and normalized exponential function softmax in turn to obtain the final normalized attention score; The final normalized attention score, any layer node feature conversion weight matrix and any layer output feature in the graph attention network are summed within the range of adjacent nodes of any node in the spatiotemporal adjacency matrix to obtain the next layer output feature of any layer in the graph attention network; Calculate the average of the next layer output features of any layer in all graph attention networks in K attention heads to obtain the output value under the multi-head attention mechanism.
[0009] According to a method for detecting forged speaking faces based on facial key points provided by the present invention, the bidirectional GRU long-term forgery trace encoder comprises: The output values under the multi-head attention mechanism are respectively input into two layers of bidirectional GRU to obtain two layers of bidirectional output features, where each layer of GRU contains the number of hidden units.
[0010] According to a method for detecting a fake speaking face based on facial key points provided by the present invention, the facial key points of the face are input into the fake speaking face detection network for prediction, and the network is trained in combination with the authenticity labels of the video frames to obtain an optimized fake speaking face detection model, including: The softmax layer is used to map the feature representation of the facial key points in the video to labels for classification, and the probability of authenticity or forgery is obtained; The optimized forged speaking face detection model is obtained by training based on minimizing the cross entropy loss between the real or forged probability and the video label.
[0011] In a second aspect, the present invention further provides a forged speaking face detection system based on facial key points, comprising: An extraction module is used to obtain a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output facial key points of the face; A training module, for cascading a short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder of a bidirectional GRU, constructing a forged speaking face detection network, inputting the facial key points of the face into the forged speaking face detection network for prediction, and performing network training in combination with the authenticity labels of the video frames to obtain an optimized forged speaking face detection model; The detection module is used to input the video to be detected into the optimized forged speaking face detection model to obtain the forged speaking face detection result.
[0012] In a third aspect, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for detecting a fake speaking face based on facial key points as described above is implemented.
[0013] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for detecting forged speaking faces based on facial key points.
[0014] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-mentioned methods for detecting fake speaking faces based on facial key points.
[0015] The method and system for detecting fake speaking faces based on facial key points provided by the present invention designs a facial key point connection network through in-depth analysis of muscle movements caused by speaking behavior and fake clues caused by the deep fake speaking face video generation process, and uses a graph attention network as the backbone network to achieve the extraction of facial authenticity features while retaining the facial topological structure; at the same time, considering the importance of long-term and short-term features in video forgery detection, short-term feature modeling is achieved by establishing temporal connections in the graph network, and long-term feature modeling is achieved by using a recurrent neural network, which can provide accurate and robust fake speaking face detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a flow chart of a method for detecting a fake speaking face based on facial key points provided by the present invention; Figure 2 It is the overall model structure diagram provided by the present invention; Figure 3 It is a design diagram of the network structure of facial key points based on muscle movement provided by the present invention; Figure 4 The present invention provides a facial key point network structure design based on forgery clues - lip-jaw distance map; Figure 5 The present invention provides a facial key point network structure design based on forgery clues - lip-nose distance map; Figure 6 It is the facial key point network structure design based on forgery clues provided by the present invention - lip closure map; Figure 7 It is a design diagram of a facial key point network structure based on muscle movement and forgery clues provided by the present invention; Figure 8It is a block diagram of the spatiotemporal graph attention block provided by the present invention; Fig. 9 It is a structural schematic diagram of a forged speaking face detection system based on facial key points provided by the present invention; Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] To address the problem that the existing technology of face deep fake detection is not robust in the compressed scenario of social communication platforms, a face deep fake detection method based on facial key points is proposed.
[0020] Figure 1 FIG. 1 is a flow chart of a method for detecting a fake speaking face based on facial key points provided by an embodiment of the present invention. Figure 1 As shown, including: Step 100: Obtain a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output facial key points of the face; Step 200: A short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder based on a bidirectional GRU are cascaded to construct a forged speaking face detection network, the facial key points of the face are input into the forged speaking face detection network for prediction, and the network is trained in combination with the authenticity labels of the video frames to obtain an optimized forged speaking face detection model; Step 300: input the video to be detected into the optimized forged speaking face detection model to obtain the forged speaking face detection result.
[0021] The model proposed in the embodiment of the present invention includes two key ideas: a detailed analysis of facial muscle movements and deep fake traces is carried out to build a facial key point adjacency network; for the short-term and long-term features in the video, a spatiotemporal graph neural network combined with a recurrent neural network is used to accurately extract the key point features of the video face. This method of comprehensive use of spatiotemporal information and neural network technology makes the facial deep fake detection method more robust in complex compression scenarios.
[0022] Specifically, the key steps include: The overall model structure is as follows Figure 2As shown in the figure, it is mainly divided into three parts: facial key point adjacency network design, short-term forgery trace encoder based on graph attention mechanism, and long-term forgery trace encoder based on bidirectional GRU.
[0023] The first part is the extraction of facial key points from video frames.
[0024] First, the location information of 68 facial key points extracted using the OpenFace open source tool is compiled as , i=1-68, OpenFace uses deep learning and computer vision technology to capture facial images and process them to locate facial areas. Then, it uses the trained model to identify and mark the locations of 68 key points. The location information of these key points is usually expressed in the form of two-dimensional coordinates (x, y). These 68 key points include eyes, eyebrows, nose, mouth, and specific locations of facial contours.
[0025] Then, by analyzing facial muscle movements and deep fake traces, an adjacency network based on speaking facial key points is constructed.
[0026] The facial key point network is constructed based on the muscle movements of the speaking face and the forgery traces brought by the deep fake video generation process. Compared with the fully connected structure, the construction of the facial key point network reduces the amount of calculation while retaining the topological structure of the face, achieving better transmission and aggregation of information between key points.
[0027] (1) Facial key point network construction scheme based on muscle movement Different emotions and sounds result in different muscle movements. When expressing anger, eyebrows often come together and stare. Arched eyebrows may indicate surprise or displeasure, while blinking eyes may indicate excitement or anxiety. When speaking, the shape of the mouth changes with different phonemes. For example, when making closed sounds such as "p" and "b", the lips are closed, while when making fricative sounds such as "f" and "v", the lips may be slightly open.
[0028] Facial muscle movements are considered to consist of a set of discrete components, which are called action units (AUs). The present invention adds key points and connections of the forehead (AU4, AU9), eyes (AU1, AU2, AU5, AU6), nose (AU9) and mouth area (AU12-AU28) to the facial key point network. The connections between key points are defined according to their semantic relationships. These key areas represent key muscles related to facial pronunciation and emotions, and can capture unnatural facial muscle activities, thereby providing clues for face forgery detection.
[0029] Specifically, for the 68 facial key points extracted using OpenFace, the key points used include the forehead key points ( , =17-21, 22-26), eye key points ( , =36-41, 42-47), nose area ( , =27-35) and mouth area ( , =48-59, 60-67). Connect these areas to represent the facial features and the movement of the facial contour, such as Figure 3 shown.
[0030] (2) Facial key point network construction scheme based on forgery clues From the perspective of forgery traces, different face forgery schemes have different vulnerabilities. In the videos generated by face replacement, there is inconsistency in the pitch angle between the inner face and the outer face. The inner face refers to the facial features of the central facial features of the character in the video, while the outer face refers to the outer contour of the character. For example, when the character lowers his head, the pitch angles of the inner face and the outer face may be inconsistent, resulting in the exposure of forgery traces. As for deep forgery technology such as lip synchronization, since human lip movement is highly correlated with speech, it is often difficult to perfectly simulate the real lip movement when performing lip synchronization forgery, which also provides certain clues for detection. It can be identified by detecting indicators such as the degree of lip opening and the lip-jaw distance. In forgeries such as face manipulation, it is easy to have inconsistencies in head movement and facial expressions, resulting in forgery traces. There are also cases where the 3D model cannot perfectly capture the facial features of the character, resulting in forgery traces.
[0031] In order to more intuitively observe the differences between the real and fake videos at each facial connection, Figure 4 , Figure 5 and Figure 6 The lip-jaw distance ( 8- 57 ), lip-nose distance ( 33 - 51 ) and lip opening and closing distance ( 62 - 66 ). The first line of each set of curves is a real video, the horizontal axis represents each video frame of the entire video, and the vertical axis represents the distance between key points.
[0032] exist Figure 4As shown in the figure, it is a curve graph of lip-jaw distance. The lip-jaw distance curve of the real video (represented by the black line) shows obvious peak-to-valley changes, which reflects the natural movement of the lips and jaw when people speak or make facial expressions in real situations. However, in the fake video (represented by the colored lines), this peak-to-valley change becomes less obvious or even disappears completely. This difference suggests that the fake video may have technical deficiencies in simulating the movement of the lips and jaw, or that the forger deliberately smoothed these changes in order to cover up the traces of forgery. A similar phenomenon also appears in the curve of nose-lip distance, such as Figure 5 As shown in Figure 2. In real videos, the change in nose-lip distance can reflect the relative movement between the nose and lips when a person makes different expressions and speaks. In fake videos, this change is often weakened or eliminated, making the facial movements in fake videos appear more rigid and unnatural. Furthermore, observe the curve of the lips opening and closing. Figure 6 In the figure, the lip opening and closing curves of the real video show obvious opening and closing movements, which is consistent with the natural movement of the lips when people speak. However, in the fake video, the opening and closing of the lips becomes less obvious, lacking the obvious opening and closing movements in the real video. This difference may be due to the technical limitations of the fake video in simulating lip movements, or the faker deliberately reduced the amplitude of lip movements to avoid detection.
[0033] In view of the above analysis and observation, this section adds key points and their connections that are highly correlated with forgery traces based on the facial muscle movement of the human face, including the lip-jaw distance ( , = 8, 57), lips opening and closing ( , = 62, 66), nose-lip distance ( , = 33, 51), facial contour ( , = 0 16). This facial network construction scheme fully considers the inherent forgery traces caused by common forgery methods such as face replacement, lip synchronization and face manipulation, helping the network to better capture forgery clues and achieve forged face recognition. The final facial key point network structure design based on muscle movement and forgery clues is as follows Figure 7 As shown, Figure 7 The middle blue connections are the muscle movement-based facial network, and the orange connections are the falsification cue-based facial network.
[0034] The second part is to build and train a facial key point feature representation network based on the spatiotemporal graph attention mechanism.
[0035] The embodiment of the present invention introduces a Graph Attention Network (GAT) and a Gate Recurrent Unit (GRU) to capture short-term and long-term forgery traces.
[0036] (1) Short-term forgery trace encoder based on graph attention mechanism In any given At this moment, each landmark point can be To indicate that The value range of is from 1 to 68, corresponding to the key points at different positions on the face. and Represents the positions of all key points on the horizontal and vertical axes respectively, which can be integrated into a horizontal axis sequence and the ordinate sequence From a formal point of view, the graph structure is composed of a set of nodes. and edge set composed of . Node Collection Contains a set of graph nodes, represented by The edge set E represents the connection relationship between a series of facial key points. In order to describe these connection relationships more clearly, a size of The adjacency matrix of , where the elements Reflects the node and nodes If there is no connection between two nodes, then The value of is 0. Here, in order to capture short-term forgery traces, we use represents the time length of the local sequence. The spatiotemporal adjacency matrix is expressed as , by enumerating all possible neighbors in the local spatiotemporal neighborhood, the adjacency matrix is constructed as shown in formula (1).
[0037] (1) in Indicates size The cross-time graph of The relationship between the node with the current timestamp and its neighbors is established on the adjacent frames. For example, the node will be connected to In the processing of short-term facial key point sequences, the feature set of each facial key point in each video frame is used as input data to form a feature tensor .in, Represents the feature dimension of each node, represents the frame number, and represents the number of nodes. Consider the input sequence as the size of each time step A sliding window is used to generate a sequence of local facial key points at each time step, which can be expressed as formula (2): (2) In order to better distinguish similar facial movements, the attention mechanism is introduced to dynamically learn the importance of each node to other nodes based on the spatiotemporal relationship between nodes. First, for the lth layer of the graph attention network, its output feature is represented as , in , represents the input dimension of node features, Represents the number of graph network nodes. In this embodiment =68. In order to calculate the attention weight of the through edge, the node features are first converted into high-dimensional features, and the attention scores of adjacent key points are calculated through the self-attention mechanism. The calculation process is shown in the following formula (3): (3) in, , convert node features into high-dimensional features, Represents the weight matrix transformed by the feature The dimension of the high-dimensional feature space to which node features are mapped; Calculate the function for the attention mechanism. Then activate it with softmax and LeakyReLu to calculate the final normalized attention score , that is, formula (4): (4) Through the above calculated attention scores, we can get the model in the first The output of the layer is shown in formula (5). represents the activation function, Indicated in In the node Adjacent nodes. Indicates Layer node feature transformation weight matrix.
[0038] (5) The multi-head attention mechanism allows the model to learn attention coefficients in multiple representation subspaces, thereby enhancing the robustness of the self-attention learning process. An independent attention mechanism By performing transformations in multiple heads, the model can more comprehensively capture different features in the input data, improving the diversity and accuracy of information processing. The layer output is represented as , through The average of the attention head inputs is used as the final output value of the node, such as Figure 8 shown.
[0039] (2) Long-term forgery trace encoder based on bidirectional GRU In order to achieve the modeling of long-term forgery traces, the embodiment of the present invention uses a two-layer bidirectional GRU block to model long-term temporal information. Through the construction of the above short-term forgery traces, the output feature vector of the spatiotemporal graph attention network can be expressed as The short-term features are processed by two layers of bidirectional GRU for long-term modeling. Output features ,in , represents the number of hidden units in each layer, as shown in formula (6): (6) Finally, a softmax layer is used to map the feature representation of the facial key points in the video to the label for classification. A vector sequence maps to the probability of being real or fake . Then, the training objective of the model is to minimize its difference with the video label The cross entropy loss .
[0040] The forged speaking face detection system based on facial key points provided by the present invention is described below. The forged speaking face detection system based on facial key points described below and the forged speaking face detection method based on facial key points described above can be referenced to each other.
[0041] Fig. 9 FIG. 1 is a schematic diagram of the structure of a forged speaking face detection system based on facial key points provided by an embodiment of the present invention. Fig. 9 As shown, it includes: an extraction module 91, a training module 92 and a detection module 93, wherein: The extraction module 91 is used to obtain a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output the facial key points of the face; the training module 92 is used to cascade a short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder of a bidirectional GRU to construct a forged speaking face detection network, input the facial key points of the face into the forged speaking face detection network for prediction, and perform network training in combination with the authenticity labels of the video frames to obtain an optimized forged speaking face detection model; the detection module 93 is used to input the video to be detected into the optimized forged speaking face detection model to obtain a forged speaking face detection result.
[0042] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10 As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030 and a communication bus 1040, wherein the processor 1010, the communication interface 1020 and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 may call the logic instructions in the memory 1030 to execute a method for detecting a fake speaking face based on facial key points, the method comprising: obtaining a sample video frame of a facial image, extracting the sample video frame of the facial image to construct a facial key point adjacency network, and outputting facial key points of the face; cascading a short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder of a bidirectional GRU to construct a fake speaking face detection network, inputting the facial key points of the face into the fake speaking face detection network for prediction, combining the authenticity labels of the video frame for network training, and obtaining an optimized fake speaking face detection model; inputting the video to be detected into the optimized fake speaking face detection model to obtain a fake speaking face detection result.
[0043] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0044] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the forged speaking face detection method based on facial key points provided by the above methods. The method includes: obtaining a sample video frame of a face image, extracting the sample video frame of the face image to construct a facial key point adjacency network, and outputting facial key points of the face; cascading a short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder based on a bidirectional GRU to construct a forged speaking face detection network, inputting the facial key points of the face into the forged speaking face detection network for prediction, combining the authenticity labels of the video frames for network training, and obtaining an optimized forged speaking face detection model; inputting the video to be detected into the optimized forged speaking face detection model to obtain a forged speaking face detection result.
[0045] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the forged speaking face detection method based on facial key points provided by the above-mentioned methods, the method comprising: obtaining a sample video frame of a facial image, extracting the sample video frame of the facial image to construct a facial key point adjacency network, and outputting facial key points of the face; cascading a short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder based on a bidirectional GRU to construct a forged speaking face detection network, inputting the facial key points of the face into the forged speaking face detection network for prediction, combining the authenticity labels of the video frames for network training, and obtaining an optimized forged speaking face detection model; inputting the video to be detected into the optimized forged speaking face detection model to obtain a forged speaking face detection result.
[0046] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0047] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting fake speaking faces based on facial key points, characterized in that: include: Obtain a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output facial key points of the face; A short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder based on a bidirectional gated recurrent unit GRU are cascaded to construct a forged speaking face detection network, the facial key points of the face are input into the forged speaking face detection network for prediction, and the network is trained in combination with the authenticity labels of the video frames to obtain an optimized forged speaking face detection model; The video to be detected is input into the optimized forged speaking face detection model to obtain a forged speaking face detection result.
2. The method for detecting fake speaking faces based on facial key points according to claim 1, characterized in that: Acquire a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output facial key points of the face, including: The OpenFace open source tool is used to extract the location information of multiple facial key points of the face image sample video frame; Extracting forehead key points, eye key points, nose area and mouth area landmarks from the plurality of facial key point position information based on facial muscle movement of the human face to form a facial muscle movement network; Based on the forgery clues, the lip-jaw distance, lip opening and closing, nose-lip distance and facial contour are added to the facial muscle movement network of the human face to form the facial key point adjacency network.
3. The method for detecting fake speaking faces based on facial key points according to claim 1, characterized in that: The short-term forgery trace encoder of the graph attention mechanism includes: Obtain the coordinate pairs corresponding to the key points of each face at any time, and integrate all the coordinate pairs to form a horizontal coordinate sequence and a vertical coordinate sequence; Constructing a graph structure from the abscissa sequence and the ordinate sequence, wherein the graph structure includes a node set and an edge set; Determine an adjacency matrix of the graph structure, wherein a non-zero element in the adjacency matrix is a connection strength between any two nodes, and a zero element is no connection between any two nodes; Representing the adjacency matrix based on the local sequence time length to obtain a spatiotemporal adjacency matrix, wherein the spatiotemporal adjacency matrix includes a connection relationship between a current timestamp node and neighboring nodes; Taking the feature set of each facial key point in the spatiotemporal adjacency matrix as input data to form a feature tensor, the feature tensor includes the feature dimension of each node, the number of frames and the number of nodes, and taking the time length of the local sequence as a sliding window to generate a local facial key point sequence; Introducing an attention mechanism into the local facial key point sequence to form a graph attention network, obtaining output features of any layer in the graph attention network, converting the output features of any layer into high-dimensional features, calculating the high-dimensional features using the attention mechanism, and obtaining attention scores of adjacent key points; The attention score is calculated by LeakyReLu activation and normalized exponential function softmax in turn to obtain the final normalized attention score; The final normalized attention score, any layer node feature conversion weight matrix and any layer output feature in the graph attention network are summed within the range of adjacent nodes of any node in the spatiotemporal adjacency matrix to obtain the next layer output feature of any layer in the graph attention network; Calculate the average of the next layer output features of any layer in all graph attention networks in K attention heads to obtain the output value under the multi-head attention mechanism.
4. The method for detecting fake speaking faces based on facial key points according to claim 3, characterized in that: The bidirectional GRU long-term forgery trace encoder comprises: The output values under the multi-head attention mechanism are respectively input into two layers of bidirectional GRU to obtain two layers of bidirectional output features, where each layer of GRU contains the number of hidden units.
5. The method for detecting fake speaking faces based on facial key points according to claim 1, characterized in that: Inputting the facial key points of the human face into the fake speaking face detection network for prediction, and combining the authenticity labels of the video frames for network training to obtain an optimized fake speaking face detection model, including: The softmax layer is used to map the feature representation of the facial key points in the video to labels for classification, and the probability of authenticity or forgery is obtained; The optimized forged speaking face detection model is obtained by training based on minimizing the cross entropy loss between the real or forged probability and the video label.
6. A fake speaking face detection system based on facial key points, characterized in that: include: An extraction module is used to obtain a sample video frame of a face image, extract the sample video frame of the face image to construct a facial key point adjacency network, and output facial key points of the face; A training module, for cascading a short-term forgery trace encoder based on a graph attention mechanism and a long-term forgery trace encoder of a bidirectional GRU, constructing a forged speaking face detection network, inputting the facial key points of the face into the forged speaking face detection network for prediction, and performing network training in combination with the authenticity labels of the video frames to obtain an optimized forged speaking face detection model; The detection module is used to input the video to be detected into the optimized forged speaking face detection model to obtain the forged speaking face detection result.
7. The forged speaking face detection system based on facial key points according to claim 6, characterized in that: The extraction module is specifically used for: The OpenFace open source tool is used to extract the location information of multiple facial key points of the face image sample video frame; Extracting forehead key points, eye key points, nose area and mouth area landmarks from the plurality of facial key point position information based on facial muscle movement of the human face to form a facial muscle movement network; Based on the forgery clues, the lip-jaw distance, lip opening and closing, nose-lip distance and facial contour are added to the facial muscle movement network of the human face to form the facial key point adjacency network.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the forged speaking face detection method based on facial key points as claimed in any one of claims 1 to 5 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting a fake speaking face based on facial key points as claimed in any one of claims 1 to 5 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for detecting a fake speaking face based on facial key points as claimed in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Gesture recognition method and device based on space-time diagram convolutional neural network
CN112329525A
Method, system and device for realizing deep counterfeit face identification, processor and computer readable storage medium thereof
CN116012958A
Face-changing video tampering detection method and device based on face key point graph modeling
CN116434297A
Time series data anomaly detection method combining graph learning and double attention mechanism
CN118779804A
Cited By
Face forgery detection method and system based on space-time hyperbolic hypergraph
CN120318889A
A face forgery detection method and system based on spatio-temporal hypergraph
CN120318889B