Face spoofing detection method based on audiovisual emotional consistency
By extracting high-order emotional features from audio and video based on the audiovisual emotion consistency method, the problem of low detection accuracy and high computational complexity in existing technologies is solved, and efficient and accurate video-level face forgery detection is achieved.
Patent Information
- Application Number
- CN202511368689.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Existing technologies for video-level face forgery detection suffer from low detection accuracy, high computational complexity, inability to meet real-time interaction requirements, and neglect of the emotional consistency between audio and video.
By using a method based on audiovisual emotional consistency, high-order emotional features of audio and video are extracted. A multimodal emotional feature deep extractor and a cross-attention module are used, combined with micro-expression and contextual emotional cue extraction modules, to achieve synchronous verification of emotional expression consistency.
It significantly improves the ability to combat partial forgery, achieves efficient face forgery detection, and can accurately identify forged content in real-time interactive scenarios.
Smart Images

Figure CN120853242B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a face forgery detection method, in particular to a face forgery detection method based on audio-visual emotional consistency. BACKGROUND
[0002] With the rapid development of artificial intelligence in the field of computer vision, face forgery technology has been widely used in social media, such as: there are a large number of AI face changing or AI voice changing videos in platforms such as Douyin and bilibili. However, deep forgery of faces or voices often poses potential security risks, such as using face forgery videos for fraud or infringing on portrait rights. Therefore, face forgery detection technology is developed for entertainment and security fields to identify the authenticity of face images or videos.
[0003] Image-level face forgery detection technology mainly focuses on forgery traces in a single visual modality, such as forgery details in the RGB domain or frequency domain during upsampling. When facing video-level forgery detection, this type of method can only obtain the final result by judging each frame of the video, which is often time-consuming and inaccurate.
[0004] Current video-level face forgery detection technology adds a temporal consistency constraint based on image-level detection technology:
[0005] The prior art discloses an audio-visual forgery detection method and device (CN 114596609 A), which extracts features from face images and audio data in the video data to be tested through the image network branch and the audio network branch of the dual-flow network, respectively, and obtains inter-frame consistency features of the face images and time consistency features of the audio data based on the feature extraction results, so as to input the inter-frame consistency features of the face images and the time consistency features of the audio data into a prediction network to realize true or false detection of the video data to be tested. However, this technology has some defects: the dual-flow network separates the audio-visual modalities during feature extraction, only performs fusion prediction in the backend, and lacks the underlying correlation between modalities. When the single modality consistency of the forged content is very high, the detection accuracy is low. The inter-frame consistency features cannot capture rendering flaws, and the time consistency detection can be deceived by asynchronous synthesis strategies.
[0006] Multi-modal face forgery detection technology synchronously analyzes face liveness, voiceprint, lip language, micro-expression, etc. Typical strategies include: micro-expression + rPPG blood flow + breathing pattern; sensor noise + frequency domain flaws + light consistency; lip language synchronization + voiceprint + gesture behavior. Strategy 1 has the shortcoming of rPPG blood flow signal extraction delay, strategy 2 is limited by hardware and scene and relies heavily on original imaging data, and strategy 3 has high computational complexity. These three strategies cannot meet the detection needs in real-time interactive scenarios.
[0007] Other multi-modal fusion face forgery detection methods disclosed in the prior art also often focus on the consistency of lip movement and audio, the identity consistency of audio-visual, however, what is easily overlooked is that the audio and video should have the same emotional expression, that is, the video and audio should have the same emotion, and in addition, the emotional changes of the two should also be synchronous. SUMMARY
[0008] The present application provides a face forgery detection method based on audio-visual emotional consistency, aiming to mine emotional clues in audio and video, and realize efficient face forgery detection by judging the consistency of emotional expression and the synchronism of emotional change.
[0009] One of the technical solutions adopted by the present application is to provide a face forgery detection method based on audio-visual emotional consistency, comprising the following steps:
[0010] S1. Obtain a face video to be detected and its synchronous audio;
[0011] S2. Input the face video and its synchronous audio into a preprocessing module, and output corresponding video frame sequence and mel spectrum graph;
[0012] S3. Extract initial image features of the video frame sequence and initial audio features of the mel spectrum graph through an encoding module;
[0013] S4. Input the initial image features and initial audio features into a deep feature extraction module, the deep feature extraction module comprising at least one multi-modal emotional feature depth extractor, and performing:
[0014] Video-based emotional feature extraction: guided by the initial audio features, extract high-order video emotional features synchronized with audio from the initial image features;
[0015] Audio-based emotional feature extraction: guided by the initial image features, extract high-order audio emotional features synchronized with video from the initial audio features;
[0016] S5. Input the high-order video emotional features and high-order audio emotional features into an audio-visual emotional feature fusion module to generate audio-visual emotional discrimination features;
[0017] S6. Perform by a classification module:
[0018] Predict a first emotional tendency according to the high-order audio emotional features;
[0019] Predict a second emotional tendency according to the high-order video emotional features;
[0020] predicting a third emotional tendency according to the audio-visual emotional discrimination feature;
[0021] S7. Comparing the consistency of the first emotional tendency, the second emotional tendency and the third emotional tendency, and outputting a forgery detection result.
[0022] Further, the operation of extracting the video-based emotional feature includes:
[0023] A1, inputting the initial image feature into the micro-expression emotional clue extraction module to output a first intermediate video emotional feature;
[0024] A2, inputting the initial image feature and the initial audio feature into the cross-attention module to generate an audio emotional feature guided by the video emotional feature;
[0025] A3, inputting the audio emotional feature output by step A2 into the contextual emotional clue extraction module to output an audio emotional auxiliary feature;
[0026] A4, fusing the first intermediate video emotional feature and the audio emotional auxiliary feature through affine transformation to generate the high-order video emotional feature.
[0027] Further, the operation of extracting the audio-based emotional feature includes:
[0028] B1, inputting the initial audio feature into the contextual emotional clue extraction module to output a first intermediate audio emotional feature;
[0029] B2, inputting the initial audio feature and the initial image feature into the cross-attention module to generate a video emotional feature guided by the audio emotional feature;
[0030] B3, inputting the video emotional feature output by step B2 into the micro-expression emotional clue extraction module to output a video emotional auxiliary feature;
[0031] B4, fusing the first intermediate audio emotional feature and the video emotional auxiliary feature through affine transformation to generate the high-order audio emotional feature.
[0032] Further, the calculation formula of the cross-attention module is:
[0033] where d is the dimension of the key vector .
[0034] In step A2, A is the initial image feature and B is the initial audio feature;
[0035] In step B2, A is the initial audio feature and B is the initial image feature.
[0036] Further, the operation formula of the affine transformation is:
[0037] ;
[0038] In step A4, F is a first intermediate video emotion feature, and obtained by two linear layers from the audio emotion auxiliary feature, and Output is a generated high-order video emotion feature;
[0039] In step B4, F is a first intermediate audio emotion feature, and obtained by two linear layers from the video emotion auxiliary feature, and Output is a generated high-order audio emotion feature.
[0040] Further, the process of the micro-expression emotion clue extraction module processing input features is:
[0041] The input features are sequentially subjected to a first deep separable convolution unit, a first layer normalization unit, a state space dual calculation module, a second deep separable convolution unit, a second layer normalization unit and a feedforward neural network unit, and finally output processed features;
[0042] In step A1, the input of the micro-expression emotion clue extraction module is an initial image feature, and the output is a first intermediate video emotion feature;
[0043] In step B3, the input of the micro-expression emotion clue extraction module is the video emotion feature output in step B2, and the output is a video emotion auxiliary feature.
[0044] Further, the process of the state space dual calculation module processing input features is:
[0045] C1, linearly transforming the input features through a first linear layer;
[0046] C2, performing a deep separable convolution operation on the transformed features to generate new features;
[0047] C3, converting the new features into an input matrix X through a SiLU activation function;
[0048] C4, in the state space dual calculation, performing:
[0049] , wherein A is a state transition matrix, X is an input matrix, B is an input state matrix, and C is an output matrix;
[0050] C5, sequentially subjecting the output Y to layer normalization processing and second linear layer transformation, and finally output processed features.
[0051] Further, the process of the context sentiment clue extraction module processing the input features is as follows:
[0052] The input features are sequentially input into the xLSTM system, the memory attention network and the gate circuit system, and finally the processed features are output.
[0053] In step A3, the input of the context sentiment clue extraction module is the audio sentiment feature output in step A2, and the output is the audio sentiment auxiliary feature.
[0054] In step B1, the input of the context sentiment clue extraction module is the initial audio feature, and the output is the first intermediate audio sentiment feature.
[0055] Further, in the xLSTM system, the input audio features are sequentially arranged in time sequence and denoted as A 1 …A t-1 , A t …A T In the memory attention network, the attention calculation formula is:
[0056] ,
[0057] ,
[0058] wherein, is the input audio feature at t-1 and t in the xLSTM system, i.e. A t-1 and A t , is a neural network, is obtained by the neural network , is the Hadamard product, is and is obtained by the Hadamard product.
[0059] Further, the audio-visual sentiment feature fusion module performs the following operations:
[0060] D1, input the base feature F and the high-order video sentiment feature into a first cross-attention module to generate a fused high-order video sentiment feature;
[0061] D2, input the base feature F and the high-order audio sentiment feature into a second cross-attention module to generate a fused high-order audio sentiment feature;
[0062] D3, input the fused high-order video sentiment feature and the fused high-order audio sentiment feature into a third cross-attention module to generate the audio-visual sentiment discrimination feature.
[0063] Further, the base feature F is generated by the following way:
[0064] The base feature F dimension is consistent with the high-order video emotional feature and the high-order audio emotional feature, the value of each dimension is sampled from a uniform distribution, and a tensor is constructed.
[0065] Further, the first cross-attention module, the second cross-attention module and the third cross-attention module perform:
[0066] In the first cross-attention module, the key and the value are provided by the high-order video emotional feature, the query is provided by the base vector, and the fused high-order video emotional feature is generated;
[0067] In the second cross-attention module, the key and the value are provided by the high-order audio emotional feature, the query is provided by the base vector, and the fused high-order audio emotional feature is generated;
[0068] In the third cross-attention module, the key and the value are provided by the fused high-order video emotional feature, the query is provided by the fused high-order audio emotional feature, and the audio-visual emotional discriminative feature is generated.
[0069] Further, the classification module generates the first emotional tendency, the second emotional tendency and the third emotional tendency by a three-branch fully connected network, and each branch shares the bottom network parameters.
[0070] Further, the deep extraction module comprises K multi-modal emotional feature deep extractors connected in series (K≥2), the output of the previous extractor is taken as the input of the next extractor, and multi-level emotional feature extraction is realized.
[0071] The face forgery detection method based on audio-visual emotional consistency has the following advantages:
[0072] 1. A multi-modal emotional feature deep extractor based on audio-visual is designed, which is guided by the audio feature when extracting the video feature and guided by the video feature when extracting the audio feature, so as to mine the cross-modal emotional synchronization clues (high-order video emotional feature and high-order audio emotional feature) and realize the synchronization verification of the emotional expression of the two.
[0073] 2. By comparing the consistency of the first emotional tendency, the second emotional tendency and the third emotional tendency, a multiple verification of emotional consistency is constructed, which significantly improves the ability to resist local forgery. BRIEF DESCRIPTION OF DRAWINGS
[0074] Figure 1 is a structural schematic diagram of a face forgery detection model in the first embodiment of the application;
[0075] Figure 2 is a structural schematic diagram of the emotion feature extraction module mainly based on video in the first embodiment of the present application;
[0076] Figure 3 is a structural schematic diagram of the emotion feature extraction module mainly based on audio in the first embodiment of the present application;
[0077] Figure 4 is a structural schematic diagram of the micro-expression emotion cue extraction module in the first embodiment of the present application;
[0078] Figure 5 is a structural schematic diagram of the state space dual calculation module in the first embodiment of the present application;
[0079] Figure 6 is a structural schematic diagram of the context emotion cue extraction module in the first embodiment of the present application;
[0080] Figure 7 is a structural schematic diagram of the audio-visual emotion feature fusion module in the first embodiment of the present application. DETAILED DESCRIPTION
[0081] The preferred embodiments of the present application will be described in detail below with reference to the accompanying drawings, so that the advantages and features of the present application can be more easily understood by those skilled in the art, and the protection scope of the present application can be more clearly defined.
[0082] The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and cannot be used to limit the protection scope of the present application.
[0083] In the first embodiment of the present application, a face forgery detection model is constructed, which is used to implement a face forgery detection method based on audio-visual emotion consistency.
[0084] Please refer to Figure 1 , the structure of the face forgery detection model is as follows:
[0085] The preprocessing module is used to obtain a video frame sequence and a mel spectrum diagram respectively according to the face video to be detected and its synchronous audio;
[0086] The encoding module is used to generate initial image features and initial audio features according to the video frame sequence and the mel spectrum diagram;
[0087] The depth extraction module includes K multi-modal emotion feature depth extractors (K≥2) connected in series, and the output of a previous extractor is used as the input of a subsequent extractor. The initial image features and the initial audio features generated by the encoding module are subjected to multi-level bidirectional mutual guidance of emotion feature extraction, so as to generate final high-order audio emotion features and high-order video emotion features;
[0088] an audiovisual emotion feature fusion module, configured to generate an audiovisual emotion discriminant feature according to the high-order audio emotion feature and the high-order video emotion feature;
[0089] a classification module, configured to respectively predict three emotion tendencies according to the high-order audio emotion feature, the high-order video emotion feature and the audiovisual emotion discriminant feature, and compare the consistency of the three emotion tendencies, and output a result of the forgery detection.
[0090] In the encoding module:
[0091] The video frame sequence and the mel spectrum graph are processed by using a video encoder and an audio encoder:
[0092] The function of the video encoder can be implemented by using an existing mainstream convolutional neural network or an optimized version thereof, such as ResNet, SCNet, Xception, etc., and the architecture thereof is composed of multiple convolutional units and down-sampling units, the convolutional units and the down-sampling units are alternately connected in series, and the initial image feature is generated.
[0093] The function of the audio encoder can be implemented by using an existing mainstream sequence encoder, such as RNN or LSTM, etc., and the initial audio feature is generated.
[0094] The structure of the multi-modal emotion feature deep extractor is as follows:
[0095] a video-based emotion feature extraction module, configured to extract a high-order video emotion feature synchronized with audio;
[0096] an audio-based emotion feature extraction module, configured to extract a high-order audio emotion feature synchronized with video;
[0097] The video-based emotion feature extraction module and the audio-based emotion feature extraction module each include a micro-expression emotion clue extraction module, a cross-attention module, a context emotion clue extraction module and an affine transformation fusion module.
[0098] For the first multi-modal emotion feature deep extractor connected in series, the input thereof is and , and the following is performed:
[0099] video-based emotion feature extraction:
[0100] A1, input the initial image feature into the micro-expression emotion clue extraction module, and output a first intermediate video emotion feature;
[0101] A2, input the initial image feature and the initial audio feature into the cross-attention module, and generate an audio emotion feature guided by the video emotion feature;
[0102] A3, input the audio emotion feature output by step A2 into the contextual emotion clue extraction module, and output an audio emotion auxiliary feature;
[0103] A4, fuse the first intermediate video emotion feature and the audio emotion auxiliary feature through affine transformation to generate a high-order video emotion feature;
[0104] Audio-based emotion feature extraction:
[0105] B1, input the initial audio feature into the contextual emotion clue extraction module to output a first intermediate audio emotion feature;
[0106] B2, input the initial audio feature and the initial image feature into the cross-attention module to generate a video emotion feature guided by the audio emotion feature;
[0107] B3, input the video emotion feature output by step B2 into the micro-expression emotion clue extraction module to output a video emotion auxiliary feature;
[0108] B4, fuse the first intermediate audio emotion feature and the video emotion auxiliary feature through affine transformation to generate a high-order audio emotion feature.
[0109] Wherein, the calculation formula of the cross-attention module is:
[0110] Wherein d is the dimension of the key vector ;
[0111] In step A2, A is the initial image feature , and B is the initial audio feature ;
[0112] In step B2, A is the initial audio feature , and B is the initial image feature .
[0113] Wherein, the operation formula of affine transformation is:
[0114] ;
[0115] In step A4, F is the first intermediate video emotion feature, and are obtained by two linear layers from the audio emotion auxiliary feature, and Output is the generated high-order video emotion feature;
[0116] In step B4, F is the first intermediate audio emotion feature, and The video emotion auxiliary feature is obtained by two linear layers, and the output is the generated high-order audio emotion feature.
[0117] Please refer to Figure 2 and Figure 3 , for any subsequent multimodal emotion feature depth extractor in series, the input is , respectively representing the high-order video emotion feature and the high-order audio emotion feature output by the last multimodal emotion feature depth extractor, performing:
[0118] Video-based emotion feature extraction:
[0119] E1, the high-order video emotion feature output by the last multimodal emotion feature depth extractor is input into the micro-expression emotion clue extraction module, and the first intermediate video emotion feature is output.
[0120] E2, the high-order video emotion feature output by the last multimodal emotion feature depth extractor and the high-order audio emotion feature are input into the cross-attention module to generate video emotion feature guided audio emotion feature .
[0121] E3, the audio emotion feature output by step E2 is input into the context emotion clue extraction module to output audio emotion auxiliary feature .
[0122] E4, the first intermediate video emotion feature and the audio emotion auxiliary feature are fused by affine transformation to generate high-order video emotion feature .
[0123] Audio-based emotion feature extraction:
[0124] F1, the high-order audio emotion feature output by the last multimodal emotion feature depth extractor is input into the context emotion clue extraction module to output the first intermediate audio emotion feature .
[0125] F2, the high-order audio emotion feature output by the last multimodal emotion feature depth extractor and the high-order video emotion feature are input into the cross-attention module to generate audio emotion feature guided video emotion feature .
[0126] F3, the video emotion feature output by step F2 is fused with the audio emotion feature output by step F1 to obtain a first intermediate audio emotion feature The input micro-expression emotion clue extraction module outputs a video emotion auxiliary feature
[0127] F4, the first intermediate audio emotion feature is fused by affine transformation to obtain a high-order audio emotion feature and the video emotion auxiliary feature .
[0128] The calculation formula of the cross-attention module is as follows:
[0129] where d is the dimension of the key vector .
[0130] In step E2, A is the high-order video emotion feature output by the previous multi-modal emotion feature depth extractor , and B is the high-order audio emotion feature output by the previous multi-modal emotion feature depth extractor .
[0131] In step F2, A is the high-order audio emotion feature output by the previous multi-modal emotion feature depth extractor , and B is the high-order video emotion feature output by the previous multi-modal emotion feature depth extractor .
[0132] The calculation formula of the affine transformation is as follows:
[0133] .
[0134] In step E4, F is the first intermediate video emotion feature , and are obtained from the audio emotion auxiliary feature by two linear layers, and Output is the generated high-order video emotion feature .
[0135] In step F4, F is the first intermediate audio emotion feature , and are obtained from the video emotion auxiliary feature by two linear layers, and Output is the generated high-order audio emotion feature .
[0136] Please refer to Figure 4 The micro-expression emotional clue extraction module comprises a first deep separable convolution unit, a first layer normalization unit, a state space dual calculation module, a second deep separable convolution unit, a second layer normalization unit and a feedforward neural network unit.
[0137] The state space dual calculation module is configured to capture the dynamic relationship of the features in the time and space dimensions, mine key features related to the micro-expression, and enhance the feature mining and learning ability of the module by using the long sequence deep feature extraction capability of the spatial state model.
[0138] Referring to Figure 5 The state space dual calculation module performs:
[0139] C1, linearly transforming the input features through a first linear layer;
[0140] C2, performing a deep separable convolution operation on the transformed features to generate new features;
[0141] C3, converting the new features into an input matrix X through a SiLU activation function;
[0142] C4, performing in the state space dual calculation:
[0143] wherein A is a state transition matrix, X is an input matrix, B is an input state matrix, and C is an output matrix;
[0144] C5, sequentially performing layer normalization processing and second linear layer transformation on the output Y, and finally outputting the processed features.
[0145] Referring to Figure 6 The contextual emotional clue extraction module comprises an xLSTM system, a memory attention network and a gating circuit system.
[0146] In the xLSTM system, the input audio features are sequentially arranged in time sequence and denoted as … , … , or According to the time sequence arrangement, as , represents the audio emotional clue after k multi-modal emotional feature deep extractors at t time;
[0147] In the memory attention network, the attention calculation formula is: , wherein is the input audio feature at t-1 and t time in the xLSTM system, i.e. and , For the neural network, the feature splicing of adjacent time is received as input, and the attention coefficient is output, which quantifies the cross-view interaction correlation of different view memory dimensions, For Through the neural network Get, For Hadamard product, For With Get through Hadamard product;
[0148] Through the gate circuit system, it contains 3 gate circuit units , And , respectively generate the reserved gate signal, the update gate signal and the update proposal signal, wherein the reserved gate decides how much information of the memory at t-1 time is reserved, the update gate decides how much information of the current update proposal signal is integrated into the memory, and the two jointly generate the memory state at t time, realizing the dynamic storage and update of the context emotional clue information in the time dimension.
[0149] Please refer to Figure 7 , the structure of the audio-visual emotion feature fusion module is as follows:
[0150] The first cross-attention module is used to generate the fused high-order video emotion feature based on the basic feature F and the high-order video emotion feature;
[0151] The second cross-attention module is used to generate the fused high-order audio emotion feature based on the basic feature F and the high-order audio emotion feature;
[0152] The third cross-attention module is used to generate the audio-visual emotion discrimination feature based on the fused high-order video emotion feature and the fused high-order audio emotion feature;
[0153] Among them, the dimension of the basic feature F is consistent with the high-order video emotion feature and the high-order audio emotion feature, and the value of each dimension is sampled from a uniform distribution to form a tensor;
[0154] In the first cross-attention module, the key and value are provided by the high-order video emotion feature, and the query is provided by the basic vector;
[0155] In the second cross-attention module, the key and value are provided by the high-order audio emotion feature, and the query is provided by the basic vector;
[0156] In the third cross-attention module, the key and value are provided by the fused high-order video emotion feature, and the query is provided by the fused high-order audio emotion feature.
[0157] In the classification module:
[0158] The first sentiment tendency P1, the second sentiment tendency P2 and the third sentiment tendency P3 are respectively generated by the first classifier, the second classifier and the third classifier, the input size of the full connection layer of each classifier is the size of the classification feature, and the output size is 1, which respectively represents the predicted sentiment tendency;
[0159] The fourth classifier is used for consistency comparison of the above sentiment tendencies, the input size of the full connection layer thereof is the size of the classification feature, and the output size is 2, which respectively represents the probability of being predicted as real and fake;
[0160] The consistency comparison method is that |P1-P3|, |P2-P3| are respectively compared, and finally two consistency errors P=||P1-P3|, |P2-P3|| are obtained, the value is less than a threshold value, and is true, and is false if greater than the threshold value.
[0161] The training method of the above face forgery detection model comprises the following steps:
[0162] A training set is obtained, the training set includes real and face video and audio data generated by existing forgery technology, and corresponding labels, wherein the labels take values of 0 and 1, wherein 0 represents real data, and 1 represents fake false data;
[0163] The real and face video and audio data generated by the existing forgery technology are input into the pre-constructed face forgery detection model to obtain a model prediction result;
[0164] According to the prediction result of the fourth classifier and the corresponding label, a cross-entropy loss function is calculated:
[0165] , wherein, represents the real label corresponding to the input, represents the model prediction result;
[0166] At the same time, according to the prediction results of the first, second and third classifiers and the corresponding labels, a multi-task learning loss function is calculated:
[0167] , wherein N is the total number of training samples, m represents the task type, v, a respectively represent the video sentiment prediction result and the audio sentiment prediction result, e represents the audio-visual comprehensive sentiment prediction result, vae represents the audio-visual comprehensive sentiment consistency discrimination prediction result, is a cross-entropy loss function, represents the real label corresponding to the input, represents the model prediction result;
[0168] The face forgery detection model is iteratively updated and trained based on a gradient descent method, and the detection model at the minimum loss function of the model is taken as the trained detection model.
[0169] The obtained data set includes directly downloading a public data set and constructing a private data set:
[0170] The public data set includes DFDC, FakeAVCeleb, etc.
[0171] The private data set is constructed by the following steps:
[0172] Face videos are obtained from a video website or face videos are collected;
[0173] The face is detected by using a Dlib tool;
[0174] The face is aligned by using an affine transformation;
[0175] The face region is cropped after setting a fixed margin (such as 1.3) for the face region;
[0176] The cropped face image is scaled to a fixed pixel size (such as 256 pixels in length and width);
[0177] Meanwhile, the audio file corresponding to the video is converted by using Librosa tool, and the audio is converted by frame division, windowing and fast Fourier transform, and then a mel spectrum graph is drawn.
[0178] The face forgery detection method based on audio-visual emotion consistency has the following beneficial effects:
[0179] 1. A multi-modal emotion feature depth extractor based on audio-visual is designed, from the requirement of audio-visual emotion expression consistency and change synchronization, when extracting video features, the audio features are used as a guide, when extracting audio features, the video features are used as a guide, to mine cross-modal emotion synchronization clues (high-order video emotion features and high-order audio emotion features), and realize the synchronous verification of the emotion expression of the two;
[0180] 2. In the micro-expression and context emotion clue extraction module, an efficient long short-term memory network and a state space model are used to realize the interaction and mining of audio-visual level emotion clues;
[0181] 3. The audio-visual emotion feature fusion module based on attention bottleneck combines the multi-task learning goal, realizes the fine-grained prediction of single-modal emotion and the comprehensive discrimination of multi-modal emotion;
[0182] 4. By comparing the consistency of the first emotion tendency, the second emotion tendency and the third emotion tendency, a multiple verification of emotion consistency is constructed, which significantly improves the ability to resist local forgery.
[0183] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0184] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing device or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0185] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0187] The above description is only preferred embodiment of the present application, it should be pointed out that for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A face forgery detection method based on audiovisual emotional consistency, characterized in that, Includes the following steps: S1. Obtain the face video to be detected and its synchronized audio; S2. Input the face video and its synchronized audio into the preprocessing module, and output the corresponding video frame sequence and Mel spectrogram; S3. Extract the initial image features of the video frame sequence and the initial audio features of the Mel spectrogram using the encoding module; S4. Input the initial image features and initial audio features into the depth extraction module, which includes at least one multimodal emotion feature depth extractor, and execute: Video-based emotion feature extraction: Guided by the initial audio features, extract high-order video emotion features synchronized with the audio from the initial image features; Audio-based emotion feature extraction: Guided by the initial image features, extract high-order audio emotion features synchronized with the video from the initial audio features; S5. Input the high-order video emotion features and high-order audio emotion features into the audiovisual emotion feature fusion module to generate audiovisual emotion discrimination features; S6. Executed via the classification module: Predict the first emotional tendency based on the aforementioned high-order audio emotional features; Predict the second sentiment tendency based on the aforementioned high-order video sentiment features; Predict the third emotional tendency based on the aforementioned audiovisual emotion discrimination features; S7. Compare the consistency of the first emotional tendency, the second emotional tendency and the third emotional tendency, and output the forgery detection result; in, The video-based emotion feature extraction operation includes: A1. Input the initial image features into the micro-expression emotion clue extraction module and output the first intermediate video emotion features; A2. Input the initial image features and initial audio features into the cross-attention module to generate audio emotion features guided by video emotion features; A3. Input the audio emotional features output in step A2 into the contextual emotional cue extraction module to output audio emotional auxiliary features; A4. The higher-order video emotion features are generated by fusing the first intermediate video emotion features and the audio emotion auxiliary features through affine transformation. The operation of extracting emotional features, primarily based on audio, includes: B1. Input the initial audio features into the contextual sentiment cue extraction module and output the first intermediate audio sentiment features; B2. Input the initial audio features and initial image features into the cross-attention module to generate video emotion features guided by audio emotion features; B3. Input the video emotion features output in step B2 into the micro-expression emotion cue extraction module, and output the video emotion auxiliary features; B4. The higher-order audio emotional features are generated by fusing the first intermediate audio emotional features and the video emotional auxiliary features through affine transformation. The audiovisual emotion feature fusion module performs the following operations: D1. Input the basic feature F and the higher-order video sentiment feature into the first cross-attention module to generate a fused higher-order video sentiment feature. D2. Input the basic feature F and the higher-order audio emotion feature into the second cross-attention module to generate a fused higher-order audio emotion feature. D3. Input the fused high-order video sentiment features and the fused high-order audio sentiment features into the third cross-attention module to generate the audiovisual sentiment discrimination features.
2. The face forgery detection method based on audiovisual emotional consistency according to claim 1, characterized in that, The calculation formula for the cross-attention module is as follows: , where d is the key vector K B dimensionality; In step A2, A represents the initial image features, and B represents the initial audio features; In step B2, A is the initial audio feature, and B is the initial image feature.
3. The face forgery detection method based on audiovisual emotional consistency according to claim 1, characterized in that, The formula for the affine transformation is as follows: ; In step A4, F represents the sentiment feature of the first intermediate video. and The audio emotion-aiding features are obtained through two linear layers, and the output is the generated high-order video emotion features; In step B4, F represents the first intermediate audio emotional feature. and The video emotion-aiding features are obtained through two linear layers, and Output is the generated high-order audio emotion features.
4. The face forgery detection method based on audiovisual emotional consistency according to claim 1, characterized in that, The process by which the micro-expression emotion cue extraction module processes input features is as follows: The input features are sequentially processed through a first depthwise separable convolutional unit, a first-layer normalization unit, a state-space dual computation module, a second depthwise separable convolutional unit, a second-layer normalization unit, and a feedforward neural network unit, and finally output the processed features. In step A1, the input of the micro-expression emotion cue extraction module is the initial image features, and the output is the first intermediate video emotion features; In step B3, the input to the micro-expression emotion cue extraction module is the video emotion features output in step B2, and the output is the video emotion auxiliary features.
5. The face forgery detection method based on audiovisual emotional consistency according to claim 4, characterized in that, The process by which the state-space dual computation module processes input features is as follows: C1. Perform a linear transformation on the input features through the first linear layer; C2. Perform depthwise separable convolution on the transformed features to generate new features; C3. Convert the new features into an input matrix X using the SiLU activation function; C4. Execute in state-space dual computation: Where A is the state transition matrix, X is the input matrix, B is the input state matrix, and C is the output matrix; C5. The output Y is processed sequentially through layer normalization and the second linear layer transformation to finally output the processed features.
6. The face forgery detection method based on audiovisual emotional consistency according to claim 1, characterized in that, The process by which the contextual sentiment cue extraction module processes input features is as follows: The input features are sequentially passed through an xLSTM system, a memory attention network, and a gating circuit system, and finally the processed features are output. In step A3, the input of the contextual emotion cue extraction module is the audio emotion feature output in step A2, and the output is the audio emotion auxiliary feature; In step B1, the input of the contextual sentiment cue extraction module is the initial audio features, and the output is the first intermediate audio sentiment features.
7. The face forgery detection method based on audiovisual emotional consistency according to claim 1, characterized in that, The deep extraction module contains K multimodal sentiment feature deep extractors (K≥2) connected in series, with the output of the previous extractor serving as the input of the next extractor.
8. The face forgery detection method based on audiovisual emotional consistency according to claim 7, characterized in that, For any subsequent multimodal sentiment feature deep extractor in the cascaded sequence, its input is: , These represent the high-order video sentiment features and high-order audio sentiment features output by the previous multimodal sentiment feature deep extractor, respectively. The following steps are performed: Emotion feature extraction based primarily on videos: E1. Utilize the high-order video sentiment features output by the previous multimodal sentiment feature deep extractor. Input the micro-expression emotion cue extraction module and output the emotion features of the first intermediate video. ; E2. Utilize the high-order video sentiment features output from the previous multimodal sentiment feature deep extractor. With higher-order audio emotional features Input the cross-attention module to generate audio sentiment features guided by video sentiment features. ; E3. Analyze the audio emotional features output in step E2. Input the contextual sentiment cue extraction module and output audio sentiment auxiliary features. ; E4. Fusing emotional features from the first intermediate video using affine transformation. With audio emotion Generate high-order video sentiment features ; Emotional feature extraction, primarily based on audio: F1. Develop the high-order audio emotion features output from the previous multimodal emotion feature deep extractor. Input the contextual sentiment cue extraction module and output the sentiment features of the first intermediate audio. ; F2. Develop the high-order audio emotion features output from the previous multimodal emotion feature deep extractor. With higher-order video sentiment features Input cross-attention module to generate video sentiment features guided by audio sentiment features. ; F3. Analyze the video emotion features output in step F2. Input micro-expression emotion cue extraction module, output video emotion auxiliary features ; F4. Fuse the emotional features of the first intermediate audio through affine transformation. With video emotion-assisted features Generate high-order audio emotional features .
9. A face forgery detection method based on audiovisual emotional consistency according to claim 8, characterized in that, The calculation formula for the cross-attention module is as follows: , where d is the key vector K B dimensionality; In step E2, A represents the high-order video sentiment features output by the previous multimodal sentiment feature deep extractor. B represents the high-order audio sentiment feature output by the previous multimodal sentiment feature deep extractor. ; In step F2, A is the high-order audio emotion feature output by the previous multimodal emotion feature deep extractor. B represents the high-order video sentiment features output by the previous multimodal sentiment feature deep extractor. .
10. A face forgery detection method based on audiovisual emotional consistency according to claim 8, characterized in that, The formula for the affine transformation is as follows: ; In step E4, F represents the sentiment feature of the first intermediate video. , and Audio Emotional Auxiliary Features The output is obtained through two linear layers, and the high-order video sentiment features are generated as output. ; In step F4, F represents the first intermediate audio emotional feature. , and Video Emotional Auxiliary Features The high-order audio sentiment features are obtained through two linear layers, with Output being the generated high-order audio sentiment features. .
11. A face forgery detection method based on audiovisual emotional consistency according to claim 8, characterized in that, The process by which the micro-expression emotion cue extraction module processes input features is as follows: The input features are sequentially processed through a first depthwise separable convolutional unit, a first-layer normalization unit, a state-space dual computation module, a second depthwise separable convolutional unit, a second-layer normalization unit, and a feedforward neural network unit, and finally output the processed features. The process by which the state-space dual computation module processes the input features is as follows: C1. Perform a linear transformation on the input features through the first linear layer; C2. Perform depthwise separable convolution on the transformed features to generate new features; C3. Convert the new features into an input matrix X using the SiLU activation function; C4. Execute in state-space dual computation: Where A is the state transition matrix, X is the input matrix, B is the input state matrix, and C is the output matrix; C5. The output Y is processed sequentially through layer normalization and the second linear layer transformation to finally output the processed features.
12. The face forgery detection method based on audiovisual emotional consistency according to claim 8, characterized in that, The process by which the contextual sentiment cue extraction module processes input features is as follows: The input features are sequentially passed through an xLSTM system, a memory attention network, and a gating circuit system, and finally the processed features are output. In the xLSTM system, the input audio features are arranged sequentially in time and denoted as follows: … , … , or Based on its temporal arrangement as a function in the xLSTM system , representing the audio emotional cue after passing through k multimodal emotional feature deep extractors at time t; In the aforementioned memory-attention network, the attention calculation formula is: , ,in, The input audio features at times t-1 and t in the xLSTM system are, i.e. and , For neural networks, for Through neural networks get, For Hadamard products, for and Obtained through the Hadamard product.
Citation Information
Patent Citations
Audio-visual forgery detection method and device
CN114596609A
A video emotion recognition method integrating facial expression recognition and voice emotion recognition
CN109409296A
Counterfeit face video detection method and device based on multi-modal behavior consistency, electronic equipment, storage medium and program product
CN120580567A