Audio and video depth forgery detection method based on quality perception and multi-scale alignment

By employing a visual quality perception and multi-scale alignment-based audio and video depth forgery detection method, the problem of insufficient robustness in existing technologies is solved, achieving higher detection accuracy and adaptability, and making it applicable to a variety of real-world scenarios.

CN121009341APending Publication Date: 2025-11-25NANJING UNIV OF SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511122479.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing methods for detecting deepfake audio and video lack robustness in complex real-world environments, struggle to adapt to visual quality degradation, lack multi-scale cross-modal semantic and physiological level alignment features, and are not accurate enough in judging uncertain inputs, resulting in a high false positive rate.

Method used

A visual quality perception mechanism is introduced to generate a spatial reliability mask. Through multi-scale cross-modal fusion and self-supervised uncertainty calibration loss, the robustness and accuracy of the detection model are enhanced.

Benefits of technology

It improves the model's detection performance in low-quality and complex environments, reduces the false alarm rate, and enhances its cross-dataset generalization ability and practical deployment reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009341A_ABST
    Figure CN121009341A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video depth forgery detection method based on quality perception and multi-scale alignment, and the method comprises the following steps: coding a synchronous audio and video sequence, and obtaining a frame-level visual feature, a facial action unit and a phoneme-level voice representation; a visual quality evaluation module is introduced to generate a spatial reliability mask, and quality weighting is carried out on the visual features; designing a global-local multi-scale cross-modal alignment mechanism, performing bidirectional cross-attention modeling on voice and face dynamic synchronization globally, and performing physiological coupling alignment on phonemes and face action units locally; and an uncertainty perception reasoning and calibration scheme is provided, adaptive temperature scaling is carried out according to quality and consistency, and uncertainty calibration is carried out by self-supervision loss. According to the method, the problems of insufficient robustness and excessive self-confidence misjudgment of an existing method in a low-quality video and high-synchronization counterfeit scene are solved, and the cross-dataset generalization capability and the actual deployment reliability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of video content security and digital media forensics. Background Technology

[0002] In recent years, artificial intelligence generative models, particularly deep learning-based generative adversarial networks (GANs) and diffusion models, have made groundbreaking progress in the field of audio and video synthesis, enabling the generation of highly realistic facial images, speech, and entire video clips. These generative technologies have driven the development of multiple industries, including virtual reality, film and television production, and telecommunications. However, their misuse has also spawned a large amount of "deepfake" content, especially audio and video deepfakes, posing a serious threat to public safety, media dissemination, political stability, and personal privacy.

[0003] Deepfake audio and video content typically creates videos that closely resemble real people by synthesizing facial expressions, lip movements, and voice content. These videos are difficult to detect with the naked eye and are extremely deceptive. With the widespread adoption of this technology and the lowering of barriers to entry, deepfake videos can be used for illegal purposes such as online fraud, public opinion manipulation, identity theft, and defamation, raising widespread public concerns about the authenticity of digital media. Therefore, developing efficient, reliable, and practically adaptable deepfake audio and video detection technologies has become a global research hotspot.

[0004] Most existing detection methods rely on the temporal synchronization between audio and video, assuming that the speech content and facial movements in real videos are highly consistent in time and semantics. For example, some methods determine whether a video is forged by comparing the alignment of lip movements with the speech waveform; others utilize cross-modal attention mechanisms to model the semantic consistency between audio and video to identify potential anomalies. While these methods have achieved good results on standard test sets, they still face the following challenges in complex real-world environments: First, real videos themselves suffer from visual quality degradation, especially in low-quality acquisition scenarios such as video conferencing, social media, and mobile device shooting, often accompanied by blurring, compression, occlusion, and low resolution. These problems severely weaken the extractability of facial motion features, rendering synchronous detection algorithms that rely on high-quality visual input ineffective.

[0005] Second, with the development of generative techniques, modern forgery models can achieve highly accurate audio-visual alignment. The synthesized speech and facial movements are highly consistent in time and semantics, greatly challenging traditional detection mechanisms based on audio-visual alignment consistency. Such forgeries can often "fool" conventional detection models and are more aggressive.

[0006] Third, existing methods mostly employ single-scale or global modeling strategies, neglecting the physiological coupling between speech elements and fine-grained facial movements (such as lip opening and closing, muscle changes, etc.), resulting in insufficient ability to detect high-level and detailed forgeries.

[0007] Fourth, most existing methods lack adaptability and uncertainty modeling mechanisms for changes in the quality of input data, making it difficult to make stable judgments on fuzzy, abnormal, or boundary samples. This can easily lead to overconfident misjudgments, affecting the practicality and robustness of the model.

[0008] Based on the above analysis, there is an urgent need for a deepfake detection method for audio and video with the following capabilities: (1) the ability to perceive and adapt to visual quality degradation and filter low-confidence image regions; (2) the ability to capture cross-modal multi-scale, fine-grained semantic and physiological hierarchical alignment features; (3) the ability to evaluate and express the model's predictive confidence in uncertain inputs; and (4) the ability to generalize well across datasets and the feasibility of practical deployment. This invention addresses the core bottlenecks in current technologies by proposing a deepfake detection method for audio and video based on quality perception and multi-scale alignment. By introducing a visual quality assessment mechanism, a multi-scale cross-modal fusion structure, and uncertainty calibration loss, it effectively enhances the detection model's ability to identify forged content and its adaptability to complex real-world environments. This method possesses high innovation and practical value, meeting the urgent needs of society and industry for ensuring the authenticity of audio and video content. Summary of the Invention

[0009] This invention proposes a method for detecting deepfake audio and video based on quality perception and multi-scale alignment. The purpose of this invention is as follows:

[0010] [1] In order to improve the robustness and accuracy of audio and video deepfake detection in real and complex environments, this invention proposes to introduce a visual quality perception mechanism. By dynamically evaluating the quality of each frame of the image, a spatial reliability mask is generated, and feature information of high-confidence regions is extracted in a focused manner, effectively suppressing interference from blurred, occluded, and compressed regions. This design enables the model to maintain stable detection performance in practical application scenarios such as low resolution and compression distortion.

[0011] [2] To enhance the model's ability to recognize highly synchronized and semantically natural forged audio and video content, this invention constructs a multi-scale cross-modal speech and image feature fusion strategy, taking into account both global synchronization relationships and physiological-level fine-grained coupling information between speech and lip movements. This mechanism helps to identify forged videos that appear highly realistic on the surface but differ at the micro-temporal or local action levels, thereby improving the accuracy and robustness of detection.

[0012] [3] To improve the model's adaptability when faced with visually degraded or semantically ambiguous samples, this invention introduces a self-supervised uncertainty calibration mechanism, enabling the model to output a reasonable confidence level when facing ambiguous, low-quality, or boundary samples, thereby effectively reducing false alarm rates and overconfident erroneous predictions. This mechanism enhances the model's practicality and deployability in open and uncontrollable environments.

[0013] Technical means

[0014] This invention addresses the challenge of intelligent inspection by proposing a method for detecting deepfake audio and video based on quality perception and multi-scale alignment, thus tackling the difficulties in complex environments. Details are as follows:

[0015] S1: Encode synchronized audio and video, use a lightweight quality assessment module to generate a spatial reliability mask, dynamically weight visual features, highlight clear and task-relevant areas, suppress interference caused by blurring, occlusion and compression distortion, and improve robustness and discriminative power under low-quality / real shooting conditions from the source.

[0016] S2: At the global scale, the overall synchronization relationship between speech and facial dynamics is characterized by bidirectional cross-attention. At the local scale, the phoneme-level speech representation is physiologically coupled and aligned with the facial action unit to form a coarse-to-fine consistent representation, which is used to identify fake samples that are highly synchronized on the surface but have deviations in the deep motion patterns.

[0017] S3: Joint visual quality and cross-modal consistency estimation generate adaptive temperature coefficients, perform temperature scaling on classification outputs and implement conservative inference; during the training phase, a self-supervised uncertainty calibration loss is introduced to guide the model to output higher uncertainty on low-quality / fuzzy / abnormal samples, reduce overconfident misjudgments, and improve cross-dataset generalization and practical deployment reliability.

[0018] In a preferred embodiment of the present invention, step S1 further includes the following steps:

[0019] S101: The input consists of multimodal data composed of synchronized video and audio streams. Video frames are processed by a visual encoder to extract their spatial feature maps, which reflect static appearance information such as facial structure, local texture, and motion trajectories. Simultaneously, the same video frame is further analyzed by a facial motion encoder to extract dynamic behavioral features of facial expression changes and muscle movements. On the other hand, the audio stream is processed by an audio encoder to extract its speech semantic features. These audio features are temporally aligned with the video frames, providing auxiliary information highly correlated with the spoken content. These three features are concatenated and fed as input into a lightweight convolutional neural network. This network is designed to learn the visual quality and semantic importance of various regions in the video image, ultimately outputting a two-dimensional spatial mask map. The mask value ranges from 0 to 1, measuring the reliability of information at each pixel location. High-quality regions receive higher mask values ​​and are thus given higher weight in subsequent processing; while blurred, occluded, or task-irrelevant regions are automatically weakened. This multi-source semantic-driven mask generation mechanism effectively overcomes the drawback of existing methods that "treat all regions the same," and significantly improves the robustness of the model in visual degradation scenarios.

[0020]

[0021] Among them, M t F represents the spatial quality mask for the t-th frame. t It is the image feature extracted by the visual encoder, FAU t It is a facial movement feature. It is an audio feature aligned with that frame. This represents a learnable quality evaluation convolutional network, where σ is the sigmoid activation function. Indicates feature concatenation operation;

[0022] S102: To further improve the reliability of spatial masks, this invention designs a mask adjustment strategy based on model prediction uncertainty. In actual videos, some regions may be semantically ambiguous, or the model may lack confidence in their discrimination during training. In this case, directly using the original mask can easily lead to the model over-reliance on unreliable regions when detecting forgeries. Therefore, this invention introduces an auxiliary uncertainty estimation branch network that shares input features with the quality mask generator. This branch outputs an uncertainty score map of the same size as the mask. The higher the uncertainty score, the greater the difficulty for the model to judge the region, and the higher the risk; therefore, its weight in the mask should be reduced accordingly. Finally, the original mask value is penalized and scaled according to the uncertainty score at that location to form a new quality-adjusted mask. This mechanism dynamically combines the model's own prediction confidence information, improving the model's tolerance to uncertain regions and its sensitivity to high-confidence regions. The calculation formula for this adjustment process is as follows:

[0023]

[0024] in, Let be the adjusted mask value at pixel position (i, j) in frame t. This is the initial mask value at that position. The value represents the uncertainty score (the higher the value, the less reliable the region is), and β is the uncertainty suppression factor, used to control the intensity of regulation.

[0025] S103: Considering the temporal fluctuations in image quality in real videos, such as instantaneous quality degradation in some frames due to camera shake, compression errors, or sudden changes in facial expressions, the mask generated for a single frame may be unstable and discontinuous. To improve the consistency and smoothness of the mask in the temporal dimension, this invention proposes a cross-frame semantic consistency-driven temporal enhancement mechanism. This mechanism takes the current frame as the center and selects multiple neighboring frames before and after it to form a sliding window. The model extracts the visual features of all these frames and calculates the semantic distance between the neighboring frames and the current frame using a semantic similarity function (such as cosine similarity). Based on this, the neighboring frames are weighted using a softmax function, giving higher weights to frames that are semantically closer to the current frame. Finally, the initial masks of all neighboring frames are fused according to this weight to obtain the temporally enhanced mask of the current frame. This method not only eliminates the impact of sudden quality degradation on the current mask but also enhances the continuity and semantic coherence of the mask in the temporal axis. The formula is shown below:

[0026]

[0027] in, M represents the time-enhanced mask for the t-th frame. l It is the initial mask of the l-th frame, Ft F l The visual feature representations of frames t and l are respectively, sim(F t F l ) represents the semantic similarity between two frames, w l is the softmax normalized fusion weight, and k represents the radius of the sliding window.

[0028] As a preferred embodiment of the present invention, S2 includes the following steps:

[0029] S201: At a macroscopic scale, this invention employs a bidirectional cross-attention mechanism for global alignment of audio and visual streams. Simultaneously, it introduces a gating term jointly formed by quality masking and uncertainty estimation to suppress interference from low-quality visual regions and high-uncertainty periods. Specifically, the global audio embedding is used as a query to retrieve global visual representations, and the global visual embedding is used as a query to retrieve global audio representations. When calculating the attention distribution, additional soft constraints on quality and uncertainty are injected, focusing attention more on clear, credible, and semantically relevant visual regions and their matching audio segments, thereby obtaining a semantically robust globally synchronized representation. This leads to the audio-guided global alignment output of the video. The formal expression is as follows:

[0030]

[0031] Among them, Q a,t K represents the global query vector extracted from the audio side with reference to frame t, used to retrieve the most relevant visual cues to the current speech on the visual side; v,t With V v,t Let M represent the key and value, respectively, composed of the quality-weighted visual features of the t-th frame, used to "point to" the audio query and aggregate the corresponding visual semantics; t It is the spatial reliability mask obtained from the aforementioned steps, vec(M) t Expanding it in order of visual features to align with attention scoring, log(·) is used as a logarithmic prior to award points to high-quality regions and deduct points for low-quality regions in attention, and λ controls the strength of this quality prior's influence on attention distribution; u t Represents the overall uncertainty score for frame t, 1u t As a uniform inhibition term across all visual positions to “flatten” the attention distribution under high uncertainty, γ controls the strength of this uncertainty gating. It is a standard scaling factor to stabilize the dot product attention; softmax(·) produces normalized weights for the visual position, which are ultimately summed with V. v,t Multiply to get Its meaning is "a global visual representation obtained by audio guidance under quality and uncertainty gating".

[0032] Global audio alignment output for video guidance The formal expression is as follows:

[0033]

[0034] Among them, Q v,t This represents a global query extracted from the visual side at frame t, used to find the audio segment that best matches the current visual scene; K a,t With V a,t These are the key and value on the audio side, carrying the acoustic and semantic information that is retrieved and aggregated by visual queries; Also used for attention scaling, softmax(·) outputs the weight distribution with respect to audio location and is compared with V. a,t Aggregation Its meaning is "global audio representation obtained through visual guidance". The two formulas work together to achieve bidirectional cross-modal global synchronization alignment: the former emphasizes "where the speech is looking", and the latter emphasizes "what the picture is hearing", complementing quality prior and uncertainty gating;

[0035] S202: At the microscale, this invention places phoneme-level speech representations and facial action units in the same alignment space. Through joint modeling of physiological coupling priors and learnable coupling matrices, it captures the temporal and detailed coordination patterns of lip-sync and speech. To this end, firstly, phoneme representations and facial action unit representations are projected into a shared space of the same dimension, and a prior matrix obtained through statistical physiological knowledge or data-driven methods is introduced to characterize the relationship that "certain phonemes are more likely to trigger certain facial actions" in a weakly supervised manner. Then, this prior is adaptively corrected using learnable parameters, thereby forming an alignment weight with "prior + learning" dual constraints in local attention, enhancing the ability to detect near-surface high-synchronization forgeries. The formal expression is as follows:

[0036]

[0037] Among them, P t For phoneme representation, U t Represented by facial motion units; The phonemes and facial action units are represented by MLP projection; W is the bilinear mapping parameter; Π is the physiological coupling prior matrix (which can be derived from statistical co-occurrence or expert knowledge base); Δ is a learnable prior correction term to adapt to individual differences and cross-domain variations; α > 0 controls the prior injection intensity; For local alignment of attention; This provides fine-grained aligned output of facial action unit information. This local module explicitly integrates the physiological mapping of "phoneme-action" into attention calculation, enabling the discovery of deep inconsistencies in seemingly well-synchronized forged samples. S203: To collaboratively fuse two types of evidence—global synchronization and local physiological coupling—this invention proposes a multi-scale consistency scoring and run-length adaptive weighting strategy. First, a global consistency score and a local consistency score are established for the t-th frame. Then, the fusion weights of the two are adaptively assigned based on the quality and uncertainty of that frame. Subsequently, the temporal alignment of the phoneme sequence and the facial action unit sequence is constrained in the temporal dimension using soft alignment regularization, thereby improving consistency robustness under conditions of speech rate variation, slight drift, and near-synchronous forgery. The formal expression is as follows:

[0038]

[0039]

[0040] in, The global consistency score is measured using the cosine similarity of the bidirectional aligned output.

[0041] Local consistency score (mean of maximum alignment strength from phoneme to facial action unit); u t The uncertainty is for the t-th frame; The frame average quality (e.g., pixel mean) is obtained from the quality mask Mt statistics; η > 0 is a temperature coefficient used to control the steepness of the adaptive weights; These represent the fusion weights for global and local evidence, which dynamically change with quality / uncertainty; s t Let be the overall consistency score for frame t; softDTW(·,·) is a differentiable dynamic temporal warping regularization used to encourage temporal alignment and smoothing between phoneme sequences and facial action unit sequences, improving robustness to speech rate changes and micro-drifts. Through the linkage of frame-level score aggregation and temporal regularization, this invention achieves triple consistency modeling of "strong global synchronization + deep local coupling + stable temporal alignment".

[0042] As a preferred embodiment of the present invention, step S3 includes the following steps:

[0043] S301: To achieve "conservative and robust" decision-making during the inference phase, this invention jointly models video-level visual quality, cross-modal consistency, and model prediction uncertainty to obtain a temperature coefficient dynamically correlated with sample credibility, and smooths the classification loss using temperature scaling. Specifically, video-level statistics are first obtained from frame-level uncertainty and spatial quality, and then the consistency score obtained from multi-scale alignment is combined to calculate the video-level temperature. Finally, a temperature-scaled softmax function is used to output the classification probability, thereby automatically reducing classification aggressiveness under conditions of "low quality, low consistency, and high uncertainty."

[0044]

[0045] In the above formula, z represents the unnormalized logits output by the video-level classification head. This represents the aggregate amount of frame-level uncertainty in the time dimension. This represents the average video quality obtained from the spatial reliability mask statistics. γ represents the aggregation of the consistency scores output by the multi-scale alignment module over time. u γ q γ s τ is an adjustment coefficient for sensitivity to uncertainty, quality, and consistency. (vid) The temperature coefficient is video-level and adapts to the sample confidence level. This represents the classification probability scaled to temperature.

[0046] S302: To further suppress overconfidence on untrusted samples, this invention introduces a "prior mixing" mechanism based on temperature scaling: when a sample shows higher uncertainty, worse quality, or lower consistency, the predicted probability is linearly mixed with a neutral prior distribution in a proportional manner, thereby achieving a controllable "retreat" in the probability space to reduce the risk of false alarms and improve generalization in open environments.

[0047]

[0048] In the above formula, σ(·) is the Sigmoid function, K is the prior mixing weight that increases with increasing uncertainty and decreasing quality and consistency, and α u α q α s To control for the influence of three types of factors on the mixing intensity, π is a neutral prior distribution (e.g., a uniform distribution with equal distributions across all categories). This is the final conservative probability output; when the sample confidence is low, increasing κ makes the output move closer to the prior, achieving "evidence-based conservatism".

[0049] S303: During the training phase, this invention constructs a self-supervised uncertainty target using quality and consistency as signals. By measuring the rule that "the worse the image quality / the weaker the alignment → the more uncertain the expression," it guides the model to learn to give more cautious confidence representations for blurry or anomalous samples. Simultaneously, it incorporates temperature-scaled and prior-mixed probabilities into the classification loss, thereby integrating the "conservative inference" strategy into end-to-end training, balancing accuracy and calibration.

[0050]

[0051]

[0052] In the above three formulas, u ★ For the self-supervised uncertainty target, it decreases and is limited to [0, 1] as the average quality and consistency increase; This is to aggregate and predict the uncertainty of the entire video using the model. Aligning forecast uncertainty with the target to achieve "doubt upon seeing a discrepancy"; The probabilities obtained after conservativeization in steps seven and eight are used to calculate the standard cross-entropy loss. β q ,β s , λ unc For the corresponding weights and balance coefficients, The total loss is used to optimize classification accuracy and uncertainty calibration in a coordinated manner under a unified objective.

[0053] Unlike existing technologies, the above methods achieve the following beneficial effects:

[0054] The visual quality perception mechanism significantly improves the model's reliability in distinguishing blurred and occluded regions, while the multi-scale cross-modal alignment strategy enhances the ability to model the micro-semantic relationships between speech and facial motion. The uncertainty perception mechanism effectively reduces the risk of overconfident predictions of boundary samples. In summary, this invention not only achieves comprehensive leadership in standard evaluation metrics but also possesses significant advantages in practicality, stability, and security, making it suitable for deployment in various real-world audio and video forgery detection scenarios. Attached Figure Description

[0055] Figure 1 This is the overall flowchart of the present invention. Detailed Implementation

[0056] To explain in detail the technical content, structural features, objectives, and effects of the technical solution, the following description is provided in conjunction with specific embodiments and accompanying drawings.

[0057] Please refer to the following: Figure 1As shown in the figure, this embodiment provides a method for detecting audio and video depth forgery based on quality perception and multi-scale alignment, including the following steps:

[0058] S1: Encode synchronized audio and video, use a lightweight quality assessment module to generate a spatial reliability mask, dynamically weight visual features, highlight clear and task-relevant areas, suppress interference caused by blurring, occlusion and compression distortion, and improve robustness and discriminative power under low-quality / real shooting conditions from the source.

[0059] S2: At the global scale, the overall synchronization relationship between speech and facial dynamics is characterized by bidirectional cross-attention. At the local scale, the phoneme-level speech representation is physiologically coupled and aligned with the facial action unit to form a coarse-to-fine consistent representation, which is used to identify fake samples that are highly synchronized on the surface but have deviations in the deep motion patterns.

[0060] S3: Joint visual quality and cross-modal consistency estimation generate adaptive temperature coefficients, perform temperature scaling on classification outputs and implement conservative inference; during the training phase, a self-supervised uncertainty calibration loss is introduced to guide the model to output higher uncertainty on low-quality / fuzzy / abnormal samples, reduce overconfident misjudgments, and improve cross-dataset generalization and practical deployment reliability.

[0061] In the above embodiments, S1 further includes the following steps:

[0062] S101: The input consists of multimodal data composed of synchronized video and audio streams. Video frames are processed by a visual encoder to extract their spatial feature maps, which reflect static appearance information such as facial structure, local texture, and motion trajectories. Simultaneously, the same video frame is further analyzed by a facial motion encoder to extract dynamic behavioral features of facial expression changes and muscle movements. On the other hand, the audio stream is processed by an audio encoder to extract its speech semantic features. These audio features are temporally aligned with the video frames, providing auxiliary information highly correlated with the spoken content. These three features are concatenated and fed as input into a lightweight convolutional neural network. This network is designed to learn the visual quality and semantic importance of various regions in the video image, ultimately outputting a two-dimensional spatial mask map. The mask value ranges from 0 to 1, measuring the reliability of information at each pixel location. High-quality regions receive higher mask values ​​and are thus given higher weight in subsequent processing; while blurred, occluded, or task-irrelevant regions are automatically weakened. This multi-source semantic-driven mask generation mechanism effectively overcomes the drawback of existing methods that "treat all regions the same," and significantly improves the robustness of the model in visual degradation scenarios.

[0063]

[0064] Among them, M t F represents the spatial quality mask for the t-th frame. t It is the image feature extracted by the visual encoder, FAU t It is a facial movement feature. It is an audio feature aligned with that frame. This represents a learnable quality evaluation convolutional network, where σ is the sigmoid activation function. Indicates feature concatenation operation;

[0065] S102: To further improve the reliability of spatial masks, this invention designs a mask adjustment strategy based on model prediction uncertainty. In actual videos, some regions may be semantically ambiguous, or the model may lack confidence in their discrimination during training. In this case, directly using the original mask can easily lead to the model over-reliance on unreliable regions when detecting forgeries. Therefore, this invention introduces an auxiliary uncertainty estimation branch network that shares input features with the quality mask generator. This branch outputs an uncertainty score map of the same size as the mask. The higher the uncertainty score, the greater the difficulty for the model to judge the region, and the higher the risk; therefore, its weight in the mask should be reduced accordingly. Finally, the original mask value is penalized and scaled according to the uncertainty score at that location to form a new quality-adjusted mask. This mechanism dynamically combines the model's own prediction confidence information, improving the model's tolerance to uncertain regions and its sensitivity to high-confidence regions. The calculation formula for this adjustment process is as follows:

[0066]

[0067] in, Let be the adjusted mask value at pixel position (i, j) in frame t. This is the initial mask value at that position. The value represents the uncertainty score (the higher the value, the less reliable the region is), and β is the uncertainty suppression factor, used to control the intensity of regulation.

[0068] S103: Considering the temporal fluctuations in image quality in real videos, such as instantaneous quality degradation in some frames due to camera shake, compression errors, or sudden changes in facial expressions, the mask generated for a single frame may be unstable and discontinuous. To improve the consistency and smoothness of the mask in the temporal dimension, this invention proposes a cross-frame semantic consistency-driven temporal enhancement mechanism. This mechanism takes the current frame as the center and selects multiple neighboring frames before and after it to form a sliding window. The model extracts the visual features of all these frames and calculates the semantic distance between the neighboring frames and the current frame using a semantic similarity function (such as cosine similarity). Based on this, the neighboring frames are weighted using a softmax function, giving higher weights to frames that are semantically closer to the current frame. Finally, the initial masks of all neighboring frames are fused according to this weight to obtain the temporally enhanced mask of the current frame. This method not only eliminates the impact of sudden quality degradation on the current mask but also enhances the continuity and semantic coherence of the mask in the temporal axis. The formula is shown below:

[0069]

[0070] in, M represents the time-enhanced mask for the t-th frame. l It is the initial mask of the l-th frame, F t F l The visual feature representations of frames t and l are respectively, sim(F t F l ) represents the semantic similarity between two frames, w l is the softmax normalized fusion weight, and k represents the radius of the sliding window.

[0071] In the above embodiments, S2 further includes the following steps:

[0072] S201: At a macroscopic scale, this invention employs a bidirectional cross-attention mechanism for global alignment of audio and visual streams. Simultaneously, it introduces a gating term jointly formed by quality masking and uncertainty estimation to suppress interference from low-quality visual regions and high-uncertainty periods. Specifically, the global audio embedding is used as a query to retrieve global visual representations, and the global visual embedding is used as a query to retrieve global audio representations. When calculating the attention distribution, additional soft constraints on quality and uncertainty are injected, focusing attention more on clear, credible, and semantically relevant visual regions and their matching audio segments, thereby obtaining a semantically robust globally synchronized representation. This leads to the audio-guided global alignment output of the video. The formal expression is as follows:

[0073]

[0074] Among them, Q a,tK represents the global query vector extracted from the audio side with reference to frame t, used to retrieve the most relevant visual cues to the current speech on the visual side; v,t With V v,t Let M represent the key and value, respectively, composed of the quality-weighted visual features of the t-th frame, used to "point to" the audio query and aggregate the corresponding visual semantics; t It is the spatial reliability mask obtained from the aforementioned steps, vec(M) t Expanding it in order of visual features to align with attention scoring, log(·) is used as a logarithmic prior to award points to high-quality regions and deduct points for low-quality regions in attention, and λ controls the strength of this quality prior's influence on attention distribution; u t Represents the overall uncertainty score for frame t, 1u t As a uniform inhibition term for all visual locations to “flatten” the attention distribution under high uncertainty, γ Controlling the strength of this uncertainty gating; It is a standard scaling factor to stabilize the dot product attention; softmax(·) produces normalized weights for the visual position, which are ultimately summed with V. v,t Multiply to get Its meaning is "a global visual representation obtained by audio guidance under quality and uncertainty gating".

[0075] Global audio alignment output for video guidance The formal expression is as follows:

[0076]

[0077] Among them, Q v,t This represents a global query extracted from the visual side at frame t, used to find the audio segment that best matches the current visual scene; K a,t With V a,t These are the key and value on the audio side, carrying the acoustic and semantic information that is retrieved and aggregated by visual queries; Also used for attention scaling, softmax(·) outputs the weight distribution with respect to audio location and is compared with V. a,t Aggregation Its meaning is "global audio representation obtained through visual guidance". The two formulas work together to achieve bidirectional cross-modal global synchronization alignment: the former emphasizes "where the speech is looking", and the latter emphasizes "what the picture is hearing", complementing quality prior and uncertainty gating;

[0078] S202: At the microscale, this invention places phoneme-level speech representations and facial action units in the same alignment space. Through joint modeling of physiological coupling priors and learnable coupling matrices, it captures the temporal and detailed coordination patterns of lip-sync and speech. To this end, firstly, phoneme representations and facial action unit representations are projected into a shared space of the same dimension, and a prior matrix obtained through statistical physiological knowledge or data-driven methods is introduced to characterize the relationship that "certain phonemes are more likely to trigger certain facial actions" in a weakly supervised manner. Then, this prior is adaptively corrected using learnable parameters, thereby forming an alignment weight with "prior + learning" dual constraints in local attention, enhancing the ability to detect near-surface high-synchronization forgeries. The formal expression is as follows:

[0079]

[0080] Among them, P t For phoneme representation, U t Represented by facial motion units; The phonemes and facial action units are represented by MLP projection; W is the bilinear mapping parameter; Π is the physiological coupling prior matrix (which can be derived from statistical co-occurrence or expert knowledge base); Δ is a learnable prior correction term to adapt to individual differences and cross-domain variations; α > 0 controls the prior injection intensity; For local alignment of attention; This provides fine-grained aligned output of facial action unit information. The local module explicitly integrates the physiological mapping of "phoneme-action" into attention calculation, enabling the discovery of deep inconsistencies in seemingly well-synchronized forged samples. S203: To collaboratively fuse two types of evidence—global synchronization and local physiological coupling—this invention proposes a multi-scale consistency scoring and run-length adaptive weighting strategy. First, a global consistency score and a local consistency score are established for the t-th frame. Then, the fusion weights are adaptively assigned based on the quality and uncertainty of that frame. Subsequently, the temporal alignment of the phoneme sequence and the facial action unit sequence is constrained in the temporal dimension using soft alignment regularization, thereby improving consistency robustness under conditions of speech rate variation, slight drift, and near-synchronous forgery. The formal expression is as follows:

[0081]

[0082]

[0083] in, The global consistency score is measured using the cosine similarity of the bidirectional aligned output.

[0084] Local consistency score (mean of maximum alignment strength from phoneme to facial action unit); u t The uncertainty is for the t-th frame; For the quality mask M t The statistically obtained average frame quality (e.g., average pixel value); η > 0 is a temperature coefficient used to control the steepness of the adaptive weights; These represent the fusion weights for global and local evidence, which dynamically change with quality / uncertainty; s t Let be the overall consistency score for frame t; softDTW(·,·) is a differentiable dynamic temporal warping regularization used to encourage temporal alignment and smoothing between phoneme sequences and facial action unit sequences, improving robustness to speech rate changes and micro-drifts. Through the linkage of frame-level score aggregation and temporal regularization, this invention achieves triple consistency modeling of "strong global synchronization + deep local coupling + stable temporal alignment".

[0085] In the above embodiments, S3 further includes the following steps:

[0086] S301: To achieve "conservative and robust" decision-making during the inference phase, this invention jointly models video-level visual quality, cross-modal consistency, and model prediction uncertainty to obtain a temperature coefficient dynamically correlated with sample credibility, and smooths the classification loss using temperature scaling. Specifically, video-level statistics are first obtained from frame-level uncertainty and spatial quality, then the consistency score obtained from multi-scale alignment is combined to calculate the video-level temperature, and finally, a temperature-scaled softmax function is used to output the classification probability, thereby automatically reducing classification aggressiveness under conditions of "low quality, low consistency, and high uncertainty." Details are as follows:

[0087]

[0088] In the above formula, z represents the unnormalized logits output by the video-level classification head. This represents the aggregate amount of frame-level uncertainty in the time dimension. This represents the average video quality obtained from the spatial reliability mask statistics. γ represents the aggregation of the consistency scores output by the multi-scale alignment module over time. u ,γ q γ s τ is an adjustment coefficient for sensitivity to uncertainty, quality, and consistency. (vid) The temperature coefficient is video-level and adapts to the sample confidence level. The classification probability scaled by temperature;

[0089] S302: To further suppress overconfidence on untrusted samples, this invention introduces a "prior mixing" mechanism based on temperature scaling: when a sample shows higher uncertainty, worse quality, or lower consistency, the predicted probability is linearly mixed with a neutral prior distribution in a proportional manner, thereby achieving a controllable "retreat" in the probability space to reduce the risk of false alarms and improve generalization in open environments.

[0090]

[0091] In the above formula, σ(·) is the Sigmoid function, K is the prior mixing weight that increases with increasing uncertainty and decreasing quality and consistency, and α u α q α s To control for the influence of three types of factors on the mixing intensity, π is a neutral prior distribution (e.g., a uniform distribution with equal distributions across all categories). The final conservative probability output is calculated; when the sample confidence is low, K is increased to make the output closer to the prior, achieving "evidence-based conservatism"; S303: During the training phase, this invention constructs a self-supervised uncertainty target using quality and consistency as signals. By measuring the rule that "the worse the image quality / the weaker the alignment, the more uncertain it should be," the model is guided to learn to give more cautious confidence expressions on blurry or abnormal samples. At the same time, the probability after temperature scaling and prior mixing is used in the classification loss, thereby incorporating the "conservative inference" strategy into end-to-end training, taking into account both accuracy and calibration.

[0092]

[0093] In the above three formulas, u ★ For the self-supervised uncertainty target, it decreases and is limited to [0, 1] as the average quality and consistency increase; This is to aggregate and predict the uncertainty of the entire video using the model. Aligning forecast uncertainty with the target to achieve "doubt upon seeing a discrepancy"; The probabilities obtained after conservativeization in steps seven and eight are used to calculate the standard cross-entropy loss. β q ,β s , λ unc For the corresponding weights and balance coefficients, The total loss is used to optimize classification accuracy and uncertainty calibration in a coordinated manner under a unified objective.

[0094] To verify the effectiveness and universality of the proposed deepfake detection method for audio and video, this invention conducted model-based experimental evaluations on two representative public datasets: FakeAVCeleb and DFDC. FakeAVCeleb is one of the current mainstream benchmarks for audio and video forgery detection, covering various pairings of real and forged videos, making it quite challenging. DFDC (Deepfake Detection Challenge), on the other hand, is a large-scale deepfake competition dataset hosted by Facebook, covering a wide range of forgery techniques and video quality, and is closer to real-world scenarios.

[0095] On the FakeAVCeleb dataset, this invention achieved a classification accuracy of 99.5% and an area under the curve (AUC) of 99.8%, significantly outperforming existing visual or audio-video fusion baseline methods. Compared to pure visual models such as Xception (67.9% / 70.5%) and LipForensics (80.1% / 82.4%), this invention improved accuracy by nearly 20% and AUC by 17%, respectively. Furthermore, it achieved a further improvement of over 1% over the current leading audio-video fusion method, AVFF, fully demonstrating its leading performance in feature fusion and forgery detection.

[0096] On the DFDC dataset, this invention also performs exceptionally well, achieving an accuracy of 95.5% and an AUC of 97.2%, surpassing the detection levels of existing methods such as AVoiD-DF (91.4% / 94.8%). This result demonstrates that the proposed quality-aware mechanism, multi-scale alignment structure, and uncertainty control strategy retain good robustness and generalization ability when facing different data distributions, different forgery strategies, and fluctuations in visual quality.

[0097] The experimental results are as follows:

[0098] Table 1 Comparison results between FakeAVCeleb and the DFDC dataset.

[0099]

Claims

1. A method for detecting audio and video depth forgery based on quality perception and multi-scale alignment, characterized in that: It includes the following steps: S1: Encode synchronized audio and video, use a lightweight quality assessment module to generate a spatial reliability mask, dynamically weight visual features, highlight clear and task-relevant areas, suppress interference caused by blurring, occlusion and compression distortion, and improve robustness and discriminative power under low-quality / real shooting conditions from the source. S2: At the global scale, the overall synchronization relationship between speech and facial dynamics is characterized by bidirectional cross-attention. At the local scale, the phoneme-level speech representation is physiologically coupled and aligned with the facial action unit to form a coarse-to-fine consistent representation, which is used to identify fake samples that are highly synchronized on the surface but have deviations in the deep motion patterns. S3: Joint visual quality and cross-modal consistency estimation generate adaptive temperature coefficients, apply temperature scaling to classification outputs, and implement conservative inference; During the training phase, a self-supervised uncertainty calibration loss is introduced to guide the model to output higher uncertainty on low-quality / fuzzy / abnormal samples, reduce overconfidence misjudgments, and improve cross-dataset generalization and reliability in actual deployment.

2. The audio / video depth forgery detection method based on quality perception and multi-scale alignment according to claim 1, characterized in that, S1 includes the following steps: S101: The input consists of multimodal data composed of synchronized video and audio streams. Video frames are processed by a visual encoder to extract their spatial feature maps, which reflect static appearance information such as facial structure, local texture, and motion trajectories. Simultaneously, the same video frame is further analyzed by a facial motion encoder to extract dynamic behavioral features of facial expression changes and muscle movements. On the other hand, the audio stream is processed by an audio encoder to extract its speech semantic features. These audio features are temporally aligned with the video frames, providing auxiliary information highly correlated with the spoken content. These three features are concatenated and fed as input into a lightweight convolutional neural network for processing. This network is designed to learn the visual quality and semantic importance of various regions in the video image, ultimately outputting a two-dimensional spatial mask map. The mask's value ranges from 0 to 1, measuring the reliability of information at each pixel location. High-quality regions will receive higher mask values, thus being assigned higher weights in subsequent processing. Blurred, occluded, or task-irrelevant regions will be automatically reduced. This multi-source semantic-driven mask generation mechanism effectively overcomes the drawback of existing methods that "treat all regions equally," significantly improving the model's robustness in visual degradation scenarios. Among them, M t F represents the spatial quality mask for the t-th frame. t It is the image feature extracted by the visual encoder, FAU t It is a facial movement feature. It is an audio feature aligned with that frame. This represents a learnable quality evaluation convolutional network, where σ is the sigmoid activation function. Indicates feature concatenation operation; S102: To further improve the reliability of spatial masks, this invention designs a mask adjustment strategy based on model prediction uncertainty. In actual videos, some regions may be semantically ambiguous, or the model may lack confidence in their discrimination during training. In this case, directly using the original mask can easily lead to the model over-reliance on unreliable regions when detecting forgeries. Therefore, this invention introduces an auxiliary uncertainty estimation branch network that shares input features with the quality mask generator. This branch outputs an uncertainty score map of the same size as the mask. The higher the uncertainty score, the greater the difficulty for the model to judge the region, and the higher the risk; therefore, its weight in the mask should be reduced accordingly. Finally, the original mask value is penalized and scaled according to the uncertainty score at that location to form a new quality-adjusted mask. This mechanism dynamically combines the model's own prediction confidence information, improving the model's tolerance to uncertain regions and its sensitivity to high-confidence regions. The calculation formula for this adjustment process is as follows: in, Let be the adjusted mask value at pixel position (i, j) in frame t. This is the initial mask value at that position. The value represents the uncertainty score (the higher the value, the less reliable the region is), and β is the uncertainty suppression factor, used to control the intensity of regulation. S103: Considering the temporal fluctuations in image quality in real videos, such as instantaneous quality degradation in some frames due to camera shake, compression errors, or sudden changes in facial expressions, the mask generated for a single frame may be unstable and discontinuous. To improve the consistency and smoothness of the mask in the temporal dimension, this invention proposes a cross-frame semantic consistency-driven temporal enhancement mechanism. This mechanism takes the current frame as the center and selects multiple neighboring frames before and after it to form a sliding window. The model extracts the visual features of all these frames and calculates the semantic distance between the neighboring frames and the current frame using a semantic similarity function (such as cosine similarity). Based on this, the neighboring frames are weighted using a softmax function, giving higher weights to frames that are semantically closer to the current frame. Finally, the initial masks of all neighboring frames are fused according to this weight to obtain the temporally enhanced mask of the current frame. This method not only eliminates the impact of sudden quality degradation on the current mask but also enhances the continuity and semantic coherence of the mask in the temporal axis. The formula is shown below: in, M represents the time-enhanced mask for the t-th frame. l It is the initial mask of the l-th frame, F t F l The visual feature representations of frames t and l are respectively, sim(F t F l ) represents the semantic similarity between two frames, w l is the softmax normalized fusion weight, and k represents the radius of the sliding window.

3. The audio / video depth forgery detection method based on quality perception and multi-scale alignment according to claim 2, characterized in that, S2 includes the following steps: S201: At a macroscopic scale, this invention employs a bidirectional cross-attention mechanism for global alignment of audio and visual streams. Simultaneously, it introduces a gating term jointly formed by quality masking and uncertainty estimation to suppress interference from low-quality visual regions and high-uncertainty periods. Specifically, the global audio embedding is used as a query to retrieve global visual representations, and the global visual embedding is used as a query to retrieve global audio representations. When calculating the attention distribution, additional soft constraints on quality and uncertainty are injected, focusing attention more on clear, credible, and semantically relevant visual regions and their matching audio segments, thereby obtaining a semantically robust globally synchronized representation. This leads to the audio-guided global alignment output of the video. The formal expression is as follows: Among them, Q a,t K represents the global query vector extracted from the audio side with reference to frame t, used to retrieve the most relevant visual cues to the current speech on the visual side; v,t With V v,t Let M represent the key and value, respectively, composed of the quality-weighted visual features of the t-th frame, used to be "pointed to" by the audio query and aggregated to extract the corresponding visual semantics; t It is the spatial reliability mask obtained from the aforementioned steps, vec(M) t Expanding it in order of visual features to align with attention scoring, log(·) is used as a logarithmic prior to award points to high-quality regions and deduct points for low-quality regions in attention, and λ controls the strength of this quality prior's influence on attention distribution; u t Represents the overall uncertainty score for frame t, 1u t As a uniform inhibition term across all visual positions to "flatten" the attention distribution under high uncertainty, γ controls the strength of this uncertainty gating. It is a standard scaling factor to stabilize the dot product attention; softmax(·) produces normalized weights for the visual position, which are ultimately summed with V. v,t Multiply to get Its meaning is "a global visual representation obtained under quality and uncertainty gating, guided by audio". Video-guided audio global alignment output. The formal expression is as follows: Among them, Q v,t This represents a global query extracted from the visual side at frame t, used to find the audio segment that best matches the current visual scene; K a,t With V a,t These are the key and value on the audio side, carrying the acoustic and semantic information that is retrieved and aggregated by visual queries; Also used for attention scaling, softmax(·) outputs the weight distribution with respect to audio location and is compared with V. a,t Aggregation Its meaning is "global audio representation obtained through visual guidance". The two methods together achieve bidirectional cross-modal global synchronization alignment: the former emphasizes "where the speech is looking", and the latter emphasizes "what the picture is hearing", complementing quality prior and uncertainty gating; S202: At the microscale, this invention places phoneme-level speech representations and facial action units in the same alignment space. Through joint modeling of physiological coupling priors and learnable coupling matrices, it captures the temporal and detailed coordination patterns of lip-sync and speech. To this end, firstly, phoneme representations and facial action unit representations are projected into a shared space of the same dimension, and a prior matrix obtained through statistical physiological knowledge or data-driven methods is introduced to characterize the relationship that "certain phonemes are more likely to trigger certain facial actions" in a weakly supervised manner. Then, this prior is adaptively corrected using learnable parameters, thereby forming an alignment weight with "prior + learning" dual constraints in local attention, enhancing the ability to detect near-surface high-synchronization forgeries. The formal expression is as follows: Among them, P t For phoneme representation, U t Represented by facial motion units; The phonemes and facial action units are represented by MLP projection; W is the bilinear mapping parameter; Π is the physiological coupling prior matrix (which can be derived from statistical co-occurrence or expert knowledge base); Δ is a learnable prior correction term to adapt to individual differences and cross-domain variations; α > 0 controls the prior injection intensity; For local alignment of attention; This module provides fine-grained aligned output of facial motion unit information. It explicitly incorporates the physiological mapping of "phoneme-action" into attention calculations, enabling the uncovering of deep inconsistencies even on seemingly well-synchronized fake samples. S203: To collaboratively fuse two types of evidence—global synchronization and local physiological coupling—this invention proposes a multi-scale consistency scoring and run-length adaptive weighting strategy. First, a global consistency score and a local consistency score are established for the t-th frame. Then, the fusion weights of these two scores are adaptively assigned based on the quality and uncertainty of that frame. Subsequently, the temporal alignment of the phoneme sequence and the facial action unit sequence is constrained in the temporal dimension using soft alignment regularization, thereby improving consistency robustness under conditions of speech rate variation, slight drift, and near-synchronous forgery. The formal expression is as follows: in, The global consistency score is measured using the cosine similarity of the bidirectional aligned output. Local consistency score (mean of maximum alignment strength from phoneme to facial action unit); U t The uncertainty is for the t-th frame; For the quality mask M t The statistically obtained average frame quality (e.g., average pixel value); η > 0 is a temperature coefficient used to control the steepness of the adaptive weights; These represent the fusion weights for global and local evidence, which dynamically change with quality / uncertainty; s t Let be the overall consistency score for frame t; softDTW(·,·) is a differentiable dynamic temporal warping regularization used to encourage temporal alignment and smoothing between phoneme sequences and facial action unit sequences, improving robustness to speech rate changes and micro-drifts. Through the linkage of frame-level score aggregation and temporal regularization, this invention achieves triple consistency modeling of "strong global synchronization + deep local coupling + stable temporal alignment".

4. The audio and video depth forgery detection method based on quality perception and multi-scale alignment according to claim 3, characterized in that, S3 includes the following steps: S301: To achieve "conservative and robust" decision-making during the inference phase, this invention jointly models video-level visual quality, cross-modal consistency, and model prediction uncertainty to obtain a temperature coefficient dynamically correlated with sample credibility, and smooths the classification loss using temperature scaling. Specifically, video-level statistics are first obtained from frame-level uncertainty and spatial quality, then the consistency score obtained from multi-scale alignment is combined to calculate the video-level temperature, and finally, a temperature-scaled softmax function is used to output the classification probability, thereby automatically reducing classification aggressiveness under conditions of "low quality, low consistency, and high uncertainty." Details are as follows: In the above formula, z represents the unnormalized logits output by the video-level classification head. This represents the aggregate amount of frame-level uncertainty in the time dimension. This represents the average video quality obtained from the spatial reliability mask statistics. γ represents the aggregation of the consistency scores output by the multi-scale alignment module over time. u γ q γ s τ is an adjustment coefficient for sensitivity to uncertainty, quality, and consistency. (vid) The temperature coefficient is video-level and adapts to the sample confidence level. The classification probability scaled by temperature; S302: To further suppress overconfidence on untrusted samples, this invention introduces a "prior mixing" mechanism based on temperature scaling: when a sample shows higher uncertainty, worse quality, or lower consistency, the predicted probability is linearly mixed with a neutral prior distribution in a proportional manner, thereby achieving a controllable "retreat" in the probability space to reduce the risk of false alarms and improve generalization in open environments. In the above formula, σ(·) is the Sigmoid function, K is the prior mixing weight that increases with increasing uncertainty and decreasing quality and consistency, and α u α q α s To control for the influence of three types of factors on the mixing intensity, π is a neutral prior distribution (e.g., a uniform distribution with equal distributions across all categories). This is the final conservative probability output; when the sample confidence is low, increasing K makes the output move closer to the prior, achieving "evidence-based conservatism"; S303: During the training phase, this invention constructs a self-supervised uncertainty target using quality and consistency as signals. By measuring the rule that "the worse the image quality / the weaker the alignment → the more uncertain the result," it guides the model to learn to give more cautious confidence representations on blurry or anomalous samples. At the same time, the probability after temperature scaling and prior mixing is used in the classification loss, thereby incorporating the "conservative inference" strategy into end-to-end training, taking into account both accuracy and calibration. In the above three formulas, u * For the self-supervised uncertainty target, it decreases and is limited to [0, 1] as the average quality and consistency increase; This is to aggregate and predict the uncertainty of the entire video using the model. Aligning forecast uncertainty with the target to achieve "doubt upon seeing a discrepancy"; The probabilities obtained after conservativeization in steps seven and eight are used to calculate the standard cross-entropy loss. β q ,β s , λ unc For the corresponding weights and balance coefficients, The total loss is used to optimize classification accuracy and uncertainty calibration in a coordinated manner under a unified objective.

Citation Information

Cited By

  • Hierarchical alignment method and system for multi-modal perception data set of soft manipulator

    CN121973254A

  • Hierarchical alignment method and system for soft robotic hand multi-modal perception dataset

    CN121973254B

  • AI synthetic video detection method and system based on multi-agent collaboration

    CN122313369A

  • An AI synthetic video detection method and system based on multi-agent collaboration

    CN122313369B