Multimodal speech synthesis detection method
By introducing video modal signals and cross-modal anchor maps, the problem of difficulty in identifying forged segments caused by multimodal temporal misalignment and video action masking in existing technologies is solved, achieving higher accuracy and robustness in voice forgery detection.
Patent Information
- Application Number
- CN202610034370.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2046-01-12
AI Technical Summary
Existing voice forgery detection technologies struggle to achieve stable and reliable forgery segment recognition under complex conditions such as multimodal temporal misalignment, local forgery being masked by video actions, and mutual reinforcement of forgery signals from different modalities. In particular, they cannot effectively utilize video modal information for joint detection.
By introducing video modal signals, a cross-modal anchor graph is constructed. Using graph representation learning based on anchor time drift, the fine-grained dependencies between audio and video events are captured. Combined with a cross-modal speech synthesis detection model, accurate inference of the forgery probability of audio events and the global description vector is achieved.
Robust synthetic speech detection was achieved under non-strict temporal synchronization conditions, improving the detection accuracy and generalization ability of AI synthetic speech.
Smart Images

Figure CN121483221B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis detection technology, and in particular to a multimodal speech synthesis detection method. Background Technology
[0002] As speech synthesis technology, fake audio generation technology, and cross-modal deepfake technology continue to mature with the advancement of algorithms, computing power, and data scale, current end-to-end speech cloning and high-fidelity speech synthesis models can generate highly realistic speech content based on a small number of samples in a very short time. Their timbre, speech rate, emotion, and prosody are extremely close to real speech, thus significantly weakening the effectiveness of traditional forgery detection methods that rely on acoustic statistical features, prosodic patterns, short-time spectral differences, or manually constructed forgery indicators.
[0003] In real-world communication scenarios, such deepfake audio often appears synchronously with corresponding video footage, forming a combined audio-visual forgery. However, differences in sampling rates, cross-frame processing, and encoding methods between the audio and video modalities lead to widespread non-strict synchronization, manifesting as local time drift, unstable information alignment, or mismatch of key events. These factors make it difficult for detection techniques relying solely on a single acoustic modality to achieve stable and reliable forged segment recognition under complex conditions such as multimodal temporal misalignment, local forgery being masked by video actions, and mutual reinforcement of forged signals from different modalities. This results in a significant decrease in detection accuracy and generalization ability.
[0004] In existing research, CN115171725B proposes a self-supervised method to prevent speech synthesis attacks. This method constructs speech data training samples, performs self-supervised training using a pre-trained model, expands the anti-attack dataset with various synthetic speech algorithms, and integrates a gated recurrent unit (GRU) and attention module into a synthetic detection model to improve the efficiency and generalization ability of unlabeled data utilization. This approach jointly optimizes the pre-trained model and the synthetic detection model through two-stage training, improving detection accuracy across various synthetic speech algorithm scenarios. However, this approach primarily focuses on acoustic modality for detection training and does not jointly model the speech generation process with physical synchronization features such as lip movements and facial behaviors in the video modality, thus failing to leverage multimodal information for joint detection. Summary of the Invention
[0005] This invention aims to at least partially solve one of the technical problems in the aforementioned technologies. To this end, this invention proposes a multimodal speech synthesis detection method. Existing speech forgery detection relies solely on the audio modality, failing to identify subtle anomalies caused by cross-frame splicing, fine-grained forgery, or deep forgery generation. This application simultaneously introduces video modal signals to achieve cross-modal consistency comparison, improving the ability to identify complex forgeries. Existing methods cannot represent the temporal relationships, spatial semantics, and motion changes of audio and video events using graph structures. This application constructs a cross-modal anchor graph, using audio and video events as anchors and uncertainty as edge relationships, forming a learnable cross-modal structured representation, thereby capturing the fine-grained dependencies between the two modalities. Introducing graph representation learning based on anchor point temporal drift allows learning the true alignment between different events, effectively identifying offset patterns caused by forgery. Through a cross-modal speech synthesis detection model, the forgery probability of audio events and the global description vector are jointly input into the final discriminant layer, achieving accurate inference of the forgery synthesis probability of the entire speech segment. This can locate specific suspicious audio events and improve the generalization detection effect under different scenarios and audio quality conditions.
[0006] To achieve the above objectives, this invention proposes a multimodal speech synthesis detection method, which includes the following steps: acquiring a speech audio signal and its corresponding associated video modal signal; extracting key audio detection features of the speech audio signal at different signal points, and extracting key video detection features of the video modal signal at different signal points, wherein the key audio detection features include short-time energy and fundamental frequency, and the key video detection features include lip opening and closing event detection results and lip width; dividing the speech audio signal into multiple audio events based on the key audio detection features, and dividing the video modal signal into multiple video events based on the key video detection features; calculating the uncertainty between the audio events and the video events, and constructing a cross-modal anchor graph by using the audio events and the video events as anchor points and the uncertainty between the anchor points as edge relationships; performing graph representation learning based on anchor point time drift on the cross-modal anchor graph to obtain the alignment relationship between the anchor points; inputting the alignment relationship between the anchor points into a pre-trained cross-modal speech synthesis detection model to obtain the forgery probability of the audio events in the anchor points and the forgery synthesis probability of the speech audio signal.
[0007] The multimodal speech synthesis detection method proposed in this embodiment of the invention has the advantage that, based on multimodal detection data, it can achieve robust synthesized speech detection under non-strict temporal synchronization conditions, thereby improving the detection accuracy of AI synthesized speech.
[0008] In addition, the multimodal speech synthesis detection method proposed in the above embodiments of the present invention may also have the following additional technical features:
[0009] Optionally, extracting key video detection features of the video modal signal at different signal points includes: using a facial keypoint detector to perform facial keypoint detection on the signal values of the video modal signal at different signal points to obtain a set of contour points for the upper and lower lips; calculating the lip opening and closing height based on the absolute value of the difference between the midpoint coordinates of the upper lip contour points and the midpoint coordinates of the lower lip contour points; generating lip opening and closing event detection results based on the change values of adjacent lip opening and closing heights; and generating the lip width based on the coordinate difference between the horizontal coordinates of the left contour point and the right contour point of the lip.
[0010] Optionally, the speech audio signal is divided into multiple audio events based on the key audio detection features, including:
[0011] Based on the key audio detection features, audio event segmentation determination results of the speech audio signal at different signal points are generated, wherein the determination formula for the audio event segmentation determination results is as follows:
[0012]
[0013] in, Represents speech audio signals The audio event segmentation determination result at the nth signal point. This represents a conditional statement. If any expression in the conditional statement is true, the output of the conditional statement is True; otherwise, the output of the conditional statement is False. Represents speech audio signals The short-time energy at the nth signal point, Represents speech audio signals The short-time energy mean of N signal points Represents speech audio signals The short-time energy standard deviation at N signal points Indicates audio event control parameters, Represents speech audio signals The fundamental frequency at the nth signal point Represents speech audio signals The fundamental frequency at the (n-1)th signal point Represents speech audio signals The fundamental frequency change at the nth signal point This represents the preset fundamental frequency variation threshold. This represents the logical AND operator; N represents the audio signal. The signal length;
[0014] When both sides of the logical AND operator output True, then An output of True indicates a speech audio signal. The nth signal point is used as the audio event segmentation point; if If the output is not True, it indicates a speech audio signal. The nth signal point is not used as an audio event segmentation point;
[0015] Based on the audio event segmentation points in the speech audio signal, the speech audio signal is divided into multiple non-overlapping audio signal segments, and each audio signal segment is treated as an independent audio event, and the confidence level of the audio event is calculated.
[0016] Optionally, the video modal signal is divided into multiple video events based on the key video detection features, including:
[0017] Based on the key video detection features, video event segmentation determination results are generated for video modal signals at different signal points. The determination formula for the video event segmentation results is as follows:
[0018]
[0019] in, Represents video modal signal The video event segmentation determination result at the m-th signal point. Represents video modal signal The lip opening and closing event detection result at the m-th signal point A value of 1 indicates that a lip opening / closing event was detected. A value of 0 indicates that no lip opening / closing event was detected. Represents video modal signal The lip width at the m-th signal point, Represents video modal signal The lip width at the (m-1)th signal point, This represents the preset threshold for lip width variation, and M represents the video modal signal. The signal length;
[0020] when as well as If all outputs are True, then An output of True indicates a video modal signal. The m-th signal point is used as the video event segmentation point; if If the output is not True, it indicates a video modal signal. The m-th signal point is not used as a video event segmentation point;
[0021] Based on the video event segmentation points in the video modal signal, the video modal signal is divided into multiple non-overlapping video signal segments, and each video signal segment is treated as an independent video event. The confidence level of the video event is then calculated.
[0022] Optionally, calculating the uncertainty between the audio event and the video event includes: obtaining the confidence levels corresponding to the audio event and the video event; converting the audio event and the video event into spectral sequences respectively, and calculating the similarity between the spectral sequences corresponding to the audio event and the video event as the multimodal event similarity between the audio event and the video event; converting the multimodal event similarity between the audio event and the video event into uncertainty based on the confidence levels corresponding to the audio event and the video event, wherein the lower the uncertainty, the higher the degree of event matching between the audio event and the video event, and the audio event and the video event describe similar events.
[0023] Optionally, the video events are converted into spectral sequences, including: extracting the lip opening height and lip width at each signal point in the video event and concatenating them as feature vectors of the video event; mapping the feature vectors into S pseudo-Mel energies using a nonlinear mapping method; and concatenating the S pseudo-Mel energies as the spectral sequence of the video event.
[0024] Optionally, the structure of the cross-modal anchor point diagram is as follows:
[0025]
[0026]
[0027]
[0028] in, Represents a cross-modal anchor point diagram. Represents cross-modal anchor point diagram The set of anchor points in the middle, These are, in order, audio event-anchor set and video event-anchor set. Represents audio event-anchor set The h-th anchor point in the sequence, where H represents the number of audio events. Represents a video event-anchor set The first in There are several anchor points, where G represents the number of video events. Represents cross-modal anchor point diagram The set of edge relations in the middle, Indicates anchor point The edge relationships between them.
[0029] Optionally, graph representation learning based on anchor time drift is performed on the cross-modal anchor graph to obtain the alignment relationship between anchors, including: obtaining the event timestamps corresponding to the anchors, wherein the event timestamps are the median timestamps of the continuous signal points corresponding to the anchors; calculating the anchor time offsets between different anchors based on the event timestamps; and generating the alignment relationship between anchors based on the anchor time offsets between different anchors and the edge relationships, using an attention mechanism.
[0030] Optionally, the cross-modal speech synthesis detection model includes a cross-modal alignment encoder, an audio event authentication network, a global consistency aggregation network, and a speech synthesis discriminant network. The cross-modal alignment encoder includes a normalization layer and an encoding layer, the audio event authentication network includes a fully connected layer and a discriminant layer, the global consistency aggregation network includes a weighting module and a global consistency transformation module, and the speech synthesis discriminant network includes a splicing layer and a global discriminant layer.
[0031] Optionally, the alignment relationship between the anchor points is input into a pre-trained cross-modal speech synthesis detection model to obtain the forgery probability of audio events and the forgery synthesis probability of speech audio signals in the anchor points. This includes: receiving the alignment relationship between anchor points using the cross-modal speech synthesis detection model; the normalization layer in the cross-modal alignment encoder converts the alignment relationship between anchor points into an H-row, G-column alignment relationship matrix, where each row in the alignment relationship matrix represents the alignment relationship between the anchor point as an audio event and G anchor points as video events; the encoding layer sequentially performs convolutional encoding on each row in the alignment relationship matrix to obtain the cross-modal encoded features of each anchor point as an audio event; the audio event... The fully connected layer in the anti-spoofing network extracts hidden states from cross-modal coding features and uses a discriminant layer based on a two-layer activation function to convert the hidden states into forgery probabilities, which are then used as the forgery probabilities of the audio events associated with the hidden states. The weighting module in the global consistency aggregation network uses the forgery probabilities as weights to weight the cross-modal coding features corresponding to the audio events. The global consistency transformation module projects the weighted results to obtain a global description vector. The concatenation layer in the speech synthesis discriminant network concatenates the global description vector and the forgery probabilities corresponding to each audio event. The final discriminant layer based on a fully connected layer structure discriminates the concatenation result and outputs the forgery synthesis probability of the speech audio signal. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating a multimodal speech synthesis detection method according to an embodiment of the present invention.
[0033] Figure 2 This is a schematic diagram of a cross-modal anchor point diagram according to an embodiment of the present invention. Detailed Implementation
[0034] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0035] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the invention to those skilled in the art.
[0036] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0037] refer to Figure 1 As shown, the speech synthesis detection method based on multimodal processing in this embodiment of the invention includes the following steps:
[0038] S101, acquire the voice audio signal and its corresponding associated video modal signal.
[0039] In other words, the audio signal and its associated video modal signal are acquired. The two signals may be acquired from different acquisition links or independent devices, so there is no strict sampling clock synchronization relationship. The video modal signal can be directly provided by the audio signal provider. The video modal signal is in MP4 format, and the signal value of the video modal signal at different signal points is in the form of video frames.
[0040] Specifically, the signal points in the speech audio signal and video modal signal are the discrete coding indices of the signal. The speech audio signal is represented as follows:
[0041]
[0042]
[0043] in, Represents voice audio signals. Represents video modal signals, Represents speech audio signals The signal values at N signal points, including the speech audio signal. The timestamps corresponding to the N signal points are as follows: N represents the voice audio signal. The signal length, Represents video modal signal The signal values at M signal points, where the video modal signal... The timestamps corresponding to the M signal points are as follows: M represents the video modal signal. The signal length, Represents speech audio signals The signal value at the nth signal point, Represents video modal signal The signal value at the m-th signal point.
[0044] S102, extract key audio detection features of speech audio signals at different signal points, and extract key video detection features of video modal signals at different signal points. The key audio detection features include short-time energy and fundamental frequency, and the key video detection features include lip opening and closing event detection results and lip width.
[0045] Specifically, the formula for extracting key audio detection features of speech audio signals at different signal points is as follows:
[0046]
[0047]
[0048]
[0049] in, Represents speech audio signals Key audio detection features at the nth signal point Key audio detection features, in order. The short-time energy and fundamental frequency in the middle, Represents speech audio signals The signal value at the k-th signal point, where L represents the signal time window length, is set to 4. Represents the window function. express The corresponding window function value, Indicates the fundamental frequency period Corresponding signal value sequence The autocorrelation function, Indicates the range of the fundamental frequency period. This indicates the minimum fundamental frequency period (can be set to 2). This indicates the maximum fundamental frequency period (can be set to 10). Indicates in Extracting from makes To reach the maximum As a voice audio signal The fundamental frequency at the nth signal point.
[0050] Optionally, the window function is a Hamming window function. , This represents the cosine function.
[0051] As an example, extracting key video detection features of video modal signals at different signal points includes: using a facial keypoint detector to detect facial keypoints on the signal values of the video modal signals at different signal points to obtain a set of contour points for the upper and lower lips; calculating the lip opening and closing height based on the absolute value of the difference between the midpoint coordinates of the upper lip contour points and the midpoint coordinates of the lower lip contour points; generating lip opening and closing event detection results based on the change values of adjacent lip opening and closing heights; and generating the lip width based on the coordinate difference between the horizontal coordinates of the left and right contour points of the lips.
[0052] Optionally, the facial landmark detector uses the MediaPipe Face Mesh model, which outputs 468 fine-grained facial landmarks, making it easy to directly obtain the lip contour and jaw position.
[0053] Specifically, video modal signals The detection result of the lip opening and closing event at the m-th signal point is as follows:
[0054]
[0055] in, Represents video modal signal The lip opening and closing event detection result at the m-th signal point Represents video modal signal The lip opening height at the m-th signal point, This indicates the threshold for detecting lip opening and closing height. A value of 1 indicates that a lip opening / closing event was detected. A value of 0 indicates that no lip opening / closing event was detected.
[0056] Optionally, set for The maximum absolute value of the difference between the midpoint coordinates of the upper lip contour points and the midpoint coordinates of the lower lip contour points from the M collected signal points.
[0057] Specifically, video modal signals The formula for calculating the lip width at the m-th signal point is:
[0058]
[0059] in, Represents video modal signal The lip width at the m-th signal point, Represents video modal signal The horizontal coordinate of the left contour point of the lip at the m-th signal point. Represents video modal signal The horizontal coordinates of the right contour point of the lip at the m-th signal point are... This indicates the calculation of absolute value.
[0060] It should be noted that this application, based on the short-time energy and fundamental frequency extraction formula, can stably capture speech energy mutations and periodic acoustic structures at microscale time points. It can accurately identify the presence, start time, and intensity changes of speech in speech audio signals, solving the problem that simply relying on endpoint detection cannot stably handle weak sounds, breathy sounds, etc. At the same time, this application uses a facial key point detection model to extract the upper and lower lip contours and construct lip opening and closing height and lip width indices, obtaining key features that can directly reflect the dynamic behavior of mouth shape. This allows the video modality to still provide clear pronunciation action information in silent, partially occluded, or noise-interference scenarios. Furthermore, the lip opening and closing event detection results constructed based on the changes in lip opening and closing height enable the video modality to have discrete event feature expression capabilities, which can be mutually verified with the short-time energy peak of the audio modality, thereby improving the stability and robustness of cross-modal event alignment.
[0061] S103, the speech audio signal is divided into multiple audio events based on key audio detection features, and the video modal signal is divided into multiple video events based on key video detection features.
[0062] As one embodiment, the speech audio signal is divided into multiple audio events based on key audio detection features, including: generating audio event segmentation determination results for the speech audio signal at different signal points based on the key audio detection features, wherein the determination formula for the audio event segmentation determination results is:
[0063]
[0064] in, Represents speech audio signals The audio event segmentation determination result at the nth signal point. This represents a conditional statement. If any expression in the conditional statement is true, the output of the conditional statement is True; otherwise, the output of the conditional statement is False. Represents speech audio signals The short-time energy at the nth signal point, Represents speech audio signals The short-time energy mean of N signal points Represents speech audio signals The short-time energy standard deviation at N signal points Indicates audio event control parameters, Represents speech audio signals The fundamental frequency at the nth signal point Represents speech audio signals The fundamental frequency at the (n-1)th signal point Represents speech audio signals The fundamental frequency change at the nth signal point This represents the preset fundamental frequency variation threshold. This represents the logical AND operator; N represents the audio signal. The signal length;
[0065] When both sides of the logical AND operator output True, then An output of True indicates a speech audio signal. The nth signal point is used as the audio event segmentation point; if If the output is not True, it indicates a speech audio signal. The nth signal point is not used as an audio event segmentation point;
[0066] Based on the audio event segmentation points in the speech audio signal, the speech audio signal is divided into multiple non-overlapping audio signal segments, and each audio signal segment is treated as an independent audio event, and the confidence level of the audio event is calculated.
[0067] Specifically, if the audio event segmentation points are respectively the first... and the There are 10 signal points, among which... , The audio events are then divided as follows: The key audio detection features of the audio event at each signal point are obtained, and the confidence level of the audio event is calculated. The formula for calculating the confidence level is:
[0068]
[0069]
[0070]
[0071] in, Indicates audio events Confidence level, Indicates audio events Middle start end The confidence coefficient, Indicates audio events middle end The confidence coefficient, Indicates audio events The short-time energy mean at all signal points. Indicates audio events The fundamental frequency mean at all signal points in the signal. Indicates audio events The short-time energy standard deviation at all signal points. All represent audio confidence control coefficients, set They are 0.6 and 0.4 respectively. This represents the activation function, which is the Sigmoid function, used to control the confidence level between 0 and 1.
[0072] As one embodiment, dividing the video modal signal into multiple video events based on key video detection features includes: generating video event segmentation determination results for the video modal signal at different signal points based on the key video detection features, wherein the determination formula for the video event segmentation determination results is:
[0073]
[0074] in, Represents video modal signal The video event segmentation determination result at the m-th signal point. Represents video modal signal The lip opening and closing event detection result at the m-th signal point A value of 1 indicates that a lip opening / closing event was detected. A value of 0 indicates that no lip opening / closing event was detected. Represents video modal signal The lip width at the m-th signal point, Represents video modal signal The lip width at the (m-1)th signal point, This represents the preset threshold for lip width variation, and M represents the video modal signal. The signal length;
[0075] when as well as If all outputs are True, then An output of True indicates a video modal signal. The m-th signal point is used as the video event segmentation point; if If the output is not True, it indicates a video modal signal. The m-th signal point is not used as a video event segmentation point;
[0076] Based on the video event segmentation points in the video modal signal, the video modal signal is divided into multiple non-overlapping video signal segments, and each video signal segment is treated as an independent video event. The confidence level of the video event is then calculated.
[0077] Specifically, if the video event segmentation points are respectively the first... and the There are 1 signal points, among which , The video events are then divided as follows:
[0078] The key video detection features of the video event at each signal point are obtained, and the confidence level of the video event is calculated. The formula for calculating the confidence level is:
[0079]
[0080]
[0081]
[0082] in, Indicates video event Confidence level, Indicates audio events Middle start end The confidence coefficient, Indicates audio events middle end The confidence coefficient, Indicates video event The mean of lip opening and closing event detection results at all signal points. Indicates video event The mean lip width at all signal points. Indicates video event The standard deviation of lip opening and closing event detection results at all signal points. All represent video confidence control coefficients, set The values are 0.3 and 0.7 respectively. This represents the activation function.
[0083] S104, calculate the uncertainty between audio events and video events, and construct a cross-modal anchor graph by using audio events and video events as anchor points and the uncertainty between anchor points as edge relationships.
[0084] As one embodiment, calculating the uncertainty between audio events and video events includes: obtaining the confidence levels corresponding to the audio events and video events; converting the audio events and video events into spectral sequences respectively, and calculating the similarity between the spectral sequences corresponding to the audio events and video events as the multimodal event similarity between the audio events and video events; converting the multimodal event similarity between the audio events and video events into uncertainty based on the confidence levels corresponding to the audio events and video events, wherein the lower the uncertainty, the higher the degree of event matching between the audio events and video events, and the audio events and video events describe similar events.
[0085] As a specific embodiment, audio events The process of converting to a spectral sequence is as follows:
[0086] For audio events Perform a short-time Fourier transform and calculate the power spectrum, where the audio events... The corresponding power spectrum is:
[0087]
[0088] in, Indicates audio events In the power spectrum at frequency index k, j represents the imaginary unit. , Represents an exponential function with the natural constant as its base;
[0089] Map the power spectrum to S Mel bands as a spectral sequence:
[0090]
[0091]
[0092] in, Indicates audio events The corresponding spectral sequence, This represents the Mel energy mapped to the S-th Mel band (S can be set to 10) of the power spectrum. This represents the S-th Mel filter. Indicates control parameters, settings It is 0.1. This indicates taking the logarithm.
[0093] As a specific embodiment, the video events are converted into spectral sequences, including: extracting the lip opening and closing height and lip width at each signal point in the video event and splicing them together as the feature vector of the video event; mapping the feature vector into S pseudo-Mel energies using a nonlinear mapping method; and splicing the S pseudo-Mel energies together as the spectral sequence of the video event.
[0094] Specifically, the mapping formula for the pseudo-Mel energy is:
[0095]
[0096] in, This represents the S-th pseudo-Mel energy obtained by mapping the eigenvector F. This represents the mapping weight parameter of the S-th nonlinear mapping function. Let represent the mapping constant of the S-th nonlinear mapping function.
[0097] Optionally, a cosine similarity algorithm is used to calculate the similarity between the spectral sequences corresponding to audio events and video events. The spectral similarity measure based on cosine similarity can robustly reflect the morphological consistency between events, weaken the influence of event length differences and local shifts, and obtain stable cross-modal event similarity.
[0098] Based on the confidence scores corresponding to audio and video events, the multimodal event similarity between audio and video events is converted into uncertainty. The lower the uncertainty, the higher the degree of event matching between audio and video events, indicating that the audio and video events describe similar events.
[0099] It should be noted that this application converts audio events into spectral sequences based on Mel energy, which can effectively compress high-dimensional power spectra and enhance the distinguishability of speech in energy distribution, enabling event-level speech features to have stable frequency domain representation; at the same time, it constructs a pseudo-sound spectrum domain by using lip opening and closing height and lip width features, and uses nonlinear mapping to generate pseudo-Mel energy, making the video modality comparable to the audio modality in frequency band structure, and realizing the basis for cross-modal spectrum alignment.
[0100] Specifically, the uncertainty transformation formula between audio events and video events is as follows:
[0101]
[0102] in, This indicates that the h-th audio event is related to the h-th audio event. Uncertainty between video events, This indicates that the h-th audio event is related to the h-th audio event. Multimodal event similarity between video events This represents the confidence level of the h-th audio event. Indicates the first Confidence level of a video event All represent uncertainty weighting coefficients, set The values are 0.6 and 0.4 respectively.
[0103] It should be noted that this application introduces the confidence level of events and converts similarity into uncertainty through logarithmic weighting, thereby realizing a joint evaluation mechanism based on event reliability and cross-modal consistency. This suppresses erroneous matching caused by low-confidence events and enhances the matching strength between real speech events and real lip-sync events, accurately reflecting the speech-lip-sync relationship between the two.
[0104] As an example, the structure of the cross-modal anchor point diagram is as follows:
[0105]
[0106]
[0107]
[0108] in, Represents a cross-modal anchor point diagram. Represents cross-modal anchor point diagram The set of anchor points in the middle, These are, in order, audio event-anchor set and video event-anchor set. Represents audio event-anchor set The h-th anchor point in the sequence, where H represents the number of audio events. Represents a video event-anchor set The first in There are several anchor points, where G represents the number of video events. Represents cross-modal anchor point diagram The set of edge relations in the middle, Indicates anchor point The edge relationships between them.
[0109] S105, perform graph representation learning based on anchor time drift on the cross-modal anchor graph to obtain the alignment relationship between anchors.
[0110] As an example, graph representation learning based on anchor time drift is performed on the cross-modal anchor graph to obtain the alignment relationship between anchors, including: obtaining the event timestamps corresponding to the anchors, wherein the event timestamps are the median timestamps of the continuous signal points corresponding to the anchors; calculating the anchor time offset between different anchors based on the event timestamps; and generating the alignment relationship between anchors based on the anchor time offsets between different anchors and the edge relationships, using an attention mechanism.
[0111] Specifically, anchor point The time offset between the anchor points is:
[0112] R
[0113] in, Indicates anchor point Anchor point time offset between Indicates the prior drift standard deviation. This represents the prior drift mean. anchor point The corresponding event timestamp, Indicates anchor point The corresponding event timestamps are obtained by simultaneously acquiring multiple sets of audio and video modal signals, and calculating the mean and standard deviation of the timestamp deviations between the two types of signals, which are used as the prior drift mean and prior drift standard deviation, respectively. This represents an exponential function with the natural constant as its base.
[0114] Anchor point The formula for generating the alignment relationship between them is:
[0115]
[0116] in, Indicates anchor point Alignment relationship between them.
[0117] It should be noted that by defining the event timestamp of each anchor point as the median of the timestamps of consecutive signal points, this invention can effectively eliminate local noise, boundary ambiguity, and detection jitter within the event, thereby making the event timing position stable. Furthermore, based on the advance quantization of the timestamp deviations of multiple sets of speech audio signals and video modal signals, the prior drift mean and prior drift standard deviation are obtained, which can form a statistical prior that reflects the real mouth synchronization offset law. By using a Gaussian anchor point time offset function, the natural physiological delay between acoustic generation and facial muscle movement is reflected, which can maintain matching robustness in the presence of natural mouth shape lag, pronunciation advance, frame rate differences, etc.
[0118] Simultaneously introduce attention mechanism to construct adaptive weights It can enhance the coupling strength between real speech events and real lip-sync events, suppress artifacts such as speech-mouth asynchrony, lip-sync lag, and missing audio events that are common in AI-synthesized speech, and improve the recognition accuracy.
[0119] S106, input the alignment relationship between anchor points into the pre-trained cross-modal speech synthesis detection model to obtain the forgery probability of audio events in the anchor points and the forgery synthesis probability of speech audio signals.
[0120] As an example, the cross-modal speech synthesis detection model includes a cross-modal alignment encoder, an audio event authentication network, a global consistency aggregation network, and a speech synthesis discriminant network. The cross-modal alignment encoder includes a normalization layer and an encoding layer, the audio event authentication network includes a fully connected layer and a discriminant layer, the global consistency aggregation network includes a weighting module and a global consistency transformation module, and the speech synthesis discriminant network includes a splicing layer and a global discriminant layer.
[0121] As one embodiment, the alignment relationship between anchor points is input into a pre-trained cross-modal speech synthesis detection model to obtain the forgery probability of audio events in the anchor points and the forgery synthesis probability of the speech audio signals. This includes: receiving the alignment relationship between anchor points using the cross-modal speech synthesis detection model; the normalization layer in the cross-modal alignment encoder converts the alignment relationship between anchor points into an H-row, G-column alignment relationship matrix, where each row in the alignment relationship matrix represents the alignment relationship between an anchor point as an audio event and G anchor points as video events; and the encoding layer sequentially performs convolutional encoding on each row of the alignment relationship matrix to obtain the cross-modal encoded features of each anchor point as an audio event. The fully connected layer in the event authentication network extracts hidden states from cross-modal coding features and uses a discriminant layer based on a two-layer activation function to convert the hidden states into forgery probabilities, which are then used as the forgery probabilities of the audio events associated with the hidden states. The weighting module in the global consistency aggregation network uses the forgery probabilities as weights to weight the cross-modal coding features corresponding to the audio events. The global consistency transformation module projects the weighted results to obtain a global description vector. The concatenation layer in the speech synthesis discriminant network concatenates the global description vector and the forgery probabilities corresponding to each audio event. The final discriminant layer based on a fully connected layer structure discriminates the concatenation result and outputs the forgery synthesis probability of the speech audio signal.
[0122] Specifically, align the h-th row in the relation matrix. anchor point The alignment relationship between the G anchor points used as video events, where the convolutional coding formula of the coding layer is:
[0123]
[0124] in, Indicates anchor point Cross-modal coding features, This represents a trainable encoded weight matrix. This represents the trainable encoding bias matrix.
[0125] Specifically, cross-modal coding features Convert to forgery probability The formula is:
[0126]
[0127] in, Represents cross-modal coding features Hidden states in Represents the ReLU activation function. This indicates the activation function, and the chosen activation function is the Sigmoid function. This represents the trainable hidden layer weight parameters. This represents the bias parameters of the trainable hidden layer. This represents the trainable discriminant layer weight parameters. This represents the bias parameters of the trainable discriminant layer;
[0128] Alternatively, the ReLU activation function can be replaced with the Leaky ReLU activation function or the GELU activation function.
[0129] Specifically, the formula for projecting the weighted result by the global consistency transformation module is as follows:
[0130]
[0131] in, Represents the global description vector. This indicates the weighted result. This represents the trainable projective weight parameters. This represents the trainable bias parameters. This represents the hyperbolic tangent function.
[0132] As an embodiment of this application, by collecting the alignment relationship between multiple sets of anchor points, as well as the true forgery probability (0 or 1, 1 indicates that the audio event is a forged signal segment) of each anchor point as an audio event and the true forgery synthesis probability of the speech audio signal (0 or 1, 1 indicates that the speech audio signal is a forged signal), a loss function is constructed with the objective of minimizing the difference between the true forgery probability / true forgery synthesis probability and the forgery probability of the audio event and the forgery synthesis probability of the speech audio signal output by the model. The trainable parameters in the model are trained and optimized using the gradient descent algorithm or the Adam optimizer.
[0133] Reference Figure 2 The diagram shown is a schematic of a cross-modal anchor point graph provided in an embodiment of this application. E1_1, E1_2, and E1_3 are anchor points corresponding to audio events, and there are edge relationships between them and anchor points E2_1 and E2_2 corresponding to video events. The edge relationship represents the uncertainty between events, which is comprehensively measured by the confidence and similarity of the events. The higher the edge relationship, the lower the degree of event matching between the audio event and the video event.
[0134] In summary, the multimodal speech synthesis detection method of this application effectively identifies the true start and end points of speech by synchronously determining the short-time energy mean, standard deviation, and fundamental frequency transition amplitude. This avoids the problem of missed weak sounds or false noise detection caused by relying solely on energy. After segmenting audio events, the confidence coefficients are calculated at the beginning and end of each audio event and averaged. Combined with the normalized confidence expression constructed from the short-time energy mean, fundamental frequency mean, and standard deviation within each audio event, the authenticity, continuity, and speech stability of each audio event can be quantified, enhancing the reliability assessment at the event level. The Sigmoid activation function is used to stabilize the confidence level within the range of 0 to 1, facilitating subsequent cross-modal alignment analysis. Furthermore, by constructing a cross-modal alignment-driven cross-modal speech synthesis detection model, this application achieves fine-grained consistency modeling between audio and video events, significantly improving the accuracy and robustness of forged speech detection. Specifically, the cross-modal alignment encoder uses a normalization layer to transform anchor-level cross-modal correspondences into an H×G alignment matrix, enabling the model to capture local offset features of cross-modal events in time, semantics, and motion trajectory. The encoding layer enhances the coupling and expressive capabilities between cross-modals by extracting cross-modal encoding features row by row. The audio event authentication network, based on a two-layer activation structure, maps cross-modal encoding features to audio event-level forgery probabilities, revealing forgery patterns at the event granularity level and effectively identifying hidden anomalies such as local inconsistencies, micro-forgeries, and cross-frame discontinuities. The global consistency aggregation network uses event forgery probabilities as adaptive weights, emphasizing the contribution of key suspicious events to global discrimination. It generates stable global description vectors through a projection structure, thereby achieving unified extraction of forgery features from the event level to the global level. Finally, the discriminant layer integrates global descriptions and event forgery probabilities to achieve high-precision prediction of speech forgery synthesis probability, ultimately realizing multi-level, multi-scale cross-modal consistency authentication with interpretability, real-time performance, and generalization capabilities.
[0135] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0136] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0137] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0138] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0139] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0140] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0141] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0142] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0143] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0144] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0145] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0146] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-modal based speech synthesis detection method, characterized in that, The method comprises the following steps: obtaining a speech audio signal and a corresponding associated video modal signal; extracting key audio detection features of the speech audio signal at different signal points, and extracting key video detection features of the video modal signal at different signal points, wherein the key audio detection features include short-time energy and fundamental frequency, and the key video detection features include lip opening and closing event detection results and lip width; dividing the speech audio signal into multiple audio events according to the key audio detection features, and dividing the video modal signal into multiple video events according to the key video detection features; calculating the uncertainty between the audio events and the video events, and constructing a cross-modal anchor point graph by taking the audio events and the video events as anchor points and the uncertainty between the anchor points as edge relationships; performing anchor point time drift-based graph representation learning on the cross-modal anchor point graph to obtain alignment relationships between the anchor points; inputting the alignment relationships between the anchor points into a pre-trained cross-modal speech synthesis detection model to obtain a forgery probability of the audio events in the anchor points and a forgery synthesis probability of the speech audio signal; wherein dividing the speech audio signal into multiple audio events according to the key audio detection features comprises: generating audio event segmentation judgment results of the speech audio signal at different signal points according to the key audio detection features, dividing the speech audio signal into multiple non-overlapping audio signal segments according to the audio event segmentation judgment results, taking each audio signal segment as an independent audio event, and calculating the confidence of the audio event; wherein dividing the video modal signal into multiple video events according to the key video detection features comprises: generating video event segmentation judgment results of the video modal signal at different signal points according to the key video detection features, dividing the video modal signal into multiple non-overlapping video signal segments according to the video event segmentation judgment results, taking each video signal segment as an independent video event, and calculating the confidence of the video event; wherein performing anchor point time drift-based graph representation learning on the cross-modal anchor point graph to obtain alignment relationships between the anchor points comprises: obtaining event timestamps corresponding to the anchor points, wherein the event timestamp is the median of the timestamps of the continuous signal points corresponding to the anchor points; calculating anchor point time offsets between different anchor points according to the event timestamps; generating alignment relationships between the anchor points based on the anchor point time offsets between different anchor points and the edge relationships; wherein the cross-modal speech synthesis detection model comprises a cross-modal alignment encoder, an audio event forgery judgment network, a global consistency aggregation network, and a speech synthesis discrimination network, wherein the cross-modal alignment encoder comprises a standardization layer and an encoding layer, the audio event forgery judgment network comprises a fully connected layer and a discrimination layer, the global consistency aggregation network comprises a weighting module and a global consistency conversion module, and the speech synthesis discrimination network comprises a concatenation layer and a global discrimination layer.
2. The multi-modal based speech synthesis detection method of claim 1, wherein, extracting key video detection features of the video modal signal at different signal points comprises: The face key point detector is used to detect the face key points of the signal values of the video modal signal at different signal points, so as to obtain a contour point set of the upper lip and the lower lip. The opening and closing height of the lip is calculated according to the absolute value of the difference between the midpoint coordinates of the contour points of the upper lip and the midpoint coordinates of the contour points of the lower lip. The opening and closing event detection result of the lip is generated according to the change value of the adjacent opening and closing height of the lip. The lip width is generated according to the coordinate difference between the horizontal direction coordinates of the left contour point of the lip and the horizontal direction coordinates of the right contour point.
3. The multi-modal based speech synthesis detection method of claim 1, wherein, The determination formula of the audio event segmentation determination result is: wherein, represents a speech audio signal represents an audio event segmentation decision result at the nth signal point, represents a conditional predicate, if the formula in the conditional predicate is true, the output of the conditional predicate is True, otherwise the output of the conditional predicate is False, represents a speech audio signal represents a short-time energy at the nth signal point, represents a speech audio signal represents a short-time energy mean value at N signal points, represents a speech audio signal represents a short-time energy standard deviation at N signal points, represents an audio event control parameter, represents a speech audio signal represents a fundamental frequency at the nth signal point, represents a speech audio signal represents a fundamental frequency at the n-1th signal point, represents a speech audio signal represents a fundamental frequency variation value at the nth signal point, represents a preset fundamental frequency variation threshold value, represents a logical AND operator; N represents a signal length of a speech audio signal ; When both sides of the logical and operator output True, then the output is True, indicating that the nth signal point of the speech audio signal is an audio event segmentation point; if the output is not True, indicating that the nth signal point of the speech audio signal is not an audio event segmentation point; According to the audio event segmentation points in the speech audio signal, the speech audio signal is divided into a plurality of mutually non-overlapping audio signal segments, and each audio signal segment is regarded as an independent audio event, and the confidence of the audio event is calculated.
4. The multi-modal based speech synthesis detection method of claim 1, wherein, The determination formula of the video event segmentation determination result is: in, Represents video modal signal The video event segmentation determination result at the m-th signal point. Represents video modal signal The lip opening and closing event detection result at the m-th signal point A value of 1 indicates that a lip opening / closing event was detected. A value of 0 indicates that no lip opening / closing event was detected. Represents video modal signal The lip width at the m-th signal point, Represents video modal signal The lip width at the (m-1)th signal point, This represents the preset threshold for lip width variation, and M represents the video modal signal. The signal length; When and are both output as True, then is output as True, indicating that the mth signal point of the video modality signal is a video event segmentation point; if is not output as True, it indicates that the mth signal point of the video modality signal is not a video event segmentation point; According to the video event segmentation points in the video modal signal, the video modal signal is divided into a plurality of mutually non-overlapping video signal segments, and each video signal segment is regarded as an independent video event, and the confidence of the video event is calculated.
5. The multi-modal based speech synthesis detection method of claim 2, wherein, The uncertainty between the audio event and the video event is calculated, including: The confidence corresponding to the audio event and the video event is obtained; The audio event and the video event are respectively converted into a frequency spectrum sequence, and the similarity between the frequency spectrum sequences corresponding to the audio event and the video event is calculated as the multi-modal event similarity between the audio event and the video event; According to the confidence corresponding to the audio event and the video event, the multi-modal event similarity between the audio event and the video event is converted into uncertainty, wherein the lower the uncertainty, the higher the event matching degree between the audio event and the video event, and the audio event and the video event describe similar events.
6. The multi-modal based speech synthesis detection method of claim 5, wherein, The video event is converted into a frequency spectrum sequence, including: The lip opening and closing height and the lip width at each signal point in the video event are extracted and spliced as a feature vector of the video event; The feature vector is mapped into S pseudo-Mel energies by using a nonlinear mapping method; The S pseudo-Mel energies are spliced as the frequency spectrum sequence of the video event.
7. The multi-modal based speech synthesis detection method of claim 1, wherein, The structure of the cross-modal anchor point graph is: in, Represents a cross-modal anchor point diagram. Represents cross-modal anchor point diagram The set of anchor points in the middle, These are, in order, audio event-anchor set and video event-anchor set. Represents audio event-anchor set The h-th anchor point in the sequence, where H represents the number of audio events. Represents a video event-anchor set The first in There are several anchor points, where G represents the number of video events. Represents cross-modal anchor point diagram The set of edge relations in the middle, Indicates anchor point The edge relationships between them.
8. The multi-modal based speech synthesis detection method of claim 7, wherein, The alignment relationship between the anchor points is input into a pre-trained cross-modal speech synthesis detection model to obtain the forgery probability of the audio event and the forgery synthesis probability of the speech audio signal, including: The cross-modal speech synthesis detection model receives the alignment relationship between the anchor points, the standardization layer in the cross-modal alignment encoder converts the alignment relationship between the anchor points into an alignment relationship matrix with H rows and G columns, wherein each row in the alignment relationship matrix represents the alignment relationship between the anchor point as the audio event and the G anchor points as the video event, and the encoding layer convolves and encodes each row in the alignment relationship matrix in turn to obtain the cross-modal encoding feature of each anchor point as the audio event; The full connection layer in the audio event authentication network extracts the hidden state in the cross-modal encoding feature, and adopts a discrimination layer based on a double-layer activation function to convert the hidden state into a forgery probability as the forgery probability of the audio event associated with the hidden state; The weighting module in the global consistency aggregation network takes the forgery probability as the weight to weight the cross-modal encoding feature corresponding to the audio event, and the global consistency conversion module projects the weighting result to obtain a global description vector; The splicing layer in the speech synthesis discrimination network splices the global description vector and the forgery probability corresponding to each audio event, and adopts a final discrimination layer based on a full connection layer structure to discriminate the splicing result and output the forgery synthesis probability of the speech audio signal.
Citation Information
Patent Citations
Deep forgery detection method and system based on audio and video multi-mode fusion
CN120580481A
Multi-modal global and local collaboration-based speaking face generation video detection method and device
CN120823537A