Audio deep forgery detection method and system based on sentiment analysis
Through a method based on sentiment analysis, transfer learning and random forest classifiers are used to extract high-level emotional features of speech, which solves the robustness problem of deep fake speech detection in different data sets and noisy environments, and achieves high-accuracy detection results.
Patent Information
- Application Number
- CN202510964546.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have difficulty in effectively detecting deepfake speech, especially due to insufficient robustness in different data sets and noisy environments, and the detection effects of traditional methods are not stable enough.
A method based on sentiment analysis is adopted to detect the high-level emotional features of speech by extracting them. The emotion feature vector of the three-dimensional input matrix is extracted from the pre-trained speech emotion recognition system using a transfer learning strategy, and a random forest classifier is combined for binary classification to enhance the robustness of the system in noisy environments.
It achieves effective distinction between real speech and synthesized speech, improves detection accuracy, especially the detection performance across data sets and in noisy environments, reduces system implementation complexity and enhances generalization capabilities.
Smart Images

Figure CN120708625A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security technology, and more specifically to a method and system for detecting audio deep forgery based on sentiment analysis. Background Art
[0002] In recent years, with the rapid development of deep learning technology, audio and video deepfakes have made significant progress. In particular, in the field of speech synthesis, text-to-speech (TTS) and voice conversion (VC) technologies have become capable of generating highly realistic artificial speech, making it easier than ever to fake someone else's voice. While this technological advancement has had positive impacts in areas such as entertainment and communication, it also poses serious security risks.
[0003] Today's hyperconnected world allows multimedia content to be disseminated globally in real time, facilitating the rapid spread of fake content. Numerous cases have been reported using deepfake technology to commit fraud, spread false information, and damage personal reputations, seriously impacting the security and reliability of voice recognition and identity verification systems, as well as the public information environment.
[0004] There are various approaches to detecting deepfakes in audio. Traditional methods primarily rely on analyzing low-level features of speech, such as linear filter bank features and long- and short-term prediction features. For example, some research has fed linear filter banks into ResNet networks to generate embedding vectors for classifying real and fake speech; others have leveraged long-term features to distinguish between authentic and forged audio tracks; others have detected deepfakes based on long- and short-term prediction features, and have used traces left by time scaling to identify forged audio signals.
[0005] However, with the continuous advancement of deepfake technology, detection methods that rely solely on low-level acoustic features face increasing challenges. Modern text-to-speech (TTS) and voice-cognition (VC) systems are now able to closely mimic the fundamental acoustic properties of human speech, rendering detection methods based on these features less effective. These methods often exhibit unstable performance, particularly in cross-dataset scenarios and in noisy environments.
[0006] Recent research has shown that while deepfake technology can effectively synthesize low-level speech features, it still struggles to reproduce more complex speech characteristics, particularly emotional expressions. Speech emotion recognition (SER), a rapidly developing research field, has become effective in extracting emotion-related semantic features from speech. This provides new insights into detecting deepfakes at a semantic level.
[0007] Some research has explored the use of semantic features for deepfake analysis in audio and video, demonstrating the feasibility of this approach. However, there is currently a lack of systematic methods for deepfake detection using emotional features of speech, especially research on its robustness across diverse datasets and noisy environments.
[0008] Therefore, how to provide an audio deep fake detection method and system based on sentiment analysis to achieve effective detection of deep fake voices is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0009] In view of this, the present invention provides an audio deep fake detection method and system based on sentiment analysis. Based on the shortcomings of deep fake technology in emotional expression, the method extracts high-level emotional features of speech for detection, which can effectively distinguish real speech from synthesized speech.
[0010] In order to achieve the above object, the present invention adopts the following technical solutions:
[0011] In one aspect, the present invention provides an audio deepfake detection method based on sentiment analysis, comprising:
[0012] Step 1: Preprocess the input speech and construct a three-dimensional input matrix based on the preprocessed input speech;
[0013] Step 2: extracting the emotional feature vector of the three-dimensional input matrix from a pre-trained speech emotion recognition system using a transfer learning strategy;
[0014] Step 3: Perform synthetic speech detection based on the emotional feature vector to determine the authenticity of the input speech.
[0015] Preferably, the pretreatment includes:
[0016] Converting the input speech into a monophonic input speech;
[0017] Resampling the monophonic input speech to obtain input speech at a standard sampling frequency;
[0018] Filtering the input speech at the standard sampling frequency using a bandpass filter;
[0019] Normalize the filtered input speech;
[0020] The normalized input speech is subjected to speech interception or padding to obtain an input speech of standard speech length.
[0021] Preferably, constructing a three-dimensional input matrix based on the pre-processed input speech includes:
[0022] Performing a short-time Fourier transform on the preprocessed input speech to obtain a short-time Fourier transform amplitude spectrum;
[0023] Converting the short-time Fourier transform amplitude spectrum into a Mel spectrum;
[0024] Performing logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum graph;
[0025] Calculating the first-order discrete derivative and the second-order discrete derivative of the logarithmic Mel-spectrogram along the frequency axis;
[0026] The logarithmic Mel-spectrogram, the first-order discrete derivative, and the second-order discrete derivative are stacked along a third dimension to form a three-dimensional matrix.
[0027] Preferably, the speech emotion recognition system used in step 2 is a pre-trained 3D convolutional recurrent neural network structure, including a 3D convolutional layer, a linear layer, a bidirectional long short-term memory network, an attention layer and a dense layer.
[0028] Preferably, a transfer learning strategy is used to extract the emotion information of the three-dimensional input matrix and the emotion feature vector of the identified emotion from a pre-trained speech emotion recognition system, including:
[0029] The 3D convolution receives the three-dimensional input matrix and extracts local time-frequency features to capture short-term acoustic patterns in the speech signal;
[0030] The linear layer receives the output of the 3D convolutional layer and performs nonlinear transformation to enhance the expressiveness of features;
[0031] The bidirectional long short-term memory network receives the features transformed by the linear layer, models the long-term temporal dependency of the speech, and captures and characterizes the emotional dynamic features of the input speech that change over time;
[0032] The attention layer receives the temporal features output by the bidirectional long short-term memory network and performs weighted fusion to generate the emotion feature vector.
[0033] Preferably, the step 3 uses a random forest classifier to perform binary classification on the emotional feature vector to obtain the authenticity of the input speech.
[0034] In another aspect, the present invention provides an audio deepfake detection system based on sentiment analysis, comprising:
[0035] A processing module, configured to pre-process the input speech and construct a three-dimensional input matrix based on the pre-processed input speech;
[0036] An emotion feature extraction module, configured to extract an emotion feature vector of the three-dimensional input matrix using a pre-trained speech emotion recognition system;
[0037] The speech detection module is used to perform synthetic speech detection based on the emotional feature vector to determine the authenticity of the input speech.
[0038] It can be seen from the above technical solution that compared with the existing technology, the present invention discloses a method and system for detecting audio deep fakes based on sentiment analysis, which has the following beneficial effects:
[0039] 1. Taking advantage of the shortcomings of deep fake technology in emotional expression, the method analyzes the high-level emotional features of speech for detection, which can effectively distinguish real speech from synthesized speech, especially for deep fake speech generated by TTS and hybrid TTS / VC, with a high detection accuracy.
[0040] 2. A transfer learning strategy is adopted to extract emotional features using a pre-trained speech emotion recognition system. This eliminates the need to redesign complex feature extraction networks for deepfake detection tasks, reducing the complexity of system implementation.
[0041] 3. It performs well in cross-dataset scenarios, can detect deep fake speech from different sources and different generation algorithms, and has strong generalization ability.
[0042] 4. The synthetic speech detection module is trained through data enhancement strategy, which improves the robustness of the system in noisy environments and can still maintain good detection performance in a noisy environment with a signal-to-noise ratio of 10dB.
[0043] 5. The present invention uses a three-dimensional matrix representation to capture emotional dynamics, combines transfer learning application strategies to extract emotional features from the attention layer, and uses a single vector to holistically analyze the detection method of multiple emotional dimensions, solving the problem of insufficient robustness of existing technologies in cross-dataset and noisy environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0045] Figure 1 This is a schematic diagram of the overall process provided by the present invention.
[0046] Figure 2 This is the flowchart for step 1.
[0047] Figure 3 This is the flowchart for step 2.
[0048] Figure 4 This is the flowchart for step 3.
[0049] Figure 5 This is a structural diagram provided by the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] The embodiment of the present invention discloses one aspect, which provides an audio deep fake detection method based on sentiment analysis, such as Figure 1 As shown, including:
[0052] Step 1: Preprocess the input speech and construct a three-dimensional input matrix based on the preprocessed input speech;
[0053] Step 2: Use transfer learning strategy to extract the emotional feature vector of the three-dimensional input matrix from the pre-trained speech emotion recognition system;
[0054] Step 3: Perform synthetic speech detection based on the emotional feature vector to determine the authenticity of the input speech.
[0055] Further, if Figure 2 As shown, the preprocessing includes:
[0056] Convert the input speech into a mono input speech;
[0057] Resample the mono input speech to obtain the input speech at the standard sampling frequency; the standard sampling frequency is F s =16kHz.
[0058] Use a bandpass filter to filter the input speech at the standard sampling frequency; specifically, use a 6th-order Butterworth bandpass digital filter to filter the speech signal, with a low cutoff frequency F l =250Hz, high cutoff frequency F h =3600Hz.
[0059] The filtered input speech is normalized. Specifically, each speech is normalized using an infinite norm.
[0060] The normalized input speech is intercepted or padded to obtain the input speech of standard speech length. In this embodiment, the standard speech length L is set to cut =3s.
[0061] Further, if Figure 2 As shown, a three-dimensional input matrix is constructed based on the preprocessed input speech, including:
[0062] The short-time Fourier transform (STFT) of the preprocessed input speech is obtained by applying the short-time Fourier transform (STFT) to the input speech. The Hamming window is used to calculate the short-time Fourier transform (STFT) of the input speech. The window length L w =0.025s, step length L h = 0.01s, only the STFT amplitude spectrum is retained.
[0063] The magnitude spectrum of the short-time Fourier transform is converted into a Mel spectrum. Specifically, the STFT magnitude spectrum is processed by a Mel filter bank and scaled by applying a natural logarithm function to obtain a logarithmic Mel spectrum graph S mel , expressed as:
[0064]
[0065] Wherein, M represents the number of time windows, which is set to 300 in this embodiment, corresponding to the time resolution of 3 seconds of speech; K represents the number of Mel filter banks, which is set to 40 in this embodiment, used to capture the frequency characteristics of speech. mel It is a two-dimensional matrix representing the log-mel spectrogram of the speech.
[0066] Perform logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum graph;
[0067] Calculate the first-order discrete derivative and second-order discrete derivative of the logarithmic Mel spectrum along the frequency axis;
[0068] The logarithmic Mel spectrogram, first-order discrete derivative, and second-order discrete derivative are stacked along the third dimension to form a three-dimensional matrix, which is expressed as:
[0069]
[0070] Where, ΔS mel and ΔΔS mel They are S mel The first-order and second-order discrete derivatives along the frequency axis are used to capture the dynamic characteristics of frequency changes; X is a three-dimensional matrix formed by stacking these three matrices in the third dimension, which contains the static and dynamic characteristics of speech.
[0071] Furthermore, after stacking, the three-dimensional matrix X is subjected to Z-score normalization.
[0072] Furthermore, the speech emotion recognition system used in step 2 is a pre-trained 3D convolutional recurrent neural network structure, including a 3D convolutional layer, a linear layer, a bidirectional long short-term memory network, an attention layer, and a dense layer. In this invention, only the emotion feature vector output by the attention layer needs to be extracted.
[0073] like Figure 3 As shown in the figure, a transfer learning strategy is used to extract the emotional information of the three-dimensional input matrix and the emotional feature vector of the recognized emotion from the pre-trained speech emotion recognition system, including:
[0074] The 3D convolutional layer receives a three-dimensional input matrix and extracts local time-frequency features to capture short-term acoustic patterns in the speech signal;
[0075] The linear layer receives the output of the 3D convolutional layer and performs nonlinear transformation to enhance the expressiveness of features;
[0076] The bidirectional long short-term memory network receives the features transformed by the linear layer, models the long-term temporal dependencies of the speech, and captures and characterizes the emotional dynamics of the input speech that changes over time.
[0077] The attention layer receives the temporal features output by the bidirectional long short-term memory network and performs weighted fusion to generate a sentiment feature vector.
[0078] The sentiment feature vector characterizes the sentiment type through its position in the high-dimensional space, the sentiment intensity through its norm or relative position, and the sentiment change through its property as a weighted summary of the overall time series.
[0079] In this example, the SER system is pre-trained on the IEMOCAP dataset and can recognize four emotion categories: anger, happiness, sadness, and neutral. The training is performed using the 1st to 4th dialogue sessions of the IEMOCAP dataset, and the 5th session is used for development and testing. The Adam optimizer is used with a learning rate of l r =10 -5 The loss function is categorical cross entropy. The trained SER system achieves a balanced accuracy of 0.6 on four emotion recognition tasks.
[0080] Using the transfer learning strategy, a 256-dimensional feature vector F is extracted from the attention layer output of the SER system. x , as the emotional representation of speech. Formally, it can be expressed as formula (c), where is the feature extraction function, and x is the input speech.
[0081]
[0082] in, Represents the function mapping of features extracted from the pre-trained SER system, which maps the input speech x to a 256-dimensional feature space. This feature vector F x The emotional semantics of speech are captured, including high-level features such as emotion intensity, emotion type, and emotion variation. These features are important for distinguishing real speech from deepfakes. Since deepfake technology has shortcomings in reproducing natural emotional behavior, these emotional features can effectively reveal abnormal patterns in deepfakes.
[0083] like Figure 4 As shown, in this embodiment, a random forest classifier is used to perform binary classification on the emotional feature vector to obtain the authenticity of the input speech.
[0084] The hyperparameters of the random forest classifier are optimized by grid search on the validation set, using balanced accuracy (BA) as the evaluation metric. The hyperparameters considered include the split quality criterion (Gini impurity or information gain) and the number of learners N. RF =[10,30,100,300].
[0085] After optimization, the best random forest configuration is: the split quality standard uses information gain, the number of learners N RF = 300. Classifier output category y:
[0086] y∈{REAL,DF}
[0087] Among them, REAL represents real voice and DF represents deep fake voice.
[0088] The random forest classifier analyzes the input sentiment feature vector F x The algorithm uses patterns in the speech to determine whether the speech exhibits natural emotional characteristics. If the speech exhibits natural and consistent emotional characteristics, it tends to be classified as real speech; if the speech exhibits unusual or inconsistent emotional characteristics, it tends to be classified as a deepfake. This classification method based on high-level semantic features can overcome the limitations of traditional detection methods based on low-level acoustic features.
[0089] In order to improve the robustness of the system in noisy environments, the present invention also adopts a data augmentation strategy to add white noise of different intensities to the training data of the random forest classifier. The specific implementation method is as follows:
[0090] For the training set and validation set, white noise is added according to the two-layer probability distribution:
[0091] First layer: white noise with a signal-to-noise ratio (SNR) of 30 dB to 15 dB is added with probability p1 = 0.8;
[0092] Second layer: white noise with a signal-to-noise ratio of 15dB to 10dB is added with probability p2=0.3.
[0093] For the test set, white noise with a fixed SNR of [25, 20, 15, 10] dB was added for comparison. This design ensures that the training data contains a variety of noise levels, while the test data is obtained in a controlled environment, facilitating analysis of the results.
[0094] On the other hand, the present invention provides an audio deep fake detection system based on sentiment analysis, such as Figure 5 As shown, including:
[0095] A processing module, configured to pre-process the input speech and construct a three-dimensional input matrix based on the pre-processed input speech;
[0096] The emotion feature extraction module is used to extract the emotion feature vector of the three-dimensional input matrix from the pre-trained speech emotion recognition system using a transfer learning strategy;
[0097] The speech detection module is used to detect synthesized speech based on the emotional feature vector and determine the authenticity of the input speech.
[0098] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0099] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio deepfake detection method based on sentiment analysis, characterized in that include: Step 1: Preprocess the input speech and construct a three-dimensional input matrix based on the preprocessed input speech; Step 2: extracting the emotional feature vector of the three-dimensional input matrix from a pre-trained speech emotion recognition system using a transfer learning strategy; Step 3: Perform synthetic speech detection based on the emotional feature vector to determine the authenticity of the input speech.
2. The method for detecting deep fake audio based on sentiment analysis according to claim 1, characterized in that: The pretreatment includes: Converting the input speech into a monophonic input speech; Resampling the monophonic input speech to obtain input speech at a standard sampling frequency; Filtering the input speech at the standard sampling frequency using a bandpass filter; Normalize the filtered input speech; The normalized input speech is subjected to speech interception or padding to obtain an input speech of standard speech length.
3. The method for detecting deep fake audio based on sentiment analysis according to claim 1, characterized in that: Construct a three-dimensional input matrix based on the preprocessed input speech, including: Performing a short-time Fourier transform on the preprocessed input speech to obtain a short-time Fourier transform amplitude spectrum; Converting the short-time Fourier transform amplitude spectrum into a Mel spectrum; Performing logarithmic transformation on the Mel spectrum to obtain a logarithmic Mel spectrum graph; Calculating the first-order discrete derivative and the second-order discrete derivative of the logarithmic Mel-spectrogram along the frequency axis; The logarithmic Mel-spectrogram, the first-order discrete derivative, and the second-order discrete derivative are stacked along a third dimension to form a three-dimensional matrix.
4. The method for detecting deep fake audio based on sentiment analysis according to claim 1, characterized in that: The speech emotion recognition system used in step 2 is a pre-trained 3D convolutional recurrent neural network structure, including a 3D convolutional layer, a linear layer, a bidirectional long short-term memory network, an attention layer, and a dense layer.
5. The method for detecting deep fake audio based on sentiment analysis according to claim 4, characterized in that: A transfer learning strategy is used to extract the emotional information of the three-dimensional input matrix and the emotional feature vector of the recognized emotion from a pre-trained speech emotion recognition system, including: The 3D convolutional layer receives a three-dimensional input matrix and extracts local time-frequency features to capture short-term acoustic patterns in the speech signal; The linear layer receives the output of the 3D convolutional layer and performs nonlinear transformation to enhance the expressiveness of the features; The bidirectional long short-term memory network receives the features transformed by the linear layer, models the long-term temporal dependency of the speech, and captures and characterizes the emotional dynamic features of the input speech that change over time; The attention layer receives the temporal features output by the bidirectional long short-term memory network and performs weighted fusion to generate the emotion feature vector.
6. The method for detecting deep fake audio based on sentiment analysis according to claim 1, characterized in that: The step 3 uses a random forest classifier to perform binary classification on the emotional feature vector to obtain the authenticity of the input speech.
7. An audio deepfake detection system based on sentiment analysis, characterized in that include: A processing module, configured to pre-process the input speech and construct a three-dimensional input matrix based on the pre-processed input speech; An emotion feature extraction module, configured to extract an emotion feature vector of the three-dimensional input matrix from a pre-trained speech emotion recognition system using a transfer learning strategy; The speech detection module is used to perform synthetic speech detection based on the emotional feature vector to determine the authenticity of the input speech.