Synthetic speech detection method based on emotion feature fusion and gradient adversarial training
Through the method based on emotional feature fusion and gradient adversarial training, the problem of poor efficiency, accuracy and adaptability of synthetic speech detection in the prior art is solved, especially under noise conditions, which significantly improves performance, achieving more efficient and accurate synthetic speech detection.
Patent Information
- Application Number
- CN202510270594.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has poor detection efficiency, accuracy and adaptability when detecting synthetic speech, especially under noisy conditions, performance has significantly decreased.
A synthetic speech detection method based on emotional feature fusion and gradient adversarial training is adopted. By obtaining the audio data set and preprocessing, adding noise to data enhancement, extracting time-frequency features and emotional features, performing feature fusion, and using the Res2Net network for detection, and adding a gradient inversion layer for adversarial training.
The efficiency, accuracy and adaptability of synthetic speech detection are improved, especially under noisy conditions to detect detection performance.
Smart Images

Figure CN120108378A_ABST
Abstract
Description
Technical Field
[0001] The disclosed embodiments relate to the field of data processing technology, and in particular to a synthetic speech detection method based on emotion feature fusion and gradient adversarial training. Background Art
[0002] At present, with the development of deep learning technology, synthetic speech technology has become more and more realistic, and can simulate the voices of specific people to spread false information. By using deep fake technology to imitate the voices of social celebrities, some misleading or offensive audio recordings have been released, which have a negative impact on public opinion and social stability, and also pose a threat to the security of the automatic speaker verification system (ASV). At present, there are two main types of forgery technologies represented by text-to-speech (TTS) and voice conversion (VC). The significant development of voice deep fake technology has made forged voices more and more realistic, and it is difficult to detect them manually.
[0003] At present, the detection of synthetic speech mainly adopts the combined architecture of front-end feature extraction and back-end binary classifier. Based on this framework, existing work mainly focuses on the development of manual features, such as Mel-frequency cepstral coefficients (MFCC), linear frequency cepstral coefficients (LFCC), constant Q transform (CQT), etc., and then uses these features to train classifiers. In terms of classifiers, Gaussian mixture models (GMM), support vector machines (SVM), neural network models, and residual network (ResNet) models are widely used. However, different speech features focus on different information in speech signals. A single speech feature will have a good effect when processing specific speech signals, but it cannot fully capture all the characteristics of speech when processing complex speech signals, resulting in insufficient generalization ability of the detection system when facing unknown data. At the same time, under noisy conditions, the detection performance of a single feature may also be significantly reduced.
[0004] It can be seen that there is an urgent need for a synthetic speech detection method based on emotion feature fusion and gradient adversarial training with high detection efficiency, accuracy and adaptability. Summary of the invention
[0005] In view of this, the embodiments of the present disclosure provide a synthetic speech detection method based on emotion feature fusion and gradient adversarial training, which at least partially solves the problems of poor detection efficiency, accuracy and adaptability in the prior art.
[0006] In a first aspect, the present disclosure provides a synthetic speech detection method based on emotion feature fusion and gradient adversarial training, comprising:
[0007] Step 1, obtaining an audio data set and preprocessing it to obtain a training set, wherein the audio data set includes real speech data and synthesized speech data;
[0008] Step 2: Add noise with different signal-to-noise ratios to the training set for data enhancement;
[0009] Step 3, the training set after data enhancement is subjected to fast Fourier transform and then to Mel filter to obtain Mel spectrum and obtain time-frequency features based on it;
[0010] Step 4, inputting the Mel spectrum into the speech emotion recognition module to obtain the emotion feature;
[0011] Step 5, using the increase and decrease component method to select key speech features in the time-frequency features and the emotional features and fuse the selected features to obtain fused features;
[0012] Step 6: Input the fused features into the Res2Net network for synthetic speech detection;
[0013] Step 7, add a gradient reversal layer to the classification module, and classify the detected synthetic speech based on it, and perform adversarial training to obtain a target model to detect the speech to be detected.
[0014] According to a specific implementation of the embodiment of the present disclosure, step 1 specifically includes:
[0015] Step 1.1, pre-emphasize the high frequency part of the audio data set;
[0016] In step 1.2, the pre-emphasized audio data set is divided into frames and a Hamming window is applied to each frame.
[0017] According to a specific implementation of the embodiment of the present disclosure, the expression of the Hamming window is:
[0018]
[0019] Among them, n is the number of sample points and N is the length of the window.
[0020] According to a specific implementation of the embodiment of the present disclosure, step 3 specifically includes:
[0021] Step 3.1, the training set after data enhancement is subjected to fast Fourier transform to obtain the frequency domain signal;
[0022] Step 3.2, input the frequency domain signal into a set of Mel filters, take the logarithm of the output of each Mel filter, and obtain the logarithmic energy of the Mel spectrum;
[0023] Step 3.3, perform discrete cosine transform on the logarithmic Mel frequency spectrum to obtain Mel frequency cepstral coefficients, and retain a preset number of Mel frequency cepstral coefficients as time-frequency features.
[0024] According to a specific implementation of the embodiment of the present disclosure, step 4 specifically includes:
[0025] Step 4.1, perform convolution operation on the logarithmic Mel spectrum to obtain the convolution sequence and input it into the bidirectional long short-term memory neural network for time summarization;
[0026] Step 4.2: The output h of the bidirectional long short-term memory neural network is t Normalized through the attention layer to obtain the importance weight a t
[0027]
[0028] Among them, t represents the time step, T represents the total length of the sequence, that is, the total number of time steps, and W represents the weight matrix;
[0029] In step 4.3, the output of the bidirectional long short-term memory neural network is weighted and summed according to the importance weights to obtain the discourse-level representation and use it as the sentiment feature.
[0030] According to a specific implementation of the embodiment of the present disclosure, step 5 specifically includes:
[0031] Step 5.1, using the increase and decrease component method to calculate the average contribution of the feature components in the time-frequency feature and the emotional feature, and based on this, identify the feature component that has the greatest impact on the synthetic speech detection and delete other features;
[0032] In step 5.2, the filtered time-frequency features and emotional features are used as the horizontal axis and vertical axis of the new speech feature space respectively, and the time-frequency features and emotional features are matrix multiplied to obtain the fusion features while ensuring that the frame length and frame shift of the two are consistent.
[0033] According to a specific implementation of the embodiment of the present disclosure, the expression of the average contribution is:
[0034]
[0035] Among them, G i represents the average contribution of the i-th dimension feature, and K represents a normalization constant, which is used to scale or average the calculation results to ensure that the contribution calculation is within a reasonable range. Indicates that all feature dimensions after the i-th dimension are operated. It means to operate on all feature dimensions before the i-th dimension, p(i,j) means the recognition accuracy when the i-th to j-th dimension features are used as speech feature parameters, p(i+1,j) means the recognition accuracy when the feature dimension increases from i to i+1, and p(i,j+1) means the recognition accuracy when the feature dimension increases from j to j+1.
[0036] The synthetic speech detection scheme based on emotion feature fusion and gradient adversarial training in the embodiment of the present disclosure includes: step 1, obtaining an audio data set and preprocessing it to obtain a training set, wherein the audio data set includes real speech data and synthetic speech data; step 2, adding noise with different signal-to-noise ratios to the training set to perform data enhancement; step 3, subjecting the data-enhanced training set to fast Fourier transform and then to a Mel filter to obtain a Mel spectrum and thereby obtain time-frequency features; step 4, inputting the Mel spectrum into a speech emotion recognition module to obtain emotion features; step 5, using the increase and decrease component method to screen key speech features in the time-frequency features and emotion features and fusing the screened features to obtain fused features; step 6, inputting the fused features into a Res2Net network to perform synthetic speech detection; step 7, adding a gradient inversion layer to the classification module, and thereby classifying the detected synthetic speech, and performing adversarial training to obtain a target model to detect the speech to be detected.
[0037] The beneficial effects of the embodiments of the present disclosure are as follows: through the scheme of the present disclosure, the defects of single speech feature in synthetic speech detection are solved, and the method of feature fusion of time-frequency features and emotional features is adopted to detect synthetic speech. At the same time, a gradient reversal layer is added to perform adversarial training to optimize the final classification result, thereby improving the detection efficiency, accuracy and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 A flowchart of a synthetic speech detection method based on emotion feature fusion and gradient adversarial training provided in an embodiment of the present disclosure;
[0040] Figure 2 A system overall architecture diagram corresponding to a synthetic speech detection method based on emotion feature fusion and gradient adversarial training provided in an embodiment of the present disclosure;
[0041] Figure 3 An emotion recognition model architecture diagram provided for an embodiment of the present disclosure;
[0042] Figure 4 A workflow diagram of an attention layer provided for an embodiment of the present disclosure;
[0043] Figure 5 A residual block structure diagram provided by an embodiment of the present disclosure;
[0044] Figure 6 A Res2Net model structure diagram provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0045] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0046] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present disclosure.
[0047] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0048] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The drawings only show components related to the present disclosure rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0049] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0050] The disclosed embodiment provides a synthetic speech detection method based on emotion feature fusion and gradient adversarial training, which can be applied to the forged speech detection process in information security scenarios.
[0051] See also Figure 1 , is a flow chart of a synthetic speech detection method based on emotion feature fusion and gradient adversarial training provided by an embodiment of the present disclosure. Figure 1 and Figure 2 As shown, the method mainly comprises the following steps:
[0052] Step 1, obtaining an audio data set and preprocessing it to obtain a training set, wherein the audio data set includes real speech data and synthesized speech data;
[0053] Furthermore, the step 1 specifically includes:
[0054] Step 1.1, pre-emphasize the high frequency part of the audio data set;
[0055] In step 1.2, the pre-emphasized audio data set is divided into frames and a Hamming window is applied to each frame.
[0056] Furthermore, the expression of the Hamming window is
[0057]
[0058] Among them, n is the number of sample points and N is the length of the window.
[0059] In the specific implementation, an audio data set is obtained, including real speech data and synthetic speech data, and the training set and the validation set are divided according to a preset ratio. The original audio data is preprocessed. The spectrum is improved by pre-emphasis to enhance the high-frequency part of the signal and compensate for the loss of high-frequency signals caused by the characteristics of the human vocal organ. Frame processing is performed to divide the continuous signal into several small segments, each of which is a frame. The frame length is usually set to 20-40 milliseconds, and we set the frame length to 20ms. A Hamming window is applied to each frame signal to reduce the discontinuity at the frame boundary and suppress the leakage effect on the spectrum. The mathematical expression of the Hamming window is:
[0060]
[0061] Among them, n is the number of sample points and N is the length of the window.
[0062] The windowed signal of each sample point x(n) is:
[0063] y(n)=x(n))*w(n).
[0064] Step 2: Add noise with different signal-to-noise ratios to the training set for data enhancement;
[0065] In specific implementation, data enhancement technology is used to add Gaussian white noise with different signal-to-noise ratios to the original audio data to obtain a noisy data set.
[0066] noisy_signal=signal+noise
[0067] Among them, noisy_signal is the noisy signal, signal is the original signal, and noise is the generated noise.
[0068] For the training set and validation set, a two-layer probability distribution is used to perform noise injection. The first layer sets the probability p1 = 0.8 and randomly adds Gaussian white noise with a signal-to-noise ratio between 30dB and 15dB; the second layer sets the probability p2 = 0.3 and randomly adds Gaussian white noise with a signal-to-noise ratio between 15dB and 10dB. For the test set, the added Gaussian white noise signal-to-noise ratio SNR = [30, 25, 20, 15]dB is used to enable result analysis.
[0069] Step 3, the training set after data enhancement is subjected to fast Fourier transform and then to Mel filter to obtain Mel spectrum and obtain time-frequency features based on it;
[0070] Based on the above embodiment, step 3 specifically includes:
[0071] Step 3.1, the training set after data enhancement is subjected to fast Fourier transform to obtain the frequency domain signal;
[0072] Step 3.2, input the frequency domain signal into a set of Mel filters, take the logarithm of the output of each Mel filter, and obtain the logarithmic energy of the Mel spectrum;
[0073] Step 3.3, perform discrete cosine transform on the logarithmic Mel frequency spectrum to obtain Mel frequency cepstral coefficients, and retain a preset number of Mel frequency cepstral coefficients as time-frequency features.
[0074] In specific implementation, the time domain signal is converted into a frequency domain signal through the fast Fourier transform (FFT). The human ear's perception of frequency is logarithmic, that is, it is sensitive to changes in the low frequency band and insensitive to changes in the high frequency band. The spectrum is passed through a set of Mel filters, which simulate the human ear's perception sensitivity to different frequencies. The center frequency intervals of the Mel filters are equidistant on the Mel scale, which matches the nonlinear characteristics of the human auditory system. The logarithm of the output of each filter is taken to obtain the logarithmic energy of the Mel spectrum. The logarithmic Mel spectrum is subjected to a discrete cosine transform (DCT) to decorrelate and compress the spectral information to obtain the Mel frequency cepstrum coefficients (MFCC coefficients). The first 13 coefficients of the DCT are usually retained as MFCC features, i.e., time-frequency features.
[0075] Step 4, inputting the Mel spectrum into the speech emotion recognition module to obtain the emotion feature;
[0076] Based on the above embodiment, step 4 specifically includes:
[0077] Step 4.1, perform convolution operation on the logarithmic Mel spectrum to obtain the convolution sequence and input it into the bidirectional long short-term memory neural network for time summarization;
[0078] Step 4.2: The output h of the bidirectional long short-term memory neural network is t Normalized through the attention layer to obtain the importance weight a t
[0079]
[0080] Among them, t represents the time step, T represents the total length of the sequence, that is, the total number of time steps, and W represents the weight matrix;
[0081] In step 4.3, the output of the bidirectional long short-term memory neural network is weighted and summed according to the importance weights to obtain the discourse-level representation and use it as the sentiment feature.
[0082] In the specific implementation, the obtained Mel spectrum map is used as input to obtain the emotional features through the speech emotion recognition module. First, the convolution operation is performed on the entire logarithmic Mel spectrum; then the obtained convolution sequence is input into the LSTM for time summarization; then a series of high-level features are used as input through the attention layer to generate speech-level features; finally, the classification of speech emotions is obtained through the fully connected layer and the softmax classifier. We mainly consider the output of the final attention layer and use it as the input for subsequent synthetic speech detection through transfer learning.
[0083] The CRNN model structure is as follows Figure 3As shown in the figure, it mainly consists of several convolutional layers, a maximum pooling layer, a linear layer and an LSTM layer. Among them, the first convolutional layer has 128 feature maps, and the remaining convolutional layers have 256 feature maps. We only perform maximum pooling after the first convolutional layer, and the pooling size is 2*2. In order to improve the accuracy of the model and reduce the model parameters. Dimensionality reduction is performed by adding a linear layer, and the linear layer has 768 output units. Finally, the CNN sequence is input into the bidirectional long short-term memory neural network (BiLSTM) for temporal summary to obtain a sequence of 256-dimensional high-level feature representation.
[0084] The attention layer is used to focus on the emotion-related parts and produce discriminative utterance-level representations for SER, because not all frame-level CRNN features contribute to speech emotion classification. We adopt the attention model to evaluate the importance of a series of high-level feature representations to the final utterance-level emotion feature representation. Figure 4 As shown, we convert the LSTM output h t , normalized to obtain the importance weight a t Then, according to the weights, h t The weighted sum is performed to obtain the utterance-level representation c. Finally, the utterance-level representation is passed to a fully connected layer with 64 output units to obtain a higher-level feature representation, and speech emotion classification is performed through a softmax classifier.
[0085]
[0086] Step 5, using the increase and decrease component method to select key speech features in the time-frequency features and the emotional features and fuse the selected features to obtain fused features;
[0087] Based on the above embodiment, step 5 specifically includes:
[0088] Step 5.1, using the increase and decrease component method to calculate the average contribution of the feature components in the time-frequency feature and the emotional feature, and based on this, identify the feature component that has the greatest impact on the synthetic speech detection and delete other features;
[0089] In step 5.2, the filtered time-frequency features and emotional features are used as the horizontal axis and vertical axis of the new speech feature space respectively, and the time-frequency features and emotional features are matrix multiplied to obtain the fusion features while ensuring that the frame length and frame shift of the two are consistent.
[0090] Furthermore, the expression of the average contribution is:
[0091]
[0092] Among them, G irepresents the average contribution of the i-th dimension feature, and K represents a normalization constant, which is used to scale or average the calculation results to ensure that the contribution calculation is within a reasonable range. Indicates that all feature dimensions after the i-th dimension are operated. It means to operate on all feature dimensions before the i-th dimension, p(i,j) means the recognition accuracy when the i-th to j-th dimension features are used as speech feature parameters, p(i+1,j) means the recognition accuracy when the feature dimension increases from i to i+1, and p(i,j+1) means the recognition accuracy when the feature dimension increases from j to j+1.
[0093] In specific implementation, the increase and decrease component method is used to screen the key speech features in MFCC and emotional features, analyze the change trend of feature parameters, identify the feature components that have the greatest impact on synthetic speech detection, and remove unnecessary dimensional components. The average contribution function of the increase and decrease component method is as follows:
[0094]
[0095] The filtered features are fused. Considering that simple linear addition cannot fully exert the anti-noise ability of the two features, we use MFCC and emotional features as the horizontal and vertical axes of the new speech feature space, respectively, and ensure that the frame length and frame shift of the two are consistent, and then perform matrix multiplication on the two features to obtain the fused features.
[0096] Step 6: Input the fused features into the Res2Net network for synthetic speech detection;
[0097] In the specific implementation, the Res2Net network is used to detect real speech and synthetic speech. Its core innovation is the introduction of "shortcut connection" or "skip connection", which enables the network to learn the identity mapping, thereby effectively solving the gradient disappearance and gradient explosion problems in deep network training. There is an identity mapping relationship between the input and output of the residual module, and the formula is as follows:
[0098] F(x)=H(x)+x
[0099] Among them, F(x)F(x) represents the output of the residual module, H(x)H(x) represents the output of the convolutional layer, and xx represents the input.
[0100] Residual blocks and Res2Net models are as follows Figure 5 , Figure 6 shown.
[0101] The Res2Net network extracts multi-scale features by adding multiple receptive fields. A smaller filter group is used to replace the 3*3 filter in the residual block. After passing the 1*1 convolution, the channels are grouped and the features are evenly divided into feature subsets along the channel direction. The feature subsets have the same spatial size as the original features. Except for the first feature subset, the remaining feature subsets will undergo a 3*3 convolution; except for the second feature subset, the remaining feature subsets will undergo a 3*3 convolution; and so on. The process can be expressed as:
[0102]
[0103] Finally, 1*1 convolution is used to fuse feature information of different scales to obtain feature information with different receptive field combinations.
[0104] Step 7, add a gradient reversal layer to the classification module, and classify the detected synthetic speech based on it, and perform adversarial training to obtain a target model to detect the speech to be detected.
[0105] In specific implementation, for the classification of detected synthetic speech, a gradient reversal layer (GRL) is added before the forged type classification module. The gradient reversal layer behaves like an identity operation (i.e., directly passing the input to the output) during forward propagation, but during back propagation, it multiplies the gradient by a negative number (usually -1), and the direction of the gradient is reversed, so that the training objectives of the feature extraction module (generator) and the forged type classification module (discriminator) are opposite, forming adversarial training. The feature extraction module learns common forged features from different forged types of speech. If the features learned by the feature extraction module only contain forged information possessed by certain specific forged types of speech, the feature extraction module may not be able to distinguish the authenticity of unknown deception attacks. Therefore, by helping the feature extraction module learn common forged features from different forged types of speech, the forged type classification module can benefit the feature extraction module.
[0106] The synthetic speech detection method based on emotion feature fusion and gradient adversarial training provided in this embodiment performs synthetic speech detection by adopting the method of feature fusion of time-frequency features and emotion features, and adds a gradient reversal layer for adversarial training to optimize the final classification result, thereby solving the defects of a single speech feature for synthetic speech detection and improving detection efficiency, accuracy and adaptability.
[0107] It should be understood that various parts of the present disclosure may be implemented in hardware, software, firmware, or a combination thereof.
[0108] The above is only a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present disclosure should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A synthetic speech detection method based on emotion feature fusion and gradient adversarial training, characterized in that: include: Step 1, obtaining an audio data set and preprocessing it to obtain a training set, wherein the audio data set includes real speech data and synthesized speech data; Step 2: Add noise with different signal-to-noise ratios to the training set for data enhancement; Step 3, the training set after data enhancement is subjected to fast Fourier transform and then to Mel filter to obtain Mel spectrum and obtain time-frequency features based on it; Step 4, inputting the Mel spectrum into the speech emotion recognition module to obtain the emotion feature; Step 5, using the increase and decrease component method to select key speech features in the time-frequency features and the emotional features and fuse the selected features to obtain fused features; Step 6: Input the fused features into the Res2Net network for synthetic speech detection; Step 7, add a gradient reversal layer to the classification module, and classify the detected synthetic speech based on it, and perform adversarial training to obtain a target model to detect the speech to be detected.
2. The method according to claim 1, characterized in that The step 1 specifically includes: Step 1.1, pre-emphasize the high frequency part of the audio data set; In step 1.2, the pre-emphasized audio data set is divided into frames and a Hamming window is applied to each frame.
3. The method according to claim 2, characterized in that The expression of the Hamming window is Among them, n is the number of sample points and N is the length of the window.
4. The method according to claim 3, characterized in that The step 3 specifically includes: Step 3.1, the training set after data enhancement is subjected to fast Fourier transform to obtain the frequency domain signal; Step 3.2, input the frequency domain signal into a set of Mel filters, take the logarithm of the output of each Mel filter, and obtain the logarithmic energy of the Mel spectrum; Step 3.3, perform discrete cosine transform on the logarithmic Mel frequency spectrum to obtain Mel frequency cepstral coefficients, and retain a preset number of Mel frequency cepstral coefficients as time-frequency features.
5. The method according to claim 4, characterized in that The step 4 specifically includes: Step 4.1, perform convolution operation on the logarithmic Mel spectrum to obtain the convolution sequence and input it into the bidirectional long short-term memory neural network for time summarization; Step 4.2: The output h of the bidirectional long short-term memory neural network is t Normalized through the attention layer to obtain the importance weight a t Among them, t represents the time step, T represents the total length of the sequence, that is, the total number of time steps, and W represents the weight matrix; In step 4.3, the output of the bidirectional long short-term memory neural network is weighted and summed according to the importance weights to obtain the discourse-level representation and use it as the sentiment feature.
6. The method according to claim 5, characterized in that The step 5 specifically includes: Step 5.1, using the increase and decrease component method to calculate the average contribution of the feature components in the time-frequency feature and the emotional feature, and based on this, identify the feature component that has the greatest impact on the synthetic speech detection and delete other features; In step 5.2, the filtered time-frequency features and emotional features are used as the horizontal axis and vertical axis of the new speech feature space respectively, and the time-frequency features and emotional features are matrix multiplied to obtain the fusion features while ensuring that the frame length and frame shift of the two are consistent.
7. The method according to claim 6, characterized in that The expression of the average contribution is: Among them, G i represents the average contribution of the i-th dimension feature, and K represents a normalization constant, which is used to scale or average the calculation results to ensure that the contribution calculation is within a reasonable range. Indicates that all feature dimensions after the i-th dimension are operated. It means to operate on all feature dimensions before the i-th dimension, p(i,j) means the recognition accuracy when the i-th to j-th dimension features are used as speech feature parameters, p(i+1,j) means the recognition accuracy when the feature dimension increases from i to i+1, and p(i,j+1) means the recognition accuracy when the feature dimension increases from j to j+1.