Synthetic speech detection method and device based on background noise, and storage medium
By separating the semantics, rhythms and speaker characteristics in the speech signal, and only the background noise information is retained for synthetic speech detection, the problem of privacy leakage in traditional methods is solved, and efficient and accurate synthetic speech detection is achieved.
Patent Information
- Application Number
- CN202510726033.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Existing synthetic voice detection technologies are difficult to balance between privacy protection and user feature protection. Traditional methods may disclose user identities, affecting detection efficiency and accuracy.
The composite features are extracted by an encoder, and semantic features, pronunciation features and speaker information are separated by residual vector quantization models (RVQs) and the 3rd generation model of natural language, and only background noise information is retained for detection.
It improves the efficiency and accuracy of synthetic voice detection, ensures the balance of privacy protection and data security, adapts to different recording environments, and reduces the risk of counterattack attacks of fake voice detection.
Smart Images

Figure CN120260612A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and more specifically, relates to a synthetic speech detection method, device, and storage medium based on background noise. Background Art
[0002] Deep learning is an important branch in the field of artificial intelligence. The core idea of this technology is to simulate human intelligence, interpret data, classify data, discover potential rules, etc. through multi-layer neural networks. Deep learning technology can automatically learn complex patterns and features from large-scale data and has been widely applied in multiple fields, such as the field of computer images, natural language processing, autonomous driving, medical image analysis, speech recognition, etc. In synthetic speech detection, deep learning is used to analyze and identify subtle differences in audio. These differences may be difficult for humans to distinguish, but can indicate whether the audio is synthesized or tampered with, thus preventing forged content from being used for improper purposes and avoiding posing a security threat to society and individuals.
[0003] In the research of existing synthetic speech detection technologies, it usually focuses on the analysis of audio waveforms or spectral features, relying on extracting low-level features or high-level speech attributes from speech signals and requiring access to complete speech information. For example, a detection scheme based on deep ResNet proposed by Alzantot et al. performs score fusion for three different front-end features, namely MFCC, spectrogram, and CQCC, and trains a model to identify subtle features of synthetic speech for deepfake detection. Although these detection technologies are efficient and practical in preventing audio deepfakes, there may be a risk of exposing private speech content, especially in scenarios involving user privacy such as trade secrets and medical conditions, thus hindering the application of deepfake detection technologies. Some detection technologies attempt to develop an audio deepfake detection framework for content privacy protection. Xinfeng Li et al. proposed a new framework SafeEar for detecting without accessing audio content. It designs a neural audio codec, focuses on extracting speaker-related features, and avoids exposing content privacy by only analyzing the acoustic information of speech without relying on semantic content.
[0004] However, these detection technologies for protecting audio content privacy may still leak user identities due to individual physiological characteristics such as accents, intonations, emotions, and timbres reflected in the acoustic information, affecting the balance between efficient detection and privacy protection and making it difficult to meet the growing detection requirements. Summary of the Invention
[0005] Aiming at the deficiencies of the related technologies, the purpose of the present invention is to provide a synthetic speech detection method, device, and storage medium based on background noise, aiming to solve the technical problem that the existing synthetic speech detection methods lack content privacy protection and user feature protection for the detected audio.
[0006] To achieve the above object, the present invention provides a synthetic speech detection method based on background noise, including: S1. An encoder is used to extract a sample encoding containing composite features from the original audio, and the composite features include semantic features and acoustic features; S2. The sample encoding is input into a residual vector quantization model RVQs for feature extraction. A latent model is used to guide the first quantization output VQ1 of RVQs to extract the semantic features in the sample encoding, and the semantic features are stripped from the composite features through the residual structure of RVQs to obtain pure acoustic features; S3. A natural language 3-generation model is used to guide the second quantization output VQ2 and the third quantization output VQ3 of RVQs to extract the prosody features and speaker information in the pure acoustic features respectively, and the prosody features and speaker information are stripped from the pure acoustic features through the residual structure of RVQs to obtain pure background noise information; S4. The background noise information is input into a detection model for speech detection to determine whether the original audio is synthetic audio.
[0007] Optionally, the encoder includes a one-dimensional convolutional layer, a residual block, a downsampling layer, an LSTM layer, and a convolutional layer; The one-dimensional convolutional layer is used to extract features from the original audio signal, convert the original audio signal into an array form, and obtain preliminary features; The residual block, downsampling layer, and LSTM layer are used to extract information across the audio feature space from the preliminary features and output composite features including semantic features and acoustic features; The convolutional layer is used to convert the dimension of the composite features into the dimension of the quantization input RVQ to obtain a sample encoding containing composite features.
[0008] Optionally, step S2 specifically includes: S2.1. Map the output dimension of VQ1 to the same dimension as the semantic features extracted by the latent model; S2.2. Use the average representation of all latent model layers as a semantic supervision signal, and supervise VQ1 to learn the content representation in the sample encoding through a loss function, extract semantic features and obtain corresponding semantic tags; the loss function is:
[0009] Wherein, represents the quantization output of the VQ layer, represents the time step of the supervision signal, represents the cosine similarity, represents the sigmoid activation function, represents the projection matrix; S2.3. Use the residual structure of RVQs to subtract the semantic features corresponding to the semantic tokens from the composite features to obtain pure acoustic features.
[0010] Optionally, step S3 specifically includes: S3.1. Map the output dimension of VQ2 to the same dimension as the prosody features extracted by the natural language 3-generation model, and map the output dimension of VQ3 to the same dimension as the speaker features extracted by the natural language 3-generation model; S3.2. Use the natural language 3-generation model to extract the prosody features and speaker information in the pure acoustic features from the second quantization output VQ2 and the third-layer quantization output VQ3 respectively, and perform supervised learning on VQ2 and VQ3 through a loss function to obtain prosody tokens P and speaker tokens ; S3.3. Use the residual structure of RVQs to strip the prosody features and speaker information corresponding to the prosody tokens P and speaker tokens from the pure acoustic features to obtain pure background noise information.
[0011] Optionally, after step S3, it further includes: Transmit the pure background noise information to the remaining five quantizers of RVQs in sequence, and each quantizer refines the background noise information in sequence to enhance its feature representation.
[0012] Optionally, the detection model includes a bottleneck layer, a Transformer classifier, and an output layer; The bottleneck layer is used to reduce the dimension of the obtained background noise features through one-dimensional convolution and batch normalization; The Transformer classifier is used to pass the compressed background noise features through multiple attention heads to capture different long-range dependencies and dynamic spatial weighting information, and aggregate the output features of each attention head to obtain a unified attention spectrum; The output layer is used to compare the attention spectrum with the corresponding attention spectra of the background noise of the learned real audio and the corresponding attention spectra of the background noise of the synthetic audio to determine whether the original audio is synthetic audio.
[0013] In a second aspect, the present invention also provides a synthetic speech detection device based on background noise for performing the synthetic speech detection method according to any one of the first aspects, including: A data preprocessing module for extracting sample encodings containing composite features from the original audio by using an encoder, where the composite features include semantic features and acoustic features; An acoustic feature extraction module, which is used to input the sample encoding into a residual vector quantization model (RVQs) for feature extraction. The first quantization output VQ1 of the RVQs is guided by a latent element model to extract semantic features from the sample encoding, and the semantic features are stripped from the composite features through the residual structure of the RVQs to obtain pure acoustic features; A background feature extraction module, which is used to guide the second quantization output VQ2 and the third quantization output VQ3 of the RVQs by a natural language 3-generation model to extract prosody features and speaker information from the pure acoustic features respectively, and the prosody features and speaker information are stripped from the pure acoustic features through the residual structure of the RVQs to obtain pure background noise information; A detection module, which is used to input the background noise information into a detection model for speech detection to determine whether the original audio is synthetic audio.
[0014] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the synthetic speech detection method described in any item of the first aspect is implemented.
[0015] Through the above technical solutions conceived by the present invention, compared with the prior art, the following beneficial effects can be achieved: 1. The embodiment of the present invention provides a synthetic speech detection method based on background noise. By only analyzing background noise information for synthetic speech detection, it is determined whether the audio is deepfake. Since background noise does not contain semantic content such as context and acoustic information such as accent, intonation, emotion, and timbre, although privacy attackers can receive a series of background markers, the lack of semantic, prosody, and speaker timbre clues hinders their ability to restore understandable content and infer user identity characteristics. In the feature extraction stage, different features are extracted and stripped layer by layer to promote the hierarchical decoupling of different information; the semantic features and acoustic features are decoupled, and further the prosody, speaker features, and background noise are decoupled in the acoustic features, gradually stripping the semantic features, prosody, and speaker features, and extracting purer background noise information for deepfake detection, improving the detection efficiency and accuracy. Compared with traditional detection methods, the method of the present invention realizes an efficient and accurate synthetic speech detection method, and also ensures the balance between privacy protection and data security, providing a new technical path for the application in the field of speech deepfake detection.
[0016] 2. The embodiments of the present invention provide a method for detecting synthetic speech based on background noise, which can also improve the adaptability to different types of forged speech. Traditional methods usually rely on the acoustic features of the speech itself, while the present invention separates the noise from the speech features, avoiding the adversarial attacks of these forgery methods and significantly enhancing the detection ability of synthetic speech. In addition, it also has a certain degree of flexibility and can adapt to the requirements of different devices or application environments by further optimizing the background noise information. For the speech collected in different recording environments, the characteristics of the background noise will be different. Therefore, the method has good adaptability and can process various changing background noises by appropriately adjusting parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic flowchart of a method for detecting synthetic speech based on background noise provided by an embodiment of the present invention.
[0018] Figure 2 is a framework diagram of a method for detecting synthetic speech based on background noise proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0020] The following describes the content involved in the above embodiments in combination with a preferred embodiment.
[0021] Embodiment 1 The present invention provides a method for detecting synthetic speech based on background noise, including: S1. An encoder is used to extract a sample code containing composite features from the original audio, and the composite features include semantic features and acoustic features; S2. The sample code is input into a residual vector quantization model RVQs for feature extraction. A latent model is used to guide the first quantization output VQ1 of RVQs to extract the semantic features in the sample code, and the semantic features are stripped from the composite features through the residual structure of RVQs to obtain pure acoustic features; S3. A natural language 3-generation model is used to guide the second quantization output VQ2 and the third quantization output VQ3 of RVQs to extract the prosody features and speaker information in the pure acoustic features respectively, and the prosody features and speaker information are stripped from the pure acoustic features through the residual structure of RVQs to obtain pure background noise information; S4. Input the background noise information into a detection model for voice detection to determine whether the original audio is synthetic audio.
[0022] Reference Figure 2 , in one embodiment, the constructed model for synthetic speech detection based on background noise includes an encoder, a feature extraction module, and a detector module; the encoder is used to extract sample encodings containing composite features from the original audio; the feature extraction module includes a Residual Vector Quantization model (RVQs), a latent variable model, and a Natural Language 3rd generation model; the latent variable model guides the first quantization output VQ1 of the RVQs to extract semantic features from the sample encodings, and the Natural Language 3rd generation model guides the second quantization output VQ2 and the third quantization output VQ3 of the RVQs to extract prosodic features and speaker information from the clean acoustic features respectively, so that after the sample encodings pass through VQ1, VQ2, and VQ3 in sequence, the semantic features, prosodic features, and speaker information are stripped to obtain clean background noise information; the detection model of the detector module includes a bottleneck layer, a Transformer classifier, and an output layer. The bottleneck layer reduces the dimension of the clean background noise information, the Transformer classifier aggregates the output features of each attention head to obtain a unified attention spectrum, and the output layer judges it based on the feature difference between the background noise of the real audio and the synthetic audio, so as to determine whether the original audio is synthetic audio. If it is synthetic audio, the output result is "forged", and if it is real audio, the output result is "real".
[0023] The present invention cleverly strips other feature information in the speech, thus using the background noise as the key clue for detecting synthetic speech. The composite features are extracted by the encoder and gradually separated using the RVQs model, and finally the background noise information is obtained and detected based on this. This solution is different from traditional speech recognition methods, focusing on the background noise features in the sound that are not easily forged, and can effectively improve the detection accuracy of synthetic speech.
[0024] Optionally, the encoder includes a one-dimensional convolutional layer, a residual block, a downsampling layer, an LSTM layer, and a convolutional layer; The one-dimensional convolutional layer is used to extract features from the original audio signal, convert the original audio signal into an array form, and obtain preliminary features; The residual block, downsampling layer, and LSTM layer are used to extract information across the audio feature space from the preliminary features and output composite features containing semantic features and acoustic features; The convolutional layer is used to convert the dimension of the composite features into the dimension of the quantization input RVQ to obtain sample encodings containing composite features.
[0025] In the embodiment of the present invention, the dimension of the given original audio signal X is 1, the convolution kernel size of the one-dimensional convolution layer is 7, and the preliminary feature representation E is obtained after passing through the convolution layer. The preliminary feature representation E is input into four residual blocks, and the convolution kernel sizes in the four residual blocks are 8, 5, 4, and 2 respectively. Then, through the downsampling layer and the bidirectional LSTM (Bi-LSTM), the information in the cross-audio feature space is extracted to obtain the rich composite feature E'; the composite feature E' includes semantic features and acoustic features. The composite feature E' is passed through a convolution layer, and the exponential linear unit (ELU) and layer normalization are used in the convolution layer to enhance the non-linear representation and the stability of the model, and the dimension of the feature E' is matched to be consistent with the RVQ input dimension, which is 256.
[0026] Optionally, step S2 specifically includes: S2.1. Map the output dimension of VQ1 to the same dimension as the semantic features extracted by the latent model. S2.2. Use the average representation of all latent model layers as the semantic supervision signal, and supervise VQ1 to learn the content representation in the sample encoding through the loss function, extract semantic features and obtain corresponding semantic tags; the loss function is:
[0027] Wherein, represents the quantization output of the VQ layer, represents the time step of the supervision signal, represents the cosine similarity, represents the sigmoid activation function, represents the projection matrix; S2.3. Use the residual structure of RVQs to subtract the semantic features corresponding to the semantic tags from the composite features to obtain pure acoustic features.
[0028] The encoded feature E' is input into the quantizer RVQs for quantization. The feature dimension of the first-layer quantization output VQ1 is 1024, and it is mapped to the latent model feature dimension 768 through a linear layer to achieve feature dimension matching.
[0029] Use the semantic features F1 extracted by the latent model layer as the semantic supervision signal, and supervise the content representation learning of VQ1 by calculating the distillation loss to obtain the semantic tag of VQ1. The loss function is:
[0030] Wherein, represents the quantization output of the VQ layer, Indicates the time step of the supervision signal, represents the cosine similarity, represents the sigmoid activation function, represents the projection matrix.
[0031] Using the residual structure of RVQs, decouple the semantic features and acoustic features, and during the backpropagation process, adopt the gradient reversal technique to change the gradient direction by multiplying the acoustic feature gradient by -alpha, suppress the learning of acoustic features, and enhance the semantic labels of the extraction effect. Since RVQs have a residual structure, the semantic labels output by VQ1 corresponding semantic features can be stripped from the composite representation to obtain the pure acoustic information A without semantic information, realizing the decoupling of semantic features and pure acoustic features.
[0032] Optionally, step S3 specifically includes: S3.1. Map the output dimension of VQ2 to the same dimension as the prosody features extracted by the natural language 3-generation model, and map the output dimension of VQ3 to the same dimension as the speaker features extracted by the natural language 3-generation model; S3.2. Use the natural language 3-generation model to extract the prosody features and speaker information in the pure acoustic features from the second quantization output VQ2 and the third quantization output VQ3 respectively, and perform supervised learning on VQ2 and VQ3 through the loss function to obtain the prosody label P and the speaker label ; S3.3. Through the residual structure of RVQs, strip the prosody features and speaker information corresponding to the prosody label P and the speaker label from the pure acoustic features to obtain pure background noise information.
[0033] The output feature dimensions of VQ2 and VQ3 are 1024. Using a linear layer, map their dimensions to 256, which is consistent with the prosody feature and speaker feature dimensions of the natural language 3-generation model.
[0034] Use the prosody features F2 extracted by the natural language 3-generation model as the prosody supervision signal, and perform prosody representation supervised learning on VQ2 by calculating the distillation loss to obtain the prosody label P of VQ2; use the speaker features F3 extracted by the natural language 3-generation model as the speaker supervision signal, and perform speaker representation supervised learning on VQ3 by calculating the distillation loss to obtain the speaker label of VQ3 Further, during the backpropagation process, the gradient reversal technique is adopted to suppress the learning of other features respectively and enhance the extraction effects of the prosody marker P and the speaker marker. Based on the second-layer quantization output and the third-layer quantization output, the extraction effects further decouple the prosody, speaker features, and background noise in the pure acoustic features; based on the pure acoustic information A of the input VQ2, the prosody marker P of VQ2 and the speaker marker of VQ3 are stripped therefrom to obtain the pure background noise information B.
[0035] Optionally, after step S3, it further includes: The pure background noise information is sequentially passed to the remaining five quantizers of the RVQs, and each quantizer sequentially refines the background noise information to enhance its feature representation.
[0036] As Figure 2 shown, the obtained background noise information B is passed to the next five quantizers VQ4 to VQ8 to further refine the background noise information.
[0037] Optionally, the detection model includes a bottleneck layer, a Transformer classifier, and an output layer; The bottleneck layer is used to reduce the dimension of the obtained background noise features through one-dimensional convolution and batch normalization; The Transformer classifier is used to pass the compressed background noise features through multiple attention heads to capture different long-distance dependence relationships and dynamic spatial weighting information, and aggregate the output features of each attention head to obtain a unified attention spectrum; The output layer is used to compare the attention spectrum with the corresponding attention spectra of the background noise of the learned real audio and the corresponding attention spectra of the background noise of the synthetic audio to determine whether the original audio is a synthetic audio.
[0038] The extracted background noise feature B is passed through a bottleneck layer consisting of 5 convolutional layers with a kernel size of 1 and 1 normalization layer, reducing its dimension from 5 to 1, which facilitates the subsequent layers to operate on the compact representation B1 and avoids overfitting at the same time. The compressed background noise feature B1 is fed into a Transformer-based classifier. Through the best 8 attention heads, the model can more effectively participate in long-range feature interaction and dynamic spatial weighting, and parallelly process different aspects of the input feature space. By using sine position encoding, the output features of each attention head are aggregated to obtain the attention spectrum B1'. The aggregated attention spectrum B1' is input into the output layers of two groups of Transformer classifiers to globally process the sequence and capture global dependencies. Each group of encoders contains two feed-forward networks, multi-head self-attention, and layer normalization modules. Finally, a classification score score is output through a fully connected layer. The score score represents the similarity between the attention spectrum of the detected audio and the learned attention spectra of real audio and synthetic audio. Combining with the best threshold threshold, the output is the class that the detected audio is more similar to among the two as the classification result (using the threshold threshold to classify the output score score, classifying samples with a score higher than the threshold as synthetic audio and results lower than the threshold as real audio) to determine whether the audio is deepfake. Through experiments, it is proved that setting two groups of Transformer classifiers in this embodiment has the best effect. Using two independent Transformer classifiers can avoid a single classifier relying too much on a certain type of feature, thereby reducing the risk of overfitting. Even if the learning ability of one group of classifiers is affected by certain errors, the other group can make up for its deficiencies.
[0039] The prototype code of the method of the present invention is implemented in Python language, and feasibility and effectiveness experimental verifications are carried out on the ASVspoof 2019 corpus and the ASVspoof2021 corpus, involving the detection of samples generated by various means such as speech synthesis, voice conversion, and replay. Among them, the detection accuracy rate on the ASVspoof 2019 dataset reaches 89.68%, the detection equal error rate (eer) is as low as 0.054, and the tandem detection cost function (t-DCF) is as low as 0.159, indicating that the model has good classification performance and high detection efficiency, verifying the feasibility and effectiveness of the solution of the present invention.
[0040] The solution of the present invention promotes the hierarchical decoupling of different information by extracting and stripping different features layer by layer in the feature extraction stage; decouples semantic features and pure acoustic features, further decouples prosody, speaker features and background noise in the pure acoustic features, gradually strips semantic features, prosody and speaker features, and extracts purer background noise information for deepfake detection; performs synthetic speech detection by only analyzing background noise information, determines whether the audio is a deepfake, and improves the detection efficiency and accuracy. At the same time, since the detected audio does not contain human voice information, it hinders the ability to restore understandable content and infer user identity features, ensuring the balance of privacy protection and data security, and providing a new technical path for the application in the field of speech deepfake detection.
[0041] Embodiment 2 The present invention also provides a synthetic speech detection device based on background noise for performing the synthetic speech detection method according to any one of Embodiment 1, including: A data preprocessing module for extracting a sample encoding containing composite features from the original audio by using an encoder, where the composite features include semantic features and acoustic features; An acoustic feature extraction module for inputting the sample encoding into a residual vector quantization model RVQs for feature extraction, using a latent element model to guide the first quantization output VQ1 of RVQs to extract the semantic features in the sample encoding, and stripping the semantic features from the composite features through the residual structure of RVQs to obtain pure acoustic features; A background feature extraction module for using a NaturalSpeech3 model to guide the second quantization output VQ2 and the third quantization output VQ3 of RVQs to extract the prosody features and speaker information in the pure acoustic features respectively, and stripping the prosody features and speaker information from the pure acoustic features through the residual structure of RVQs to obtain pure background noise information; A detection module for inputting the background noise information into a detection model for speech detection to determine whether the original audio is a synthetic audio.
[0042] A synthetic speech detection device based on background noise provided by an embodiment of the present invention is used to perform a synthetic speech detection method provided by any embodiment of the present invention, and has corresponding beneficial effects.
[0043] Embodiment 3 The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the synthetic speech detection method according to any one of Embodiment 1.
[0044] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting synthetic speech based on background noise, characterized in that, Including: S1. An encoder is used to extract sample encodings containing composite features from the original audio, and the composite features include semantic features and acoustic features; S2. The sample encodings are input into a Residual Vector Quantization model RVQs for feature extraction. A latent model is used to guide the first-layer quantization output VQ1 of the RVQs to extract the semantic features in the sample encodings, and the semantic features are stripped from the composite features through the residual structure of the RVQs to obtain pure acoustic features; S3. A natural language 3-generation model is used to guide the second quantization output VQ2 and the third-layer quantization output VQ3 of the RVQs to extract the prosody features and speaker information in the pure acoustic features respectively, and the prosody features and speaker information are stripped from the pure acoustic features through the residual structure of the RVQs to obtain pure background noise information; S4. The background noise information is input into a detection model for speech detection to determine whether the original audio is synthetic audio.
2. The synthetic speech detection method according to claim 1, wherein The encoder includes a one-dimensional convolutional layer, a residual block, a downsampling layer, an LSTM layer, and a convolutional layer; The one-dimensional convolutional layer is used to extract features from the original audio signal, convert the original audio signal into an array form, and obtain preliminary features; The residual block, downsampling layer, and LSTM layer are used to extract information across the audio feature space from the preliminary features and output composite features containing semantic features and acoustic features; The convolutional layer is used to convert the dimension of the composite features into the dimension of the quantization input RVQ to obtain sample encodings containing composite features.
3. The synthetic speech detection method according to claim 1, wherein, Step S2 specifically includes: S2.
1. Map the output dimension of VQ1 to the same dimension as the semantic features extracted by the latent model; S2.
2. Use the average representation of all latent model layers as the semantic supervision signal, and supervise VQ1 to learn the content representation in the sample encodings through a loss function, extract semantic features and obtain corresponding semantic tokens; the loss function is: Among them, represents the quantization output of the VQ layer, represents the time step of the supervision signal, represents the cosine similarity, represents the sigmoid activation function, represents the projection matrix; S2.
3. Use the residual structure of the RVQs to subtract the semantic features corresponding to the semantic tokens from the composite features to obtain pure acoustic features.
4. The synthetic speech detection method according to claim 1, wherein Step S3 specifically includes: S3.
1. Map the output dimension of VQ2 to the same dimension as the prosody features extracted by the natural language 3-generation model, and map the output dimension of VQ3 to the same dimension as the speaker features extracted by the natural language 3-generation model; S3.
2. Use the natural language 3rd generation model to extract the prosody features and speaker information in the pure acoustic features from the second quantization output VQ2 and the third quantization output VQ3 respectively, and perform supervised learning on VQ2 and VQ3 through a loss function to obtain the prosody label P and the speaker label ; S3.
3. Stripping the corresponding prosodic features and speaker information of the prosody markers P and speaker markers from the clean acoustic features through the residual structure of RVQs to obtain clean background noise information. 5. A method for detecting synthetic speech based on background noise according to claim 1, characterized in that, After step S3, it further includes: The pure background noise information is sequentially passed to the remaining five quantizers of the RVQs, and each quantizer sequentially refines the background noise information to enhance its feature representation.
6. The synthetic speech detection method based on background noise according to claim 1, characterized in that The detection model includes a bottleneck layer, a Transformer classifier, and an output layer; The bottleneck layer is used to reduce the dimension of the obtained background noise features through one-dimensional convolution and batch normalization; The Transformer classifier is used to pass the compressed background noise features through multiple attention heads, capture different long-range dependency relationships and dynamic spatial weighting information, and aggregate the output features of each attention head to obtain a unified attention spectrum; The output layer is used to compare the attention spectrum with the corresponding attention spectra of the background noise of the learned real audio and the corresponding attention spectra of the background noise of the synthesized audio, and determine whether the original audio is synthesized audio.
7. A synthetic speech detection device based on background noise, which is used to execute the synthetic speech detection method according to any one of claims 1-6, characterized in that It includes: A data preprocessing module, which is used to extract sample encodings containing composite features from the original audio by using an encoder, and the composite features include semantic features and acoustic features; An acoustic feature extraction module, which is used to input the sample encoding into a residual vector quantization model RVQs for feature extraction, use a HuBERT model to guide the first quantization output VQ1 of RVQs to extract the semantic features in the sample encoding, and strip the semantic features from the composite features through the residual structure of RVQs to obtain pure acoustic features; A background feature extraction module, which is used to use a natural language 3-generation model to guide the second quantization output VQ2 and the third quantization output VQ3 of RVQs to extract the prosody features and speaker information in the pure acoustic features respectively, and strip the prosody features and speaker information from the pure acoustic features through the residual structure of RVQs to obtain pure background noise information; A detection module, which is used to input the background noise information into a detection model for speech detection to determine whether the original audio is synthesized audio.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the synthetic speech detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice content and tone separation method
CN118298802A
Deepfake detection
US20240363103A1
Generative speech model for compact data-driven speech vectors for versatile speech applications
WO2025085325A1
Cited By
Synthetic audio detection method and device, electronic equipment and medium
CN122369511A
A method, apparatus, electronic device and medium for synthesized audio detection
CN122369511B