Method and system for demixing forged voiceprints
Through the feature extraction and residual orthogonalization method based on the Transformer model, combined with the additive angle margin loss, the interference of forged voiceprints after speech conversion is solved, the voiceprint characteristics of the source speaker are restored, and the anti-forgery ability and recognition accuracy of the voice verification system are improved.
Patent Information
- Application Number
- CN202510404867.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing voice verification methods cannot effectively deal with the forged voiceprints after speech conversion, especially when facing complex speech synthesis techniques and conversion models, the real voiceprint characteristics of the source speaker cannot be restored, resulting in insufficient system anti-forgery capabilities.
The feature extraction technology based on the Transformer model is adopted, combined with the residual orthogonalization method and additive angle margin loss, and the voiceprint characteristics of the source speaker are restored and the interference of the fake voiceprint is removed.
It effectively restores the true voiceprint characteristics of the source speaker, improves the robustness and anti-forgery capabilities of the voice verification system, and enhances the accuracy and scalability of voiceprint recognition.
Smart Images

Figure CN120260575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of speech processing and deep learning, and more specifically, to a method and system for separating forged voiceprints Background Art
[0002] With the development of speech synthesis and conversion technologies, the problem of voice forgery has become increasingly serious. Especially in voiceprint recognition systems, voice conversion technologies can change the speech features of audio to make it sound like the voice of a target speaker, which makes voiceprint verification and recognition systems vulnerable to attacks. Traditional voice verification methods cannot effectively handle forged voiceprints after voice conversion, especially when facing complex speech synthesis technologies and conversion models. To address this challenge, a new method is needed to restore the true voiceprint of the source speaker to prevent maliciously forged audio from evading detection.
[0003] Most of the existing technologies focus on feature extraction and reconstruction during the voice conversion process, but lack effective solutions for restoring the voiceprint of the source speaker from voice-converted audio. Therefore, how to effectively remove the influence of the target speaker's voiceprint after voice conversion and restore the true voiceprint features of the source speaker, so as to improve the anti-forgery ability of voice verification systems, is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0004] In view of this, the present invention provides a method and system for separating forged voiceprints, which solves the problems existing in the background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for separating forged voiceprints, comprising the following steps:
[0007] Performing feature extraction on the input forged speech based on a Transformer model to obtain rough features containing the voiceprint information of the source speaker;
[0008] Decomposing the rough features using a residual orthogonalization method to restore the voiceprint features of the source speaker;
[0009] Performing dimensional normalization on the voiceprint features of the source speaker to obtain voiceprint features of a fixed length;
[0010] Using an additive angular margin loss to enhance the angular difference between the source speaker and other voiceprints, and outputting the audio data after separation.
[0011] Optionally, before feature extraction, the method for separating forged voiceprints further includes:
[0012] Preprocess the input forged speech x(t) using a band-pass filter, and obtain the discretized signal according to the following formula:
[0013] x[n] = x(nT s )
[0014] where F s is the sampling frequency and n is the index of the sampling point;
[0015] Convert the forged speech from the time domain to the frequency domain through the short-time Fourier transform:
[0016]
[0017] In the formula: X(f,t) is the time-frequency representation in the frequency domain; w[n-t] is the window function, and the Hamming window is used to smooth the signal.
[0018] Optionally, perform feature extraction on the input forged speech based on the Transformer model, specifically:
[0019] Convert the forged speech into Mel-frequency features, and use the multi-head self-attention mechanism, positional encoding, and adaptive input length technology to process the Mel-frequency features to capture the time-domain features and frequency-domain features.
[0020] Optionally, decompose the rough features using the residual orthogonalization method, specifically
[0021] Use the residual orthogonalization method to decompose the rough features into components parallel to the target speaker's voiceprint and orthogonal components, remove the parallel components related to the target speaker's voiceprint, and restore the source speaker's voiceprint features.
[0022] Optionally, perform dimensional normalization on the source speaker's voiceprint features, specifically:
[0023] Combine the global statistical pooling and adaptive normalization methods to normalize the source speaker's voiceprint features into a fixed-length output.
[0024] Optionally, the expression of the additive angular margin loss is:
[0025] L AAM = max(0, cos(θ target ) - cos(θ other ) + δ)
[0026] In the formula: θ tar get represents the angle between the source speaker and the target voiceprint, θ other represents the angle between the source speaker and other voiceprints, and δ is the additive margin.
[0027] A forged voiceprint unmixing system, comprising a feature extraction module, a difference correction module, a dimension normalization module, and a voiceprint enhancement module connected in sequence;
[0028] The feature extraction module is used to extract features from the input forged voice to obtain rough features containing the voiceprint information of the source speaker;
[0029] The difference correction module is used to remove the influence of the target speaker's voiceprint and restore the voiceprint features of the source speaker;
[0030] The dimension normalization module is used to perform dimension normalization on the voiceprint features of the source speaker to obtain voiceprint features of a fixed length;
[0031] The voiceprint enhancement module is used to enhance the angular difference between the source speaker and other voiceprints through additive angular margin loss, and output the unmixed audio data.
[0032] Optionally, the feature extraction module adopts a Transformer model to extract voiceprint features through a multi-head self-attention mechanism, position encoding, and adaptive length processing technology.
[0033] Optionally, the difference correction module decomposes the rough features through a residual orthogonalization method.
[0034] Optionally, the dimension normalization module combines global statistical pooling and adaptive normalization methods to normalize the variable-length audio into a fixed-length output.
[0035] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a forged voiceprint unmixing method and system. By adopting a Transformer-based feature extraction technology and a difference correction method, it can effectively restore the voiceprint features of the source speaker and remove the interference of forged voiceprints; combined with additive angular margin loss, it enhances the difference between the source voiceprint and other voiceprints, improving the accuracy and robustness of the voiceprint recognition system; at the same time, by adopting an adaptive input length processing technology and a global statistical pooling method, it can ensure that the system can process variable-length audio and maintain a stable output. Therefore, the present invention has strong scalability and versatility, can be widely applied to fields such as voice verification and recognition, effectively improves voice security, and reduces the risk of forged voices. Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0037] Figure 1 This is the flowchart of the method for separating forged voiceprints provided by the present invention. Detailed implementation manners
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] In order to restore the voiceprint features of the source speaker from the audio after voice conversion and enhance the security and anti-forgery ability of the voice verification system, an embodiment of the present invention discloses a method for separating forged voiceprints, as Figure 1 shown, including the following steps:
[0040] Extract features from the input forged voice based on the Transformer model to obtain rough features containing the voiceprint information of the source speaker;
[0041] Use the residual orthogonalization method to decompose the rough features to restore the voiceprint features of the source speaker;
[0042] Normalize the dimension of the voiceprint features of the source speaker to obtain voiceprint features with a fixed length;
[0043] Use the additive angular margin loss to enhance the angular difference between the source speaker and other voiceprints, and output the audio data after separation.
[0044] Based on Figure 1 the shown process, this embodiment utilizes advanced technologies such as the Transformer-based feature extraction model, residual orthogonalization difference correction, dimension normalization, and AAM loss, which can effectively remove the influence of the target speaker's voiceprint after voice conversion, restore the true voiceprint features of the source speaker, and improve the robustness and accuracy of the voice verification system when facing forged voices, thereby enhancing the anti-forgery ability of the voice verification system.
[0045] Next, the Figure 1 shown process will be elaborated in detail to further understand the implementation steps of the present invention.
[0046] (1) Sampling and preprocessing
[0047] Preprocess the input forged voice x(t) using a band-pass filter to remove unnecessary low-frequency noise and adjust the audio signal to a sampling frequency of 16 kHz. Obtain the discretized signal according to the following formula:
[0048] x[n] = x(nT s )
[0049] where F s is the sampling frequency and n is the index of the sampling points;
[0050] The forged speech is transformed from the time domain to the frequency domain through the short-time Fourier transform (STFT):
[0051]
[0052] where: X(f,t) is the time-frequency representation in the frequency domain; w[n-t] is the window function, and the Hamming window is used to smooth the signal.
[0053] (II) Feature extraction
[0054] The Transformer model can effectively capture the long-range dependencies in the audio and adapt to the variable-length audio data input. In this embodiment, based on the Transformer model, feature extraction is performed on the input forged speech, and the multi-head self-attention mechanism, positional encoding, and adaptive input length technology are used to process the Mel frequency features to capture the time-domain features and frequency-domain features, specifically including the following steps:
[0055] 1. Filter bank conversion: The forged speech X is converted into Mel frequency features (FBank) to obtain the Mel frequency feature matrix X FBank :
[0056] X FBank = FBank(X), X FBank ∈ R f×T
[0057] where: f is the number of Mel filters and T is the number of generated frames.
[0058] 2. Positional encoding: To ensure that the Transformer can process the temporal relationship of the audio and enhance the transmission of temporal information, temporal information is added to the audio features to obtain the enhanced audio feature matrix X pos :
[0059] X pos = X FBank + PosEnc(X FBank ), X pos ∈ R f×T
[0060] 3. Transformer encoding: The enhanced audio features are processed through the Transformer encoder to output the feature matrix H transformer :
[0061] Htransformer = TransformerEncoder(X pos ), H transformer ∈R d×T
[0062] where d is the dimension of the output feature and T is the time step.
[0063] It can be seen that in order to better simulate the auditory perception of the human ear, in this embodiment, a Mel frequency filter bank is used for frequency conversion to transform the spectrum into the Mel scale. Assuming that the spectrum of the signal is represented as X(f, t), the output after the Mel frequency transformation is:
[0064] M(f) = Mel(X(f, t))
[0065] where: The Mel function converts the frequency to the Mel scale, usually approximated by a combination of linear frequency and logarithmic frequency.
[0066] After obtaining the Mel frequency features, the discrete cosine transform (DCT) is further applied to process them to obtain the Mel frequency cepstral coefficients (MFCC):
[0067]
[0068] where: M is the number of Mel frequency filter banks and k is the index of the cepstral coefficients.
[0069] (III) Difference correction
[0070] The rough feature M is decomposed into the component M k parallel to the target speaker's voiceprint and the orthogonal component M ⊥ by using the residual orthogonalization method (ROB), and the parallel component related to the target speaker's voiceprint is removed to restore the voiceprint features of the source speaker.
[0071]
[0072] where: M is the input audio feature matrix, M k is the part parallel to the target voiceprint, and M ⊥ is the orthogonal part, that is, the source voiceprint restoration part.
[0073] In this embodiment, in order to counter the influence of forged voiceprints, it is necessary to extract and separate the difference features between the target speaker and the forged speech from the audio signal. Specifically, principal component analysis (PCA) is used to decompose the feature space and extract the representative principal components. Let the input feature matrix be X, and through PCA for feature dimensionality reduction, the main principal component P is obtained:
[0074] P = PCA(X)
[0075] In this way, the principal components in the feature space can help distinguish different speakers and forged voiceprints.
[0076] To further correct the influence of forged voiceprints, in this embodiment, the true voiceprint features of the source speaker are extracted by the above orthogonalization method. By performing orthogonal decomposition on the target features and forged features, the forged components are removed and the features of the source speaker are retained. The specific approach is as follows:
[0077] X ⊥ = X - P·P T ·X
[0078] In the formula: X ⊥ is the feature matrix after removing the forged components, and P is the principal component vector in the direction of the target voiceprint.
[0079] The feature matrix X obtained through orthogonalization ⊥ is fed into the reconstruction module to restore the voiceprint of the source speaker. The features of the source speaker are reconstructed through a recurrent neural network to generate the final voiceprint features.
[0080] (4) Dimension normalization
[0081] Perform dimension normalization on the voiceprint features of the source speaker, specifically: combining the global statistical pooling and adaptive normalization methods, normalize the voiceprint features of the source speaker into an output with a fixed length, so as to adapt to subsequent voiceprint verification or recognition tasks.
[0082] The global statistical pooling method maps the features from the time domain or frequency domain space to a fixed vector representation:
[0083]
[0084] In the formula: X feature are the features obtained by feature extraction, are the normalized features.
[0085] It can be seen that in this embodiment, the corrected and reconstructed features are normalized to unify the feature dimensions of all audio signals. Specifically, to achieve this goal, this embodiment uses the global statistical pooling method to perform pooling on the features to make them have a standardized range. For the feature vector x of each frame t , its normalization process is:
[0086]
[0087] In the formula: μ and σ are the mean and standard deviation of the features respectively. In this way, the unity and standardization of the features are ensured.
[0088] To further optimize the normalization process of features, an adaptive normalization method is adopted to adjust the weights of different feature dimensions. The specific method is to train an adaptive network to learn the weights of each feature dimension and perform weighted normalization:
[0089]
[0090] In the formula: W is the weighted coefficient matrix obtained by training, representing the importance of each feature dimension.
[0091] (5) Voiceprint enhancement
[0092] To further enhance the distinguishability between the source speaker and other voiceprints, this embodiment adopts an additive angular margin loss as the voiceprint enhancement loss. The design goal of this loss function is to maximize the angular difference between the source speaker and the forged voiceprint, thereby improving the distinguishability of the voiceprint. The form of this loss function is as follows:
[0093] L AAM = max(0, cos(θ target ) - cos(θ other ) + δ)
[0094] In the formula: θ target represents the angle between the source speaker and the target voiceprint, θ other represents the angle between the source speaker and other voiceprints, and δ is the additive margin, ensuring a large angular difference between the enhanced voiceprint and other voiceprints.
[0095] Corresponding to Figure 1 the method described above, an embodiment of the present invention also provides a forged speech voiceprint unmixing system for Figure 1 the specific implementation of the method in. An embodiment of the forged speech voiceprint unmixing system provided by the present invention can be applied to a computer terminal or various mobile devices, and specifically includes a feature extraction module, a difference correction module, a dimension normalization module, and a voiceprint enhancement module that are connected in sequence;
[0096] The feature extraction module is used to extract features from the input forged speech to obtain rough features containing the voiceprint information of the source speaker;
[0097] The difference correction module is used to remove the influence of the target speaker's voiceprint and restore the voiceprint features of the source speaker;
[0098] The dimension normalization module is used to perform dimension normalization on the voiceprint features of the source speaker to obtain voiceprint features of a fixed length;
[0099] The voiceprint enhancement module is used to enhance the angular difference between the source speaker and other voiceprints through the additive angular margin loss and output the unmixed audio data.
[0100] Furthermore, the feature extraction module uses a Transformer model to extract voiceprint features through the multi-head self-attention mechanism, positional encoding, and adaptive length processing technology. Among them, the audio is supplemented with time information through positional encoding to ensure the time sequence relationship; adaptive length processing is adopted to dynamically adjust the dimension of the audio features, thereby improving the processing ability for audio of different lengths.
[0101] Furthermore, the difference correction module decomposes the rough features through the residual orthogonalization method, and eliminates the influence of the target speaker's voiceprint by removing the parallel components. The adaptive orthogonalization method dynamically adjusts the orthogonalization strategy according to the features of the input audio to adapt to different types of audio data, thereby effectively restoring the voiceprint of the source speaker.
[0102] Furthermore, the dimension normalization module combines global statistical pooling and adaptive normalization methods to normalize the variable-length audio into a fixed-length output, ensuring the consistency of the output features.
[0103] In summary, this embodiment can effectively restore the true voiceprint features of the source speaker from the forged speech and enhance the anti-forgery ability of voiceprint recognition. The technical details of this method are combined with multiple technologies such as feature processing, orthogonal decomposition, dimension standardization, and enhanced loss, making the voice verification system have high recognition ability and robustness when facing forged speech.
[0104] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0105] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for separating mixed forged voiceprints, characterized in that, It includes the following steps: Extract features from the input forged speech based on the Transformer model to obtain rough features containing the voiceprint information of the source speaker; Decompose the rough features using the residual orthogonalization method to restore the voiceprint features of the source speaker; Perform dimensional normalization on the voiceprint features of the source speaker to obtain voiceprint features of a fixed length; Use the additive angular margin loss to enhance the angular difference between the source speaker and other voiceprints, and output the demixed audio data.
2. The method for separating and unmixing forged voiceprints according to claim 1, wherein Before feature extraction, the forged speech voiceprint demixing method further includes: Preprocess the input forged speech x(t) using a band-pass filter, and obtain the discretized signal according to the following formula: x[n] = x(nT s ) Among them, F s is the sampling frequency, and n is the index of the sampling point; Convert the forged speech from the time domain to the frequency domain through the short-time Fourier transform: In the formula: X(f,t) is the time-frequency representation in the frequency domain; w[n-t] is the window function, and a Hamming window is used to smooth the signal.
3. A method for separating mixed forged voiceprints according to claim 1, characterized in that, Extract features from the input forged speech based on the Transformer model, specifically: Convert the forged speech into Mel frequency features, and use the multi-head self-attention mechanism, positional encoding, and adaptive input length technology to process the Mel frequency features to capture time-domain features and frequency-domain features.
4. A method for separating mixed forged voiceprints according to claim 1, characterized in that, Decompose the rough features using the residual orthogonalization method, specifically Use the residual orthogonalization method to decompose the rough features into components parallel and orthogonal to the voiceprint of the target speaker, remove the parallel components related to the voiceprint of the target speaker, and restore the voiceprint features of the source speaker.
5. A method for separating mixed forged voiceprints according to claim 1, characterized in that, Perform dimensional normalization on the voiceprint features of the source speaker, specifically: Combine the global statistical pooling and adaptive normalization methods to normalize the voiceprint features of the source speaker into an output of a fixed length.
6. A method for separating mixed forged voiceprints according to claim 1, characterized in that, The expression of the additive angular margin loss is: L AAM = max(0, cos(θ target ) - cos(θ other ) + δ) Where: θ target represents the angle between the source speaker and the target voiceprint, and θ other represents the angle between the source speaker and other voiceprints, and δ is an additive margin.
7. A forged voiceprint demixing system, characterized in that, It includes a feature extraction module, a difference correction module, a dimensional normalization module, and a voiceprint enhancement module connected in sequence; The feature extraction module is used to extract features from the input forged speech to obtain rough features containing the voiceprint information of the source speaker; The difference correction module is used to remove the influence of the voiceprint of the target speaker and restore the voiceprint features of the source speaker; The dimensional normalization module is used to perform dimensional normalization on the voiceprint features of the source speaker to obtain voiceprint features of a fixed length; The voiceprint enhancement module is used to enhance the angular difference between the source speaker and other voiceprints through the additive angular margin loss, and output the demixed audio data.
8. A forged voiceprint demixing system according to claim 7, characterized in that, The feature extraction module uses the Transformer model to extract voiceprint features through the multi-head self-attention mechanism, positional encoding, and adaptive length processing technology.
9. A forged voiceprint demixing system according to claim 7, characterized in that, The difference correction module decomposes the rough features through the residual orthogonalization method.
10. A forged voiceprint unmixing system according to claim 7, characterized in that, The dimensional normalization module combines the global statistical pooling and adaptive normalization methods to normalize the variable-length audio into an output of a fixed length.