Method and system for voice protection, and training method of voice protection model

TW202634591AActive Publication Date: 2026-08-16NATIONAL TSING HUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW114104176
Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2026-08-16
Estimated Expiration
2045-02-04

AI Technical Summary

Technical Problem

Existing speech-to-speech technologies can replicate a speaker's voice characteristics, posing security risks for voice theft and fraud, with current detection methods being passive and ineffective in preventing speech duplication.

Method used

A speech protection method and system that trains a voice protection model to generate targeted noise based on speaker features, using an improved WaveUNet architecture to maintain speech intelligibility while preventing copying.

Benefits of technology

Effectively prevents speech from being copied while preserving speaker characteristics, achieving high protection effectiveness and speech quality, suitable for real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TA001072034_001
    Figure TWG2TA001072034_001
  • Figure TWG2TA001072034_002
    Figure TWG2TA001072034_002
  • Figure TWG2TA001072034_003
    Figure TWG2TA001072034_003
Patent Text Reader

Abstract

A training method for voice protection model is provided. The method includes: selecting a first speech and a second speech from the speech database; extracting multiple first features from the first speech; extracting multiple second features from the second speech; generating a first mixed speech based on the first speech, the first features, and the second features using the voice protection model; extracting multiple third features from the first mixed speech; calculating a first distance between the third features and the first features, and a second distance between the third features and the second features; calculating a loss function based on the first distance and the second distance; and updating the speech protection model based on the loss function. The loss function is positively correlated with the first distance, and negatively correlated with the second distance.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a speech processing technology, and more particularly to a speech protection method, system, and training method for a speech protection model. [Previous Technology]

[0002] With the advancement of deep learning technology, speech conversion technology has developed to the point where it can replicate a speaker's voice characteristics with just one sentence. This technology can alter the speaker's identity while preserving the speech content, and is being applied to positive uses such as AI voice-over and cross-language speech conversion. However, this technology also brings serious security risks. There have been reports of criminals using speech conversion technology for telephone fraud, causing significant economic losses.

[0003] Existing solutions mainly focus on deep audio fakery detection, attempting to distinguish between real and synthesized speech. However, these methods can only passively detect copied speech and cannot prevent speech theft, exhibiting significant limitations. With the continuous advancement of speech conversion technology, the effectiveness of detection methods is also facing challenges. [Summary of the Invention]

[0004] In view of this, the present invention provides a speech protection method, system, and training method for a speech protection model, which can effectively prevent speech from being copied by a speech conversion model while maintaining the speaker characteristics of the original speech. Advantageously, this protection mechanism protects speech from being copied while still maintaining the intelligibility of speech for human listeners.

[0005] A first aspect of the present invention provides a method for training a speech protection model, applicable to a system including a speech library. The method includes: selecting a first speech and a second speech from the speech library; extracting multiple first features from the first speech; extracting multiple second features from the second speech; obtaining a first mixed speech using the speech protection model based on the first speech, the multiple first features, and the multiple second features; extracting multiple third features from the first mixed speech; calculating a first distance between the multiple third features and the multiple first features, and a second distance between the multiple third features and the multiple second features; calculating a loss function based on the first distance and the second distance; and updating the speech protection model based on the loss function, wherein the loss function is positively correlated with the first distance and negatively correlated with the second distance.

[0006] In some embodiments of the first aspect, the above-mentioned method of obtaining the first mixed speech using a speech protection model based on the first speech, a plurality of first features and a plurality of second features includes: inputting the first speech, a plurality of first features and a plurality of second features into the speech protection model to generate initial noise; adding the initial noise to the first speech to obtain the first mixed speech.

[0007] In some embodiments of the first aspect, the above-mentioned addition of initial noise to the first speech to obtain the first mixed speech includes: performing a de-offset operation on the initial noise to obtain adjusted noise; and adding the adjusted noise to the first speech based on the first signal-to-noise ratio to obtain the first mixed speech.

[0008] In some embodiments of the first aspect, the first speech and the second speech mentioned above come from different speakers.

[0009] In some embodiments of the first aspect, the first speech is a time-domain signal, and extracting multiple first features from the first speech includes: converting the first speech into a Mel-spectrogram; and extracting multiple first features from the Mel-spectrogram.

[0010] A second aspect of the present invention provides a computer-based method for voice protection, comprising: receiving input voice; obtaining protective noise based on the input voice and a voice protection model; and adding the protective noise to the input voice to obtain protected voice. The aforementioned voice protection model is trained using the method for training a voice protection model proposed in the first aspect of the present invention.

[0011] In some embodiments of the second aspect, the above-mentioned method of obtaining the first mixed speech based on the first speech, a plurality of first features and a plurality of second features using a speech protection model includes: inputting the first speech, a plurality of first features and a plurality of second features into the speech protection model to generate initial noise; adding the initial noise to the first speech to obtain the first mixed speech.

[0012] In some embodiments of the second aspect, the above-mentioned addition of initial noise to the first speech to obtain the first mixed speech includes: performing a de-offset operation on the initial noise to obtain adjusted noise; and adding the adjusted noise to the first speech based on the first signal-to-noise ratio to obtain the first mixed speech.

[0013] In some embodiments of the second aspect, the first speech is a time-domain signal, and extracting multiple first features from the first speech includes: converting the first speech into a Mel-spectrogram; and extracting multiple first features from the Mel-spectrogram.

[0014] A third aspect of the present invention provides a voice protection system, comprising: an input / output element, a memory, and a processor. The input / output element is used to receive input voice. The memory is used to store at least one instruction. The processor is coupled to the input / output element and the memory, and when the processor executes at least one instruction, it is used to: obtain protective noise based on the input voice and a voice protection model; and add the protective noise to the input voice to obtain protected voice. The aforementioned voice protection model is trained by the method for training a voice protection model proposed in the first aspect of the present invention.

Implementation Method

[0016] The following will describe the training method and implementation system for implementing a voice protection model according to embodiments of the present invention with reference to the relevant figures, wherein the same elements will be described with the same reference numerals.

[0017] The following description contains specific information relating to exemplary embodiments of the present invention. The accompanying drawings and detailed description are exemplary embodiments only. However, the present invention is not limited to these exemplary embodiments. Other variations and embodiments of the invention will occur to those skilled in the art. Unless otherwise stated, the same or corresponding elements in the drawings may be indicated by the same or corresponding reference numerals. Furthermore, the drawings and illustrations in the present invention are generally not drawn to scale and are not intended to correspond to actual relative dimensions.

[0018] For the purposes of consistency and ease of understanding, the same features are indicated by reference numerals in the exemplary drawings (although not in some examples). However, features in different embodiments may differ in other respects, and therefore should not be narrowly limited to the features shown in the drawings.

[0019] Terms such as "at least one embodiment," "one embodiment," "multiple embodiments," "different embodiments," "some embodiments," and "this embodiment" indicate that the embodiments of the present invention described herein may include specific features, structures, or characteristics, but not every possible embodiment of the present invention must include such specific features, structures, or characteristics. Furthermore, repeated use of the phrases "in one embodiment" and "in this embodiment" does not necessarily refer to the same embodiment, although they may be the same. Moreover, the use of phrases such as "embodiment" in connection with "the present invention" does not mean that all embodiments of the present invention must include specific features, structures, or characteristics, and should be understood as "at least some embodiments of the present invention" including the stated specific features, structures, or characteristics. The term "coupled" is defined as a connection, whether direct or indirect through an intermediate element, and is not necessarily limited to physical connections. When the term "comprising" is used, it means "including but not limited to," which explicitly indicates an open inclusion or relationship of combinations, groups, series, and equivalents.

[0020] Furthermore, for illustrative and non-limiting purposes, specific details such as functional entities, technologies, agreements, standards, etc., are described to provide an understanding of the described technologies. In other examples, detailed descriptions of well-known methods, technologies, systems, architectures, etc., are omitted to avoid obscuring the explanatory narrative with unnecessary details.

[0021] The terms "first," "second," and "third," etc., in the specification and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or modules is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to these processes, methods, products, or apparatuses.

[0022] Figure 1 is a schematic diagram of a training method for a voice protection model according to an embodiment of the present invention; Figure 2 is a flowchart of a training method for a voice protection model according to an embodiment of the present invention.

[0023] Referring to Figures 1 and 2, the system for training the speech protection model 130 includes a speech library 100 storing multiple speech samples 110; a feature extractor 120; and a speech protection model 130. For clarity, this document labels the first speech sample 111 and the second speech sample 112 among the numerous speech samples 110 in the speech library 100. However, it should be noted that the first speech sample 111 or the second speech sample 112 is only a representative example used to illustrate the embodiments of the present invention, and they are not functionally different from the other speech samples 110. The speech protection model training method proposed in the embodiments of the present invention is applicable to system architectures containing at least two speech samples.

[0024] In step S210, the system first selects a first speech 111 and a second speech 112 from the speech library 100. Speech here refers, for example, a time-domain signal recorded as a waveform of time versus amplitude, such as a PCM format speech file with a sampling rate of 16 kHz, but this invention is not limited thereto. In some embodiments, the first speech 111 and the second speech 112 come from different speakers; for example, the first speech 111 may come from a male speaker, while the second speech 112 may come from a female speaker. In some embodiments, the system can select appropriate speech samples based on the quality and length of the speech; for example, the system can select speech samples with a high signal-to-noise ratio and a long duration to ensure the accuracy of subsequent feature extraction. Advantageously, this targeted selection strategy can improve training effectiveness.

[0025] In step S220, the system extracts multiple first features 1110 from the first speech 111. Specifically, the system can use the feature extractor 120 to extract multiple first features 1110 from the first speech 111.

[0026] In some embodiments, the system, for example, first converts the first speech 111 into a time-frequency domain Mel-spectrogram, and then uses a feature extractor 120 to extract multiple first features 1110 from the Mel-spectrogram. A Mel-spectrogram is a time-frequency domain representation based on human auditory perception, and its frequency distribution is suitable for capturing speaker characteristics, including pitch, rhythm, and timbre. Advantageously, this representation can effectively analyze the extent to which protective noise affects these features.

[0027] In some embodiments, the feature extractor 120 may include sub-modules, such as a speech conversion module and a speaker encoder. For example, the speech conversion module is responsible for converting the raw speech into a Mel spectrogram, while the speaker encoder extracts speaker features from the Mel spectrogram. The speaker encoder may be a pre-trained deep learning model, such as a Speaker Encoder.

[0028] In some embodiments, these first features 1110 may include fundamental frequency features reflecting pitch variations of the speaker, spectral features containing timbre information of the speech, and prosodic features describing the rhythm and stress pattern of the speech, etc., which are not limited to this invention.

[0029] In some embodiments, the feature extractor 120 may also extract acoustic features representing speaker identity characteristics. Advantageously, multi-dimensional feature extraction can capture the characteristics of speech more comprehensively.

[0030] In some embodiments, the feature extractor 120 employs a multi-layer neural network structure. Specifically, this structure includes multiple convolutional and pooling layers for extracting features at different levels from the Mel spectrogram. For example, shallow networks primarily extract basic acoustic features, such as frequency distribution; deeper networks are able to capture more complex speaker features, such as voice personalization features. Advantageously, this hierarchical feature extraction approach can more comprehensively represent various aspects of speech.

[0031] In step S230, the system extracts multiple second features 1120 of the second speech 112. The method by which the system extracts multiple second features 1120 of the second speech 112 in step S230 is the same as the method by which the system extracts multiple first features 1110 of the first speech 111 in step S220, so it will not be described in detail here.

[0032] In some embodiments, the feature extractor 120 for extracting multiple second features 1120 can be the same as the feature extractor 120 for extracting multiple first features 1110. Specifically, using the same feature extractor 120 can ensure the consistency of feature extraction, which is helpful for subsequent feature comparison.

[0033] In some embodiments, the feature extractor 120 for extracting multiple second features 1120 may be different from the feature extractor 120 for extracting multiple first features 1110, in order to obtain more diverse feature information. For example, the feature extractor 120 for processing the second speech 112 may focus more on the overall characteristics of the speech. The present invention does not limit the specific configuration of the feature extractor 120.

[0034] In some embodiments, the feature extractor 120 is frozen during the training of the speech protection model 130, meaning its parameters are not updated. Since the feature extractor 120 is primarily used to extract speaker features of speech, keeping its parameters fixed ensures consistency in feature extraction. Advantageously, this design provides stable feature representations, which is helpful for the training of the speech protection model 130.

[0035] In some embodiments, the voice protection model 130 employs an improved WaveUNet architecture. WaveUNet is an end-to-end deep neural network architecture capable of directly processing audio in the waveform domain, originally designed primarily for audio source separation tasks. This invention improves the WaveUNet architecture to make it more effectively applied to voice protection tasks.

[0036] In detail, the improved WaveUNet architecture introduces a conditional embedding mechanism, making it a three-input architecture. In addition to the original speech waveform input, two conditional inputs are added to the architecture: the embedding vectors of the first feature 1110 and the second feature 1120. The specific improvements include: (1) introducing a new feature processing branch in the downsampling path, which first concatenates the first feature 1110 and the second feature 1120; (2) performing dimensionality transformation on the concatenated features through a linear layer; (3) performing a second concatenation between the transformed features and the output of the original WaveUNet downsampling layer; and (4) processing the concatenation result through a one-dimensional convolutional layer before connecting it back to the original WaveUNet backbone architecture. This improvement enables the model to process both speech waveforms and speaker features simultaneously.

[0037] For example, in the downsampling path, the mixed audio is first processed through a one-dimensional convolutional layer of size 15, then downsampled, and the temporal resolution of the feature map is gradually reduced while the number of channels is increased through multiple downsampling blocks (e.g., L blocks) to extract more abstract features. In the upsampling path, the temporal resolution is gradually restored through corresponding L upsampling blocks using a one-dimensional convolutional layer of size 5. To retain more detailed information, this invention adds multiple skip connections to the architecture, directly connecting the output of each downsampling block to the corresponding upsampling block through cropping and concatenation. Advantageously, this architecture design not only retains more detailed information during the restoration of temporal resolution but also effectively utilizes speaker features for noise generation.

[0038] Advantageously, the improved WaveUNet architecture enables direct computation in the waveform domain, avoiding the dimensionality issues that may arise in the feature extraction and speech synthesis stages. Furthermore, through the design of conditional embedding, the model can self-adjust according to the input features, further improving the processing performance. Combined with temporal operations and feature parameter adjustment mechanisms, the speech protection model 130 possesses flexibility and high adaptability.

[0039] In step S240, the system obtains a first mixed speech 140 based on the first speech 111, multiple first features 1110, and multiple second features 1120 using a speech protection model 130. Specifically, the speech protection model 130 generates initial noise based on the input features, and then the system adds the initial noise to the first speech 111. In some embodiments, the generation process of the initial noise is adjusted by the first features 1110 and the second features 1120. For example, the system can adjust the spectral distribution and intensity of the noise based on these features to generate more targeted protective noise.

[0040] In some embodiments, the input of the voice protection model 130 includes a first voice 111, a first feature 1110 and a second feature 1120, and the output of the voice protection model 130 includes initial noise.

[0041] In some embodiments, the system performs offset removal on the generated initial noise to adjust the mean of the initial noise to zero, thereby obtaining adjusted noise. Then, the system adds the adjusted noise to the first speech 111 based on a first signal-to-noise ratio (SNR) to obtain the first mixed speech 140. The first SNR may be a preset SNR or a user-specified SNR, and this invention is not limited thereto.

[0042] Advantageously, this processing increases the randomness of the noise, avoiding a fixed, biased pattern. Furthermore, when the initial noise with a mean adjusted to zero is added to the first speech 111, it more evenly affects different frequency bands of the speech signal, improving the masking effect on speech features and enhancing protection reliability. Specifically, the system mixes and adjusts the noise with the first speech 111 based on a set signal-to-noise ratio (SNR), thereby achieving a balance between protection effectiveness and speech quality. For example, in situations requiring high speech quality, the user can set a higher SNR, adding weaker noise to help preserve the clarity and intelligibility of the first speech 111 in the first mixed speech 140; while in situations requiring enhanced protection, the user can choose a lower SNR to increase the noise intensity, thereby improving the masking effect on speech features. Advantageously, the flexible SNR adjustment mechanism allows the system to meet the protection needs of different users in various application scenarios.

[0043] In some embodiments, the first signal-to-noise ratio is preset to, for example, 20. Advantageously, this controlled noise generation mechanism can maintain the recognizability of speech while protecting it.

[0044] In step S250, the system extracts a plurality of third features 141 from the first mixed speech 140. In some embodiments, the feature extractor 120 for extracting the third features 141 employs the same configuration as the aforementioned feature extraction to ensure feature consistency. Specifically, the first mixed speech 140 is, for example, a time-domain signal in the form of a time-to-amplitude waveform. The system, for example, first converts the first mixed speech 140 into a Mel spectrogram and then extracts a plurality of third features 141 from it.

[0045] Advantageously, the system operates directly on the original waveform through the voice protection model 130. Compared to traditional methods that require multiple conversions between time-domain, frequency-domain features, and reconstructed time-domain signals, this avoids domain mismatch problems. In some embodiments, the voice protection model 130 can directly generate corresponding protective noise 330 based on the input original waveform and feature representation, and provides a parameterized control mechanism for noise intensity. Users can adjust the protection parameters according to actual needs. For example, for scenarios with high protection requirements, the noise intensity can be increased to enhance the protection effect; for scenarios that need to maintain voice quality, the noise intensity can be reduced to ensure clarity.

[0046] In step S260, the system calculates a first distance between the plurality of third features 141 and the plurality of first features 1110, and a second distance between the plurality of third features 141 and the plurality of second features 1120. This step aims to quantify the degree of difference between different features; in other words, "distance" here refers to a mathematical measure of the difference between two sets of feature parameters.

[0047] Specifically, the first distance is used to measure the similarity between the first mixed speech 140 and the first speech 111 (or between the third feature 141 and the first feature 1110), while the second distance is used to evaluate the similarity between the first mixed speech 140 and the second speech 112 (or between the third feature 141 and the second feature 1120). That is, a smaller first distance indicates a high feature similarity between the first mixed speech 140 and the first speech 111, while a larger second distance indicates a low feature similarity between the first mixed speech 140 and the second speech 112.

[0048] In some embodiments, multiple methods can be used to calculate the distance or similarity between features. Specifically, cosine similarity can be used to calculate the similarity between feature vectors, or Euclidean distance can be used to calculate the difference between feature vectors. For example, when using cosine similarity, a higher value indicates that the features are more similar; when using Euclidean distance, a larger value indicates that the features are more different. Advantageously, this dual calculation design can simultaneously ensure the recognizability of speech and its anti-copying effect.

[0049] In some embodiments, the system employs an adjustable weighting mechanism to balance the influence of the first distance and the second distance. Specifically, the system introduces a weighting parameter λ (where λ is a real number greater than 0) to calculate a first weighted value for the first distance and a second weighted value for the second distance. For example, when it is necessary to emphasize the preservation of original speech features, the value of λ can be decreased to enhance the contribution of the first distance; when it is necessary to emphasize the anti-copying effect, the value of λ can be increased to enhance the contribution of the second distance. Advantageously, this adjustable weighting mechanism allows the system to adjust the protection strength according to different application requirements, thereby achieving a balance between speech clarity and protection effect.

[0050] In step S270, the system calculates a loss function based on the first distance and the second distance. Specifically, the loss function is positively correlated with the first distance and negatively correlated with the second distance. In some embodiments, the system adjusts the strength of these two correlations by setting different weighting values ​​using a weighting parameter λ. For example, to emphasize maintaining the recognizability of the original speech, the weight of the first distance can be increased; to emphasize the effect of preventing duplication, the weight of the second distance can be increased. In other embodiments, these weighting parameters can be dynamically adjusted during training. Advantageously, this flexible loss function design allows the protection effect to be adjusted according to specific needs.

[0051] In some embodiments, the loss function (denoted by LCS) can be calculated as follows: where represents the feature embedding of the first mixed speech 140 (e.g., the third feature 141), represents the feature embedding of the first speech 111 (e.g., the first feature 1110), represents the feature embedding of the second speech 112 (e.g., the second feature 1120), cos represents the cosine similarity, and λ is the weight parameter. Specifically, the first term ensures the similarity between the first mixed speech 140 and the first speech 111, while the second term suppresses the similarity between the first mixed speech 140 and the second speech 112. In other words, we want them to be as far apart as possible, and as close as possible. Advantageously, using cosine similarity as the loss function is not only computationally efficient, but also effectively measures the angular relationship between vectors, making it particularly suitable for distance metrics in the feature embedding space.

[0052] In some embodiments, the system may also use mean squared error (MSE loss) as a distance metric, calculating the loss function (denoted by LMSE) as follows: where MSE represents the mean squared error function, used to measure the square of the Euclidean distance between two embedding vectors. The negative sign of the first term indicates that we want to maximize the distance between the first mixed speech 140 and the first speech 111, while the second term indicates that we want to minimize the distance between the first mixed speech 140 and the second speech 112. Advantageously, using MSE as the loss function can directly reflect the Euclidean distance between vectors, the calculation result has a clear geometric meaning, and the gradient calculation is simple, which helps to stabilize the training of the model.

[0053] In step S280, the system updates the voice protection model 130 according to the loss function. In some embodiments, the update process employs optimization algorithms such as gradient descent. Specifically, the system calculates the gradient of the loss function with respect to the model parameters and adjusts the model parameters accordingly.

[0054] In some embodiments, in each iteration, the system randomly selects a batch of training sample pairs from the training set. For example, each batch contains 4 sets of training sample pairs, where each set of sample pairs contains a first speech 111 and a second speech 112, and these speeches come from different speakers. The system does not fix a specific combination of speakers, but randomly selects different speaker pairings throughout the training process to improve the generalization ability of the model.

[0055] In some embodiments, in the next iteration, the system randomly selects new pairs of speech samples from the training set. For example, the system selects four new pairs of speech samples 110, each pair containing a new first speech 111 and a new second speech 112. In some embodiments, the system ensures that these new speech samples come from speakers different from those in the previous iteration, thereby increasing the diversity of the training data. This training sample selection strategy continues throughout the training process until a preset stopping condition is met, such as reaching a specific number of iterations or a convergence criterion for the loss function. Advantageously, this dynamic sample selection mechanism can improve the model's performance when dealing with different combinations of speakers.

[0056] In other embodiments, the system may employ more complex optimization strategies, such as adaptive learning rates, momentum methods, etc. For example, when the loss function decreases slowly, the system may automatically adjust the learning rate or use a momentum term to accelerate convergence. Advantageously, such optimization strategies, combined with random sampling training methods, can effectively improve the model's generalization performance and training efficiency.

[0057] In some embodiments, the system employs a cyclic learning rate strategy for model updates. For example, the system may use an optimizer (e.g., the Adam optimizer) and cyclically vary the learning rate between a maximum value of 0.001 and a minimum value of 0.00001. Advantageously, this learning rate adjustment strategy helps the model converge to the optimal solution more effectively.

[0058] Figure 3 is a schematic diagram of a voice protection method according to an embodiment of the present invention; Figure 4 is a flowchart of a voice protection method according to an embodiment of the present invention.

[0059] In addition to providing a training method for a voice protection model, the present invention also provides a method for actually protecting voice. Specifically, the method uses the voice protection model 130 trained through the above steps S210 to S280 to perform protection processing on the input voice.

[0060] Specifically, when a user wants to protect their voice from being stolen or forged by deepfake technology, they can first input the voice to be made public into the voice protection model 130 trained by the above steps S210 to S280, and then process the voice for public use.

[0061] In some embodiments, the application scenarios of voice protection can cover various situations requiring voice security protection. For example, before giving a public speech, a politician can process the recording of the speech using the voice protection model 130 of this invention before uploading it to an online platform to prevent their voice from being stolen and used to create deepfake videos containing false information. Similarly, before uploading audio files to social media, celebrities can also use the method of this invention to protect their voice characteristics, reducing the risk of being defrauded by others using voice conversion technology. Advantageously, the voice protection method provided by this invention not only effectively protects users' voice characteristics from being stolen by deepfake technology, but also maintains the naturalness and intelligibility of the voice. This allows users to maintain normal voice communication and usage needs while protecting their voice security.

[0062] Referring to Figures 3 and 4, in step S410, the system receives the input voice 310 that needs to be protected. In some embodiments, the input voice 310 may be voice recorded in real time or a pre-stored voice file. Specifically, the system supports a variety of common voice formats, such as WAV, MP3, etc., which are not limited to this invention. Advantageously, this flexible input support allows the protection method of the present invention to be widely applied in various scenarios.

[0063] In step S420, the system inputs the input speech 310 to the trained speech protection model 130 to obtain protective noise 330. In some embodiments, the parameters of the speech protection model 130 are fixed, meaning that for the same input speech 310, the model will produce the same protective noise 330. In other embodiments, the model may introduce randomness, such that even if the same input speech 310 is processed each time, slightly different protective noise 330 will be produced. Advantageously, this design increases the unpredictability of the protection mechanism, further improving security.

[0064] In some embodiments, the system extracts features from the input speech 310 to obtain input speech features. Specifically, the system converts the input speech 310 into a Mel spectrogram and extracts input speech features from the Mel spectrogram. Simultaneously, the system selects a speaker of a different gender from the input speech 310 from the speech library 310 as an auxiliary speech, converts the auxiliary speech into a Mel spectrogram, and extracts auxiliary speech features from the Mel spectrogram. Then, the system inputs the input speech 310, the input speech features, and the auxiliary speech features into the speech protection model 130 to generate protective noise 330. It should be noted that the system still requires the auxiliary speech features to participate in the calculation when generating the protective noise 330, which is the same as the operation during the training phase. However, in practical applications, it is not necessary to use a large number of different speech samples 110 for iterative training as in the training phase. Furthermore, the system selects an auxiliary speech of a different gender from the input speech 310 because this helps the speech protection model 130 generate more effective protective noise 330, enabling the speech features to effectively prevent speech duplication while preserving the intelligibility of the input speech 310. Advantageously, this design ensures consistent and reliable protection.

[0065] In some embodiments, the system performs a mean shift on the generated protective noise 330. For example, the system calculates the average value of the protective noise 330 along the time axis and subtracts this average value from the protective noise 330, making the mean value of the protective noise zero. Advantageously, this mean shifting process increases the randomness of the protective noise, avoiding the noise being merely a simple displacement of the original signal, and improving the reliability of the protection effect.

[0066] In step S430, the system adds protective noise 330 to the input speech 310 to obtain protective speech 340. In some embodiments, this addition process is a direct signal superposition in the time domain. For example, the system adjusts the intensity of the protective noise 330 according to a preset signal-to-noise ratio (SNR), and then adds the adjusted noise signal to the input speech 310 signal.

[0067] In some embodiments, the system calculates the intensity of the protective noise 330 as follows: where is the protective noise 330, x is the input speech 310, and z is the signal-to-noise ratio constraint parameter (noise constraint parameter), for example, z can be set to 20. After the intensity of the protective noise 330 is adjusted, the system adds the adjusted protective noise 330 to the input speech 310 to obtain the protective speech 340.

[0068] It is worth noting that the speech protection method of the present invention has several advantages. In some embodiments, since the noise is generated targeting specific speech features, it is possible to maintain the intelligibility of the speech while protecting it from illegal copying. Specifically, the protected speech remains clearly intelligible to human listeners, but it is difficult for speech conversion systems to extract effective features. In other embodiments, the protection method of the present invention has low computational complexity, which makes it suitable for real-time processing scenarios. Advantageously, these characteristics enable the method of the present invention to play an important role in practical applications.

[0069] The method of the present invention can be applied to a variety of practical scenarios. In some embodiments, it can be used to protect voice messages, such as voice memos, voice calls, etc. Specifically, the system can automatically add protective noise before the voice is stored or transmitted. In other embodiments, it can be integrated into a voice assistant or smart device to provide protection in real time while the user is speaking. Advantageously, this real-time protection mechanism can effectively prevent voice from being illegally recorded and copied.

[0070] For example, in financial industry applications, the method of the present invention is particularly suitable for voice banking services. When a customer makes a voice verification or transaction instruction via telephone, the system can immediately protect the customer's voice to prevent it from being recorded and used to impersonate the customer. Advantageously, this protection mechanism can effectively prevent voice-based fraud.

[0071] For example, in the news media industry, the method of the present invention can be used to protect the broadcasting voices of news professionals. Specifically, the system can automatically add protection processing before the news release voice files are uploaded to a public platform to prevent others from stealing the anchor's voice to create fake news. Advantageously, this application can effectively maintain the authenticity of news and the credibility of the media.

[0072] For example, in professional conference settings, the method of the present invention can be applied to protect the speech content of participants. Specifically, the system can protect voice in real time during video conferences to prevent important speech content from being extracted by others and used to create false statements.

[0073] In some embodiments, to verify the advantages of the present invention compared to the prior art, objective evaluation tests include indicators such as attack success rate (ASR) and preservation success rate (PSR). In a white-box scenario, the present invention achieves an attack success rate (ASR) of 0.94 while maintaining a preservation success rate (PSR) of 0.82, which is a significant improvement compared to the prior art's attack success rate of 0.53. In a black-box scenario, the present invention still achieves an attack success rate of 0.85. Furthermore, while maintaining a signal-to-noise ratio (SNR) as high as 19.98, the present invention achieves a speech quality perception evaluation score (PESQ) of 1.90 and a short-term objective intelligibility (STOI) of 0.81, demonstrating excellent speech quality protection capabilities and reliability.

[0074] In some embodiments, in terms of subjective evaluation, human ear recognition tests were conducted, and the results showed that the present invention can effectively balance the protective effect and speech quality. Specifically, the speech processed using the method of the present invention can be clearly identified by the subject as having the characteristics of the original speaker, while effectively preventing the speech conversion system from copying the speaker's characteristics. Advantageously, these experimental results verify the advantages of the present invention in protecting speech features and maintaining speech quality.

[0075] Figure 5 is a schematic block diagram of a computing system according to an embodiment of the present invention.

[0076] Referring to Figure 5, the method for training the voice protection model described herein is a computer-implemented method, which can be implemented on a computing system 500 having various hardware components. Similarly, the voice protection method described herein is also a computer-implemented method and can also be implemented on a computing system 500 having various hardware components. In some embodiments, the computing system 500 may be implemented as an electronic device, which includes, for example, but is not limited to, one or more of the following components: a processor (Central Processing Unit, CPU) 520, a graphics processing unit (GPU) 550, an input / output element 530, a network element 540, and a memory 510. These components can communicate and transmit data via the system bus 590. However, the present invention does not limit the specific model, quantity, and configuration of the components. Those skilled in the art can adjust, select, or add or remove components according to the specific needs and operating environment when implementing the present invention.

[0077] In some embodiments, the main computing core inside the computing system 500 is one or more processors 520. This processor 520 can be responsible for running the main computing processes and related control logic of algorithms such as voice protection. In some embodiments, the processor 520 is used to execute processing instructions (i.e., machine-executable instructions) stored in a non-volatile computer-readable medium (e.g., storage device 560).

[0078] In some embodiments, to improve the computational efficiency of the voice protection model, the computing system 500 may further include one or more graphics processors 550 designed specifically for performing massive parallel computations. These graphics processors 550 can effectively enhance the system's computational power when training and inferring the voice protection model.

[0079] In some embodiments, the computing system 500 may include a variety of input / output elements 530 for receiving user input and displaying system output. For example, such input / output elements 530 may include a keyboard, mouse, touchpad, display screen, speaker, and other types of sensing devices.

[0080] In some embodiments, the computing system 500 may also include a network element 540 for network communication. For example, this network element 540 may include a network interface card for wired or wireless network connections, or a communication module for 3G, 4G, 5G or other wireless communication technologies.

[0081] In some embodiments, the computing system 500 may include one or more memory elements 510, such as volatile memory elements like random access memory (RAM). The memory 510 stores parameters of the voice protection model, as well as other data and programs used to run algorithms such as the voice protection model.

[0082] In addition, the computing system 500 may also include one or more of the following elements: storage device 560, power management element 570 and other elements 580 (e.g., other hardware elements).

[0083] In some embodiments, the computing system 500 may include one or more storage devices 560, such as non-volatile memory elements like hard disk drives (HDDs) or solid-state drives (SSDs). These storage devices 560 can be used to store information such as the code, training data, and model parameters of the voice protection model software. Furthermore, the storage devices 560 can also be used to store intermediate results and final outputs of algorithms such as the voice protection model.

[0084] In some embodiments, the computing system 500 may include one or more power management elements 570 for providing power to various hardware components of the computing system 500 and managing their power consumption. This power management element 570 may include a battery, a power converter, and other power management devices.

[0085] In some embodiments, the computing system 500 may also include other components 580, such as cooling fans, heat sinks, and various other control and monitoring devices, but the present invention is not limited thereto.

[0086] In summary, the speech protection model training method and system proposed in this embodiment of the invention have the following technical advantages: First, unlike existing passive detection methods, this invention achieves an active speech protection mechanism through the design of a loss function based on feature distance, enabling the speech protection model to effectively learn how to generate targeted protection noise; Second, unlike the traditional two-stage processing method that requires feature extraction followed by speech synthesis, this invention adopts a direct waveform generation method, avoiding the problem of feature mismatch; Third, compared to the problem that existing technologies cannot simultaneously address the protection effect and speech intelligibility, this invention, through an improved WaveUNet architecture, achieves effective protection against speech copying while maintaining the naturalness and intelligibility of the speech; Fourth, the real-time protection mechanism provided by this invention overcomes the limitation of traditional passive detection methods that can only detect copied speech, and can more effectively prevent speech from being illegally copied and misused.

[0087] Based on the above description, it is evident that various techniques can be used to implement the concepts described in this application without departing from the scope of these concepts. Furthermore, although the concepts have been described with specific reference to certain embodiments, those skilled in the art will recognize that changes in form and detail may be made without departing from the scope of these concepts. Thus, the described embodiments are to be considered illustrative rather than restrictive in all respects. Moreover, it should be understood that this application is not limited to the specific embodiments described above, but many rearrangements, modifications, and substitutions can be made without departing from the scope of the invention. [Simplified Explanation of the Diagram]

[0015] Figure 1 is a schematic diagram of a training method for a voice protection model according to an embodiment of the present invention; Figure 2 is a flowchart of a training method for a voice protection model according to an embodiment of the present invention; Figure 3 is a schematic diagram of a voice protection method according to an embodiment of the present invention; Figure 4 is a flowchart of a voice protection method according to an embodiment of the present invention; Figure 5 is a block diagram of a computing system according to an embodiment of the present invention.

Claims

1. A method for training a speech preservation model, applicable to a system including a speech library, the method comprising: Select a first speech and a second speech from the speech database; extract multiple first features from the first speech; Extract multiple second features from the second speech; Based on the first speech, the first features, and the second features, a first mixed speech is obtained using the speech protection model; multiple third features are extracted from the first mixed speech; a first distance between the third features and the first features, and a second distance between the third features and the second features are calculated; a loss function is calculated based on the first distance and the second distance; and the speech protection model is updated based on the loss function, wherein: the loss function is positively correlated with the first distance, and the loss function is negatively correlated with the second distance.

2. The method as described in claim 1, wherein obtaining the first mixed speech using the speech preservation model based on the first speech, the first features, and the second features includes: The first speech, the first features, and the second features are input into the speech protection model to generate initial noise; The initial noise is added to the first speech to obtain the first mixed speech.

3. The method as described in claim 2, wherein adding the initial noise to the first speech to obtain the first mixed speech includes: The initial noise is de-offset to obtain the adjusted noise; And based on the first signal-to-noise ratio, the adjusted noise is added to the first speech to obtain the first mixed speech.

4. The method as described in claim 1, wherein the first speech and the second speech are from different speakers.

5. The method of claim 1, wherein the first speech is a time-domain signal, and extracting the first features of the first speech comprises: Convert the first speech into a Mel-spectrogram. And extract these first features from the Mel spectrogram.

6. A computer implementation method for voice protection, the method comprising: Receive input voice; Based on the input speech and the speech protection model, protective noise is obtained; The process involves adding protective noise to the input speech to obtain protected speech, wherein the speech protection model is trained via the following steps: selecting a first speech and a second speech from a speech database; extracting multiple first features from the first speech; extracting multiple second features from the second speech; obtaining a first mixed speech based on the first speech, the first features, and the second features using the speech protection model; extracting multiple third features from the first mixed speech; calculating a first distance between the third features and the first features, and a second distance between the third features and the second features; calculating a loss function based on the first distance and the second distance; and updating the speech protection model based on the loss function, wherein the loss function is positively correlated with the first distance and negatively correlated with the second distance.

7. The method of claim 6, wherein obtaining the first mixed speech using the speech protection model based on the first speech, the first features, and the second features comprises: The first speech, the first features, and the second features are input into the speech protection model to generate initial noise; The initial noise is added to the first speech to obtain the first mixed speech.

8. The method as described in claim 7, wherein adding the initial noise to the first speech to obtain the first mixed speech comprises: The initial noise is de-offset to obtain the adjusted noise; And based on the first signal-to-noise ratio, the adjusted noise is added to the first speech to obtain the first mixed speech.

9. The method of claim 6, wherein the first speech is a time-domain signal, and extracting the first features of the first speech comprises: Convert the first speech into a Mel-spectrogram. And extract these first features from the Mel spectrogram.

10. A voice protection system, the system comprising: An input element for receiving input voice memory for storing at least one instruction; The processor, coupled to the input element and the memory, wherein when the processor executes the at least one instruction, it is used to: obtain protective noise based on the input speech and the speech protection model; and add the protective noise to the input speech to obtain protected speech, wherein the speech protection model is trained by the following steps: selecting a first speech and a second speech from a speech library; extracting a plurality of first features of the first speech; extracting a plurality of second features of the second speech; obtaining a first mixed speech based on the first speech, the first features, and the second features using the speech protection model; extracting a plurality of third features of the first mixed speech; calculating a first distance between the third features and the first features, and a second distance between the third features and the second features; calculating a loss function based on the first distance and the second distance; and updating the speech protection model based on the loss function, wherein: the loss function is positively correlated with the first distance, and the loss function is negatively correlated with the second distance.