Speech enhancement training method, device, equipment and medium based on speaker perception
By introducing training methods and adversarial training mechanisms based on speaker perception in the speech enhancement and speaker recognition system, the problem of low system performance in noisy environments is solved, and more accurate speaker recognition and more effective noise suppression are achieved.
Patent Information
- Application Number
- CN202510344834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-24
AI Technical Summary
In noisy environments, traditional speech enhancement and speaker recognition system have problems such as accumulation of errors and conflicts between task goals, making it difficult for the system performance to reach an ideal state.
The speech enhancement training method based on speaker perception is adopted, and the noisy voice is denoised through the initial speech enhancement joint training system, and the speaker sensitive features are extracted through a shared encoder. Combined with the adversarial training mechanism, the system parameters are updated to reduce error transmission and target conflict.
Accurate training of speaker recognition tasks is achieved in noisy environments, reducing conflicts and error accumulation between speech enhancement and speaker recognition tasks, and improving the performance of the system in a low signal-to-noise ratio environment.
Smart Images

Figure CN119851671B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a speech enhancement training method, device, equipment and medium based on speaker perception. Background Art
[0002] In the field of speech processing, the performance of traditional automatic speech recognition (ASR) and speaker diarization (SD) systems has significantly decreased in noisy environments, which seriously affects the user experience of voice interaction systems. At present, the widely adopted solution in the industry is to build a "tandem" processing system: first, the noise interference in the noisy speech is filtered out by the speech enhancement module, and then the enhanced speech is input into the speaker recognition module for feature extraction and classification. Although this tandem architecture is simple and direct in engineering implementation, it has significant defects in both theory and practice. For example, the traditional tandem system has obvious error accumulation problems. The two tasks of speech enhancement and speaker recognition have natural goal conflicts. In particular, the traditional speech enhancement module usually optimizes the signal-to-noise ratio or reduces the mean square error, and tends to over-filter the changing features in the speech components; and these filtered acoustic features (such as glottal characteristics, intonation fluctuations, etc.) contain rich speaker identity information, which is crucial for downstream speaker recognition. This inherent conflict in task goals makes it difficult for the overall system performance to reach an ideal state.
[0003] In summary, how to achieve accurate training of speaker recognition tasks in noisy environments and reduce the conflict and error accumulation between speech enhancement and speaker recognition tasks is a technical problem to be solved in this field. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a speech enhancement training method, device, equipment and medium based on speaker perception, which can realize accurate training of speaker recognition tasks in noisy environments and reduce the conflict and error accumulation between speech enhancement and speaker recognition tasks. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a speech enhancement training method based on speaker perception, comprising:
[0006] Inputting a first noisy speech sample into an initial speech enhancement joint training system so that a pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, outputs a first enhanced speech sample to determine a speech enhancement loss, and inputs the first enhanced speech sample into a pre-trained shared encoder of the initial speech enhancement joint training system;
[0007] Extracting features of a speaker sensitive area of the first enhanced speech sample by using the pre-trained shared encoder to obtain a first speaker sensitive feature, and inputting the first speaker sensitive feature into a pre-trained speaker recognition module, so as to perform speaker ID classification on the first speaker sensitive feature by using the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculate a speaker recognition loss;
[0008] The adversarial training discriminator of the initial speech enhancement joint training system is used to judge whether the first speaker sensitive feature has noise, and the adversarial loss is adjusted according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss and the adversarial loss to obtain the trained target speech enhancement joint training system.
[0009] Optionally, updating the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss to obtain a trained target speech enhancement joint training system includes:
[0010] Constructing a multi-task loss function based on the speech enhancement loss, the speaker recognition loss, the adversarial loss and corresponding weight coefficients;
[0011] Based on the multi-task loss function, a step-by-step update strategy is performed on the system parameters of the initial speech enhancement joint training system to balance the speech enhancement task objectives and the speaker recognition task, thereby obtaining a trained target speech enhancement joint training system.
[0012] Optionally, the step of adjusting the adversarial loss according to the error between the noise judgment result and the noise sample label of the first noisy speech sample includes:
[0013] If the noise determination result of the first noisy speech sample is a noisy speech determination result, increasing the adversarial loss based on the gradient reversal process;
[0014] If the noise determination result of the first noisy speech sample is a noise-free speech determination result, the adversarial loss is reduced based on the gradient inversion process.
[0015] Optionally, the first noisy speech sample is a speech sample carrying a real noise sample label and a real speaker ID label;
[0016] Accordingly, the calculation of speaker recognition loss includes:
[0017] A speaker recognition loss is determined based on the speaker ID classification prediction result and the speaker ID label.
[0018] Optionally, the speech enhancement training method based on speaker perception further includes:
[0019] Freezing the initial speaker recognition module, and performing basic noise reduction training on the initial speech enhancement module using the second noisy speech sample to obtain a pre-trained speech enhancement module;
[0020] The pre-trained speech enhancement module is frozen, and the initial shared encoder and the initial speaker recognition module connected to the initial shared encoder are trained for speaker feature extraction using clean speech samples to obtain a pre-trained shared encoder and a pre-trained speaker recognition module.
[0021] Optionally, the using the second noisy speech sample to perform basic noise reduction training on the initial speech enhancement module to obtain a pre-trained speech enhancement module includes:
[0022] Acquire a clean speech sample, and perform noise adding processing on the clean speech sample to obtain a second noisy speech sample;
[0023] Inputting the second noisy speech sample into the SEGAN structure to perform a primary noise reduction process on the second noisy speech sample to obtain a process speech sample;
[0024] The time domain loss value and the frequency domain loss value between the process speech sample and the clean speech sample are calculated so that the Adam optimizer reversely updates the network parameters of the SEGAN structure according to the time domain loss value, the frequency domain loss value, and the first preset learning rate to obtain a pre-trained speech enhancement module.
[0025] Optionally, the using of clean speech samples to perform speaker feature extraction training on an initial shared encoder and an initial speaker recognition module connected to the initial shared encoder to obtain a pre-trained shared encoder and a pre-trained speaker recognition module comprises:
[0026] Inputting the clean speech sample into the initial shared encoder to train the initial shared encoder to learn to extract speaker-sensitive features, so as to obtain a pre-trained shared encoder, and output a second speaker-sensitive feature;
[0027] Inputting the second speaker sensitive feature into a speaker recognition network to train the speaker recognition network to extract a speaker embedding vector from the second speaker sensitive feature;
[0028] Calculating the classification loss value between the speaker embedding vector and the class center vector using the AM-Softmax loss function, wherein the class center vector is obtained by initializing the vector distribution information of the speaker embedding vector of each of the clean speech samples;
[0029] The network parameters of the speaker recognition network are reversely updated according to the classification loss value and the second preset learning rate by an Adam optimizer to obtain a pre-trained speaker recognition module.
[0030] In a second aspect, the present application discloses a speech enhancement training device based on speaker perception, comprising:
[0031] a speech enhancement module, configured to input a first noisy speech sample into an initial speech enhancement joint training system, so that a pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, output a first enhanced speech sample to determine a speech enhancement loss, and input the first enhanced speech sample into a pre-trained shared encoder of the initial speech enhancement joint training system;
[0032] a speaker recognition module, configured to perform feature extraction of a speaker sensitive area on the first enhanced speech sample through the pre-trained shared encoder to obtain a first speaker sensitive feature, and input the first speaker sensitive feature into the pre-trained speaker recognition module, so as to perform speaker ID classification on the first speaker sensitive feature through the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculate a speaker recognition loss;
[0033] The model training module is used to judge whether the first speaker sensitive feature has noise through the adversarial training discriminator of the initial speech enhancement joint training system, and adjust the adversarial loss according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss and the adversarial loss to obtain the trained target speech enhancement joint training system.
[0034] In a third aspect, the present application discloses an electronic device, comprising:
[0035] Memory, used to store computer programs;
[0036] A processor is used to execute the computer program to implement the steps of the aforementioned disclosed speech enhancement training method based on speaker perception.
[0037] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned disclosed speech enhancement training method based on speaker perception are implemented.
[0038] It can be seen that the present application discloses a speech enhancement training method based on speaker perception, comprising: inputting a first noisy speech sample into an initial speech enhancement joint training system so that a pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, outputs a first enhanced speech sample to determine the speech enhancement loss, and inputs the first enhanced speech sample into a pre-trained shared encoder of the initial speech enhancement joint training system; extracting features of a speaker sensitive area of the first enhanced speech sample through the pre-trained shared encoder to obtain a first speaker sensitive feature, and inputting the first speaker sensitive feature into the pre-trained shared encoder. The speaker recognition module is trained so that the speaker ID classification of the first speaker sensitive feature is performed through the pre-trained speaker recognition module, the speaker ID classification prediction result is obtained, and the speaker recognition loss is calculated; the adversarial training discriminator of the initial speech enhancement joint training system is used to judge whether the first speaker sensitive feature has noise, and the adversarial loss is adjusted according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss, and obtain the trained target speech enhancement joint training system. It can be seen that by jointly training the speech enhancement module, the shared encoder and the speaker recognition module, the negative impact of the acoustic distortion introduced by the speech enhancement module in the traditional serial system on the downstream speaker recognition task is avoided; the adversarial training mechanism further constrains the output of the speech enhancement module to ensure that the enhanced speech retains the speaker sensitive features and reduces error propagation. Furthermore, through adversarial training and multi-task loss design, the target conflict between the two tasks (speech enhancement task and speaker recognition task) is coordinated to ensure that the enhancement module retains the speaker identity information while reducing noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0040] Figure 1 A flow chart of a speech enhancement training method based on speaker perception disclosed in this application;
[0041] Figure 2 A flowchart of a specific speech enhancement training method based on speaker perception disclosed in this application;
[0042] Figure 3This is a schematic diagram of the structure of a speech enhancement training device based on speaker perception disclosed in this application;
[0043] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0045] In the field of speech processing, the performance of traditional automatic speech recognition and speaker recognition systems significantly degrades in noisy environments, seriously affecting the user experience of voice interaction systems. At present, the solution widely adopted in the industry is to build a "tandem" processing system: first, the noise interference in the noisy speech is filtered out through the speech enhancement module, and then the enhanced speech is input into the speaker recognition module for feature extraction and classification. Although this tandem architecture is simple and direct in engineering implementation, it has significant defects in both theory and practice.
[0046] First, the traditional cascade system has obvious error accumulation problems. The speech enhancement module may introduce new acoustic distortions, such as music noise or improper filtering of speech components, which will be directly passed to the downstream speaker recognition module, causing feature extraction bias and recognition errors. Especially in low signal-to-noise ratio (SNR < 5dB) environments, the instability of the speech enhancement module can cause a sharp drop in downstream task performance. Several studies have shown that in some cases, directly using the original noisy speech can even produce better speaker recognition results than using enhanced speech, which highlights the inherent defects of traditional architectures.
[0047] Secondly, there is a natural conflict between the two tasks of speech enhancement and speaker recognition. Traditional speech enhancement modules usually optimize the signal-to-noise ratio or reduce the mean square error, and tend to over-filter the changing features in the speech components; however, these filtered acoustic features (such as glottal characteristics, intonation fluctuations, etc.) contain rich speaker identity information, which is crucial for downstream speaker recognition. This inherent conflict in task objectives makes it difficult for the overall system performance to reach an ideal state.
[0048] Third, the problem of incompatible feature representation also restricts system performance. Speech enhancement modules usually output Mel-spectrogram features or time-domain waveforms. Although these representations are suitable for speech recognition tasks, they may not be the optimal input form for speaker recognition. Existing studies have shown that different tasks have significant differences in sensitivity to acoustic features: speech content recognition focuses more on the spectral envelope, while speaker identity recognition relies on fine-grained features such as pitch trajectory and harmonic structure. Traditional tandem systems lack explicit modeling of such feature differences.
[0049] In addition, the poor noise robustness of existing solutions is a common problem. Mainstream speaker recognition models (such as x-vector, d-vector, etc.) perform well on clean speech, but their performance drops sharply in complex noise environments. Research data shows that when the background noise drops from 10dB to 0dB, the speaker error rate of a typical system may increase by 3-5 times, which severely limits the deployment of existing technologies in practical application scenarios.
[0050] Although some studies in recent years have attempted to optimize cascade systems through end-to-end training strategies, these methods mainly focus on joint parameter optimization and lack a deep understanding and explicit modeling of the relationship between tasks. For example, simply connecting a pre-trained speech enhancement network to a speaker recognition network and fine-tuning it, although it has achieved some improvements on specific datasets, it has failed to solve the underlying structural problems, especially when generalizing to different noise types.
[0051] In summary, although the existing technologies have made great progress in their respective fields, they still have systematic defects in speaker recognition tasks in noisy environments. There is an urgent need for a unified framework that can fully coordinate speech enhancement and speaker feature extraction to improve the robustness and accuracy of the system in practical application scenarios.
[0052] To this end, the present invention provides a speech enhancement training scheme based on speaker perception, which can achieve accurate training of speaker recognition tasks in noisy environments and reduce conflicts and error accumulation between speech enhancement and speaker recognition tasks.
[0053] Reference Figure 1 As shown, this embodiment discloses a speech enhancement training method based on speaker perception, comprising:
[0054] Step S11: Input the first noisy speech sample into the initial speech enhancement joint training system so that the pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, outputs a first enhanced speech sample to determine the speech enhancement loss, and inputs the first enhanced speech sample into the pre-trained shared encoder of the initial speech enhancement joint training system.
[0055] In this embodiment, an initial speech enhancement joint training system is first constructed, wherein the system architecture of the joint training system adopts a cascade network structure and is composed of four core modules, specifically including: a pre-trained speech enhancement module, a pre-trained shared encoder, a pre-trained speaker recognition module, and an adversarial training discriminator. Specifically, the pre-trained speech enhancement module is connected to the pre-trained shared encoder, the pre-trained shared encoder is connected to the pre-trained speaker recognition module, and the pre-trained shared encoder is connected to the adversarial training discriminator.
[0056] The cascade network architecture proposed in the present invention consists of two modules: the speech enhancement front-end and the speaker recognition back-end, and deep fusion is achieved through a shared coding layer. The speech enhancement front-end is based on the SEGAN (Structure-Enhanced Generative Adversarial Network) structure, using a six-layer one-dimensional convolution encoder and a symmetrical transposed convolution decoder, and setting jump connections between corresponding layers to retain audio detail information. For the input 16kHz sampling rate speech waveform, the front-end module outputs an enhanced speech waveform of the same length, effectively achieving noise suppression while retaining the key time-frequency features of the speech.
[0057] In this embodiment, since the system architecture is a cascade architecture, before the system is used, each module of the system needs to be trained separately and end-to-end jointly. First, it is trained separately. The specific separate training process is as follows:
[0058] Before end-to-end joint training, the speech enhancement module and speaker recognition module need to be pre-trained separately. Specifically, the pre-training steps of each module are as follows:
[0059] The initial speaker recognition module is frozen, and the initial speech enhancement module is trained with the second noisy speech sample for basic noise reduction to obtain a pre-trained speech enhancement module; it can be understood that when training the initial speech enhancement module, the parameters of each module in the initial speaker recognition module are frozen, and the initial speech enhancement module performs speech noise reduction training on the SEGAN codec structure (including an independent encoder) based on the second noisy speech sample, that is, the codec structure based on the generative adversarial network (SEGAN) is used to receive the second noisy speech sample input and output enhanced speech, and the initial shared encoder has not yet been activated or participated in training. At this time, only the encoder inside SEGAN (for speech noise reduction) is trained, and the shared encoder, as an intermediate layer connecting the speech enhancement module and the speaker recognition module, is not used at this stage. The specific training process of the initial speech enhancement module is as follows:
[0060] Obtain a clean speech sample, add noise to the clean speech sample to obtain a second noisy speech sample; input the second noisy speech sample into the SEGAN structure to perform initial noise reduction on the second noisy speech sample to obtain a process speech sample; calculate the time domain loss value and frequency domain loss value between the process speech sample and the clean speech sample, so that the Adam optimizer reversely updates the network parameters of the SEGAN structure according to the time domain loss value, the frequency domain loss value, and the first preset learning rate to obtain a pre-trained speech enhancement module. It can be understood that the speech sample used for training is first pre-processed. Specifically, the pre-processing includes: sampling rate adjustment, audio frame processing (frame length setting processing, frame shift processing), and further, multiple types of environmental noise are selected, and the environmental noise includes but is not limited to: office noise, street noise, cafe noise, etc. Then, after the speech sample preprocessing and the environmental noise addition processing, the second noisy speech sample is obtained. For example, first, a clean speech sample after preprocessing is obtained (the sampling rate is 16kHz, the frame length is set to 25ms, and the frame shift is 10ms), and then the clean speech sample is subjected to noise addition processing. Specifically, the above-mentioned environmental noise is added to the clean speech sample according to a preset signal-to-noise ratio (-5dB to 15dB) to generate a second noisy speech sample, and the second noisy speech sample is input into the SEGAN structure, and the encoder of the SEGAN structure is trained (a 6-layer one-dimensional convolutional network, each layer has a convolution kernel size of 31, a step size of 2, and the number of channels increases from network layer to network layer: 16→32→64→128→256→512) to extract deep time-frequency features from the second noisy speech sample, and then enters the decoder of the SEGAN structure (a symmetrical 6-layer transposed convolutional network, and the number of channels decreases layer by layer: 512→256→128→64→32→16→1), and the decoder is trained to reconstruct the speech waveform after speech enhancement. Among them, jump connections are added between the corresponding layers of the encoder and the decoder to further retain high-frequency detail information and avoid excessive smoothing of speech components during the noise reduction process. Finally, the decoder of the SEGAN structure outputs an enhanced speech waveform (process speech sample) of the same length as the second noisy speech sample, and then calculates the L1 time domain loss value and the frequency domain STFT loss value between the process speech sample and the previous clean speech sample, wherein the point-by-point absolute difference between the process speech sample and the clean speech sample is calculated as the L1 time domain loss value. In this way, by calculating the time domain loss value, the process speech sample can be constrained to align with the clean speech sample in the time domain, retaining the global structure of the speech; further, the process speech sample and the clean speech sample are subjected to short-time Fourier transform, and the L1 difference of the spectral amplitude is calculated, that is, the frequency domain STFT loss value is obtained. In this way, by calculating the frequency domain loss value, the spectral fidelity of the speech sample can be enhanced and frequency domain distortions such as music noise can be reduced. Finally, the total loss of the pre-trained speech enhancement module is calculated. ,in,
[0061] ;
[0062] in, represents the time domain loss value, represents the loss value in the frequency domain;
[0063] ;
[0064] ;
[0065] Where N represents the total number of sampling points of the second noisy speech sample. For example, a 3-second 16kHz speech segment corresponds to 48,000 sampling points; Indicates The time domain waveform value of the pure speech sample at sampling points, Indicates The time domain waveform value of the process speech sample of sampling points (the value of the speech after noise reduction), Represents absolute value operation, which means calculating the absolute difference between two waveform values; Indicates the total number of time frames of the short-time Fourier transform. For example, 3 seconds of speech is divided into 25ms frame length and 10ms frame shift, resulting in 299 frames. The total number of frequency components (i.e., the number of frequency bands) in each time frame. For example, when using a 512-point FFT, Indicates the total number of frequency components in each time frame. For example, when using a 512-point FFT, =257 (symmetry preserves half the frequency), represents the time frame, Indicates frequency, Represents a clean speech sample in time frame ,frequency The spectrum amplitude value at (amplitude only, without phase), Represents the process speech sample in time frame ,frequency The spectrum amplitude value at Represents a double accumulation summation of the amplitude differences of all time frames and frequency components.
[0066] It should be noted that in the L1 time domain loss, N refers to the total number of sampling points of a single speech segment, not the number of speech segments. For example, if the input is 3 seconds of speech (48,000 points), the difference of each point is calculated. In the calculation process of the frequency domain STFT loss value, only the amplitude (amplitude) difference of the spectrum is calculated, and the phase information is ignored, because the human ear is not sensitive to phase changes, and amplitude fidelity is more important for speech quality. Finally, the frequency domain loss weight is low (0.1) to avoid excessive impact of frequency domain optimization on time domain waveform reconstruction and ensure the naturalness of speech.
[0067] In this way, the L1 time domain loss is calculated to ensure that the time domain waveform of the process speech sample is globally aligned with the clean speech sample, and the speech structure is preserved; the frequency domain STFT loss is calculated to constrain the spectral amplitude of the process speech sample to be close to the clean speech sample, reducing frequency domain distortion such as music noise. The final designed total loss balances waveform fidelity and spectral quality through weighted fusion of time domain and frequency domain losses, ultimately improving the speech enhancement effect.
[0068] Then, the Adam optimizer reversely updates the network parameters of the SEGAN structure according to the total loss and the first preset learning rate 1e-4 to obtain the pre-trained speech enhancement module. It should be noted that in the training of the initial speech enhancement module, the goal is to enable it to learn basic noise reduction capabilities without involving speaker feature preservation.
[0069] After the initial speech enhancement module is trained, the initial shared encoder and initial speaker recognition module are further trained. The specific steps are as follows:
[0070] Freeze the pre-trained speech enhancement module, and use the clean speech sample to perform speaker feature extraction training on the initial shared encoder and the initial speaker recognition module connected to the initial shared encoder to obtain the pre-trained shared encoder and the pre-trained speaker recognition module. Specifically, the clean speech sample is input into the initial shared encoder to train the initial shared encoder to learn to extract speaker sensitive features to obtain the pre-trained shared encoder, and output the second speaker sensitive feature; the second speaker sensitive feature is input into the speaker recognition network to train the speaker recognition network to extract the speaker embedding vector from the second speaker sensitive feature; the AM-Softmax loss function is used to calculate the classification loss value between the speaker embedding vector and the class center vector, wherein the class center vector is the vector distribution information of the speaker embedding vector of each clean speech sample obtained by initialization; the network parameters of the speaker recognition network are reversely updated by the Adam optimizer according to the classification loss value and the second preset learning rate to obtain the pre-trained speaker recognition module. It is understandable that the clean speech samples are subjected to data preprocessing and speaker labeling. Specifically, the original speech data without noise pollution is selected as the clean speech samples (16kHz sampling rate, 3-second segments), and each sample is associated with a unique speaker ID label. Each training batch contains P speakers × M voices (for example: P=5 speakers, M=10 voices for each speaker) to enhance the robustness of the model to intra-speaker differences. The clean speech samples are uniformly sampled at 16kHz and processed in frames (frame length 25ms, frame shift 10ms). The speech waveform is amplitude normalized (such as quantile normalization) to obtain clean speech samples with speaker ID labels to avoid training instability. The steps for training the initial shared encoder and the initial speaker recognition module using speaker ID labels are as follows:
[0071] First, the parameters of the pre-trained SEGAN module are fixed and do not participate in forward propagation or back propagation. Then, the clean speech sample skips the pre-trained speech enhancement module and directly enters the initial shared encoder. The structure of the initial shared encoder is a bidirectional GRU (Gated Recurrent Units) network (hidden dimension 256). Combined with the attention mechanism, the initial shared encoder is trained to adaptively focus on the speaker-sensitive area, and then the output result is a 512-dimensional feature vector. The output feature vector is the second speaker-sensitive feature, which specifically includes the speaker identity information. The initial speaker recognition module that enters is a dynamic convolutional network, which specifically includes 4 series-connected dynamic convolutional modules. The convolution kernel parameters of each module are dynamically generated by the speaker embedding vector through MLP (Multilayer Perceptron), and the convolution kernel parameters are no longer fixed. MLP structure: FC(512→256)→ReLU→FC(256→K×C), and the output is reshaped into the convolution kernel parameters (K is the kernel size, C is the number of channels). This allows the feature extraction process to adapt to the acoustic characteristics of different speakers. Each dynamic convolution module is equipped with Layer Normalization and PReLU activation functions to further enhance the discriminability of features. Dynamic convolution operation: convolution of input features (second speaker sensitive features) → Layer Normalization → PReLU activation, and then enter the classification layer. The fully connected layer maps the features of the dynamic convolution output to the speaker ID category space to obtain the speaker embedding vector.
[0072] Furthermore, the classification loss value AM-Softmax Loss between the speaker embedding vector and the class center vector is calculated. In this process, the speaker embedding vector involved is obtained by the output of the dynamic convolution layer (512-dimensional feature vector), and the class center vector is the reference vector of each speaker class (dynamically updated during the initial speaker recognition module training process). Therefore, the cosine similarity score between the speaker embedding vector and the class center vector is calculated first:
[0073] ;
[0074] in, represents the speaker embedding vector (normalized), Indicates The class center vector of the class (learned during training) is the angle between the pure speech sample and the center of the corresponding class.
[0075] Therefore, the specific AM-Softmax Loss is calculated as follows:
[0076] ;
[0077] in, represents the cosine similarity score between the target speaker’s speaker embedding vector and the class center vector, Represents the interval parameter (usually 0.2), which is used to increase the difference between classes. Represents the speaker category index, which is used to refer to different speaker categories. It can distinguish the category identifiers in related calculations such as the class center vector corresponding to different speaker categories. It represents the scaling factor (usually set to 30), which is used to enhance the distinguishability of classification boundaries. It can be seen that by increasing the distance between classes and reducing the distance within classes, the feature discriminability is enhanced.
[0078] Furthermore, in each training batch, samples with high classification difficulty (i.e., pure speech samples with low prediction confidence) are automatically screened and their losses are calculated first to accelerate the model's learning of difficult cases. In addition, before calculating AM-SoftmaxLoss, the 512-dimensional feature vector output by the pre-trained shared encoder is L2 normalized to prevent feature amplitude differences from affecting the classification effect.
[0079] In getting Afterwards, cosine annealing is used to balance the convergence speed and training stability, and the optimizer is used to reversely update the network parameters of the speaker recognition network according to the second preset learning rate 1e-3 to obtain a pre-trained speaker recognition module.
[0080] In this embodiment, after the training of the first three modules is completed, end-to-end joint training and fine-tuning are performed. Specifically, the first noisy speech sample is a speech sample carrying a real noise sample label and a real speaker ID label. The initial speech enhancement joint system is trained and fine-tuned using the first noisy speech sample. Specifically, the first noisy speech sample is input into a pre-trained speech enhancement module of the initial speech enhancement joint training system, and the first noisy speech sample is subjected to noise reduction processing by the pre-trained speech enhancement module, and a first enhanced speech sample after speech enhancement is output, and the speech enhancement loss is determined, that is, .
[0081] Furthermore, the first enhanced speech sample is input into the pre-trained shared encoder, and feature extraction and normalization are performed by the shared encoder to perform subsequent speaker recognition and adversarial training processes. The shared encoder designed by the present invention adopts a bidirectional GRU structure, and its hidden layer dimension is 256. Attention mechanisms are configured at the input and output ends of the GRU, respectively, to adaptively capture speaker-sensitive areas in the speech. This layer ultimately outputs a 512-dimensional feature vector, which contains both the audio quality information required for speech enhancement and the identity features required for speaker recognition, thereby realizing information sharing and feature reuse of the two tasks.
[0082] Step S12: extracting features of the speaker sensitive area of the first enhanced speech sample through the pre-trained shared encoder to obtain a first speaker sensitive feature, and inputting the first speaker sensitive feature into the pre-trained speaker recognition module, so as to perform speaker ID classification on the first speaker sensitive feature through the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculate the speaker recognition loss.
[0083] In this embodiment, speaker sensitive features are extracted from the first enhanced speech sample by a pre-trained shared encoder to obtain the first speaker sensitive features of the enhanced speech sample, wherein the extraction process of the first speaker sensitive features is consistent with the feature extraction method of the training process of the aforementioned public pre-trained shared encoder, which will not be described in detail. Then, it is divided into two branches, wherein the first speaker sensitive features are input into the pre-trained speaker recognition module to predict the speaker ID, wherein the speaker recognition process is the same as the speaker recognition process of the training process of the aforementioned public pre-trained speaker recognition module, which will not be described in detail.
[0084] In this embodiment, the speaker recognition loss is determined based on the speaker ID classification prediction result and the speaker ID label. It can be understood that, as in the above-mentioned loss obtaining process of the pre-trained speaker recognition module, the speaker ID classification prediction result predicted this time and the pre-marked speaker ID label are used to calculate the loss and determine the speaker recognition loss. .
[0085] Step S13: The adversarial training discriminator of the initial speech enhancement joint training system is used to judge whether the first speaker sensitive feature has noise, and the adversarial loss is adjusted according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss to obtain the trained target speech enhancement joint training system.
[0086] In this embodiment, another branch is to input the first speaker sensitive feature into the adversarial training discriminator, and use the adversarial training discriminator to determine whether the current speech sample where the first speaker sensitive feature is located is a noisy speech sample, and conduct adversarial training, wherein the goal of this fine-tuning (adversarial training) is to force the first enhanced speech sample generated by the pre-trained speech enhancement module to be indistinguishable from the clean speech sample in the feature space, thereby ensuring that the speaker characteristics are retained during the enhancement process. In order to achieve this goal, it is necessary to design a discriminator to determine the source of the feature (enhanced speech or clean speech), and optimize the pre-trained speech enhancement module through the adversarial training mechanism. Specifically, if in the judgment result output by the adversarial training discriminator, if the noise judgment result of the first noisy speech sample is a noisy speech judgment result, then the adversarial loss is increased based on the gradient reversal processing; if the noise judgment result of the first noisy speech sample is a non-noisy speech judgment result, then the adversarial loss is reduced based on the gradient reversal processing. Specifically, if the judgment result is that the current speech sample is a noisy speech judgment result, indicating that the binary classification result is consistent with the real noisy speech result, then due to the existence of the gradient reversal layer (GradientReversal Layer), the gradient direction is reversed (multiplied by -1). Correspondingly, the gradient direction received by the pre-trained enhancement module is: increase the adversarial loss (that is, it is hoped that the adversarial training discriminator will misjudge the enhanced speech features as pure speech). The final fine-tuning result is to control the pre-trained enhancement module to adjust its internal parameters so that the generated first speaker sensitive features are closer to the pure speech features, thereby retaining more speaker information. If the judgment result is that the current speech sample is a noise-free speech judgment result, then after the gradient is reversed, the gradient direction received by the pre-trained enhancement module is: reduce the adversarial loss (that is, the discriminator has been successfully deceived). The final fine-tuning result is that the pre-trained enhancement module does not need to significantly adjust its internal parameters, maintains the current noise reduction strategy (at this time the noise reduction effect is already good, and the speaker features are not excessively lost), and the current adversarial loss is output at this time. .
[0087] It should be noted that during the training process of the adversarial training discriminator, step-by-step updates are performed, that is, the adversarial discriminator parameters are updated every 5 batches, and the adversarial discriminator is frozen for the rest of the time. The overall learning rate of the system is adjusted to the initial lr=1e-5 for the pre-trained speech enhancement module, which decays by 1% per epoch; the initial lr=1e-3 for the pre-trained speaker recognition module is scheduled with cosine annealing.
[0088] In this way, through the above mechanism, the system forces the enhancement module to generate enhanced speech that is indistinguishable from pure speech features. The discriminator continuously optimizes to distinguish the features of the two, forming a game process, so that the pre-trained enhanced speech retains enough speaker information while reducing noise, making it impossible for the adversarial training discriminator to accurately classify.
[0089] In this embodiment, a multi-task loss function is constructed based on the speech enhancement loss, the speaker recognition loss, the adversarial loss and the corresponding weight coefficients; based on the multi-task loss function, a step-by-step update strategy is performed on the system parameters of the initial speech enhancement joint training system to balance the speech enhancement task objectives and the speaker recognition task, thereby obtaining a trained target speech enhancement joint training system. It can be understood that the multi-task loss function of the fine-tuned system is constructed based on the above three losses and the corresponding weight coefficients, so the multi-task loss function is expressed as follows:
[0090] ;
[0091] in, represents the speech enhancement loss, represents the weight of the speech enhancement task, which is 0.5; represents the speaker recognition loss, The speaker recognition task weight is 1.0. Represents resistance to loss, represents the adversarial loss weight, which is -0.2, where the negative sign indicates that the generator is optimized by gradient reversal.
[0092] In this way, the multi-task loss function obtained by the weighted sum formula coordinates the conflicting goals of speech enhancement and speaker recognition, forcing the enhanced speech features to be aligned with the pure speech; the staged weight allocation (such as speaker recognition has the highest weight) ensures that key tasks converge first.
[0093] Reference Figure 2 As shown, the present invention adopts a three-stage progressive training strategy to perform speech enhancement + speaker recognition training on the system:
[0094] Phase 1: Pre-training of speech enhancement module (1) Input noisy speech signal and perform preliminary noise reduction through SEGAN structure; (2) Calculate the time domain L1 loss and frequency domain loss between the enhanced output and the original clear speech; (3) Use Adam optimizer to update network parameters, and the initial learning rate is set to 1e-4; (4) When the validation set loss no longer decreases, save the model parameters and enter the next phase.
[0095] Phase 2: Speaker recognition module pre-training (1) Freeze the parameters of the speech enhancement module trained in the first phase; (2) Input clear speech, extract features through the speaker recognition network, and generate speaker representation; (3) Use the AM-Softmax loss function to calculate the classification loss; (4) Perform online difficult sample mining in each mini-batch; (5) Use the Adam optimizer to update the network parameters, and the initial learning rate is set to 1e-3.
[0096] The third stage: end-to-end joint optimization (1) Load the model parameters pre-trained in the first two stages; (2) Input noisy speech and the corresponding speaker label at the same time; (3) Forward propagation process: noisy speech → speech enhancement → feature extraction → speaker recognition; (4) Calculate the overall loss: including speech enhancement loss, speaker recognition loss and adversarial loss; (5) Use different learning rates to update the parameters of each module separately; (6) After training every 5 batches, update the parameters of the adversarial discriminator.
[0097] Specifically, the implementation process of the dynamic convolution module is as follows: feature extraction preparation (1) first obtain the 512-dimensional initial feature vector through the shared coding layer; (2) input the feature vector into the MLP network to generate dynamic convolution kernel parameters; (3) according to the preset convolution kernel size K and number of channels C, reshape the MLP output into the convolution kernel shape.
[0098] The dynamic convolution operation (1) uses the generated convolution kernel parameters to perform a convolution operation on the input features; (2) applies Layer Normalization to normalize the features; (3) enhances the nonlinear expression ability of the features through the PReLU activation function; (4) connects multiple dynamic convolution modules in series to form a deep feature extraction network.
[0099] The implementation process of the adversarial training mechanism is as follows: Discriminator training (1) collects a batch of shared coding layer features, including features from enhanced speech and pure speech; (2) inputs the features into the discriminator network to obtain binary classification prediction results; (3) calculates the focal loss and updates the discriminator parameters. Generator training (1) applies a gradient reversal layer on the feature transfer path; (2) calculates the adversarial loss and combines it with other loss terms; (3) back-propagates to update the parameters of the speech enhancement module.
[0100] The above technical solutions have achieved significant performance improvements in low signal-to-noise ratio scenarios: in a 5dB noise environment, the speaker recognition accuracy is improved by more than 15% compared to traditional methods; the speaker verification EER index of the enhanced speech is reduced by about 20%; and the training convergence time is reduced by about 40%. In particular, the speaker recognition accuracy can still be maintained at more than 80% in a -5dB noise environment, and it shows good robustness to various noise types.
[0101] In an intelligent conference system, participants speak in a noisy conference room environment, and the background noise includes keyboard tapping, air conditioning noise, and multiple people talking (the signal-to-noise ratio may be as low as -5dB). The system needs to enhance the voice signal in real time and accurately identify the identities of different speakers in order to automatically generate meeting minutes or respond to voice commands. The performance of the traditional series solution significantly decreases in this scenario due to error accumulation and target conflict, while the following joint training of this solution realizes the joint training of dual tasks.
[0102] Phase 1: Speech Enhancement Module (SEGAN) Pre-training: Clean speech library: Collect clean speech (no noise interference) of multiple speakers in the conference room scene, each speech is 3 seconds, 16kHz sampling rate. Noise library: Collect typical conference room noise (keyboard sound, air conditioner sound, background human voice), add it to the clean speech according to the signal-to-noise ratio (-5dB, 15dB), and generate noisy training samples. The training process includes input: noisy speech waveform (such as a speech clip containing keyboard sound) output: enhanced speech waveform (the target is the corresponding clean speech). Loss function: Time domain L1 loss: constrain the waveform of enhanced speech to align with that of clean speech; frequency domain STFT loss (weight 0.1): reduce spectral distortion (such as residual low-frequency noise of air conditioner). Training effect: After pre-training, the enhancement module can effectively suppress the keyboard tapping sound and retain the speech content.
[0103] Phase 2: Speaker Recognition Module (SD) Pre-training: Clean speech library: Use the clean speech data from Phase 1, and each speaker provides 10 speech samples. Label: Label each speech with the speaker ID (such as "Speaker A" and "Speaker B"). The training process includes input: clean speech waveform (such as the clean speech of speaker A). Shared encoder: Extract 512-dimensional speaker-sensitive features (such as pitch and intonation features) through bidirectional GRU. Dynamic convolutional network: Dynamically generate convolution kernels based on features and adaptively extract speaker identity information. Loss function: AM-Softmax Loss (Margin=0.2), increase the inter-class differences between different speakers. Training effect: The model can accurately distinguish different speakers (such as A and B) in the conference room with an accuracy rate of 95%.
[0104] Phase 3: End-to-end joint fine-tuning and adversarial training: Noisy speech: simulate conference room noise scenarios (SNR -5dB to 15dB). Clean speech and label: original speech and speaker ID paired with noisy speech. The training process includes forward propagation: noisy speech (such as speech with background conversation) → speech enhancement module → enhanced speech; enhanced speech → shared encoder → 512-dimensional features; features are input simultaneously: speaker recognition module → predict speaker ID (such as "speaker C"); adversarial discriminator → determine the source of features (enhanced speech or clean speech). Loss calculation: speech enhancement loss (weight 0.5): ensure that the enhanced speech is close to the clean speech; speaker recognition loss (weight 1.0): improve the accuracy of identity recognition; adversarial loss (weight -0.2): force the enhanced speech features to be indistinguishable from the clean speech features through gradient reversal.
[0105] Among them, the adversarial training mechanism is as follows: Clean speech path: clean speech → shared encoder → feature → adversarial discriminator (as a "positive sample"). During the training process, it is necessary to maintain a dynamic balance, that is, update the adversarial discriminator every 5 batches to prevent it from being too strong and causing the generator to crash. The final training effect is noise reduction capability: under -5dB noise, the PESQ (speech quality score) of enhanced speech is increased from 1.2 to 3.5; speaker recognition: the recognition accuracy is increased from 65% of the traditional solution to 85%; real-time: after the dynamic convolutional network and shared encoder are optimized, the inference delay is less than 50ms, meeting the real-time needs of the meeting.
[0106] Scenario verification: robustness test under sudden noise; test conditions: simulate high-intensity noise that suddenly appears in the conference room (such as the sound of chair dragging, signal-to-noise ratio -10dB). System response: the speech enhancement module quickly suppresses sudden noise and retains the high-frequency details of the speech; the shared encoder focuses on the speaker's sensitive areas (such as glottal pulses) through the attention mechanism; the dynamic convolution network adaptively adjusts parameters to accurately identify the speaker's identity. Results: speech intelligibility is improved from 0.4 to 0.8; the speaker recognition accuracy remains above 80%. Error accumulation elimination: joint training avoids distortion of the enhancement module from being passed to downstream tasks; target conflict coordination: adversarial training ensures that identity features (such as pitch trajectory) are retained during noise reduction; computational efficiency: dynamic convolution and step-by-step update strategies reduce resource consumption; strong generalization: high robustness is maintained for unseen noise types (such as coffee machine sounds). This solution provides high-precision, low-latency speech processing capabilities in noisy environments for intelligent conference systems, significantly improving user experience.
[0107] It can be seen that the present application discloses a speech enhancement training method based on speaker perception, comprising: inputting a first noisy speech sample into an initial speech enhancement joint training system so that a pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, outputs a first enhanced speech sample to determine the speech enhancement loss, and inputs the first enhanced speech sample into a pre-trained shared encoder of the initial speech enhancement joint training system; extracting features of a speaker sensitive area of the first enhanced speech sample through the pre-trained shared encoder to obtain a first speaker sensitive feature, and inputting the first speaker sensitive feature into the pre-trained shared encoder. The speaker recognition module is trained so that the speaker ID classification of the first speaker sensitive feature is performed through the pre-trained speaker recognition module, the speaker ID classification prediction result is obtained, and the speaker recognition loss is calculated; the adversarial training discriminator of the initial speech enhancement joint training system is used to judge whether the first speaker sensitive feature has noise, and the adversarial loss is adjusted according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss, and obtain the trained target speech enhancement joint training system. It can be seen that by jointly training the speech enhancement module, the shared encoder and the speaker recognition module, the negative impact of the acoustic distortion introduced by the speech enhancement module in the traditional serial system on the downstream speaker recognition task is avoided; the adversarial training mechanism further constrains the output of the speech enhancement module to ensure that the enhanced speech retains the speaker sensitive features and reduces error propagation. Furthermore, through adversarial training and multi-task loss design, the target conflict between the two tasks (speech enhancement task and speaker recognition task) is coordinated to ensure that the enhancement module retains the speaker identity information while reducing noise.
[0108] Reference Figure 3 As shown, the present invention also discloses a speech enhancement training device based on speaker perception, comprising:
[0109] The speech enhancement module 11 is used to input the first noisy speech sample into the initial speech enhancement joint training system so that the pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, output a first enhanced speech sample to determine the speech enhancement loss, and input the first enhanced speech sample into the pre-trained shared encoder of the initial speech enhancement joint training system;
[0110] A speaker recognition module 12 is used to perform feature extraction of a speaker sensitive area on the first enhanced speech sample through the pre-trained shared encoder to obtain a first speaker sensitive feature, and input the first speaker sensitive feature into the pre-trained speaker recognition module so as to perform speaker ID classification on the first speaker sensitive feature through the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculate a speaker recognition loss;
[0111] The model training module 13 is used to judge whether there is noise in the first speaker sensitive feature through the adversarial training discriminator of the initial speech enhancement joint training system, and adjust the adversarial loss according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss to obtain the trained target speech enhancement joint training system.
[0112] It can be seen that the present application discloses inputting a first noisy speech sample into an initial speech enhancement joint training system so that the pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, outputs a first enhanced speech sample, determines the speech enhancement loss, and inputs the first enhanced speech sample into the pre-trained shared encoder of the initial speech enhancement joint training system; extracts the features of the speaker sensitive area of the first enhanced speech sample by the pre-trained shared encoder to obtain a first speaker sensitive feature, and inputs the first speaker sensitive feature into a pre-trained speaker recognition module so that the first speaker sensitive feature is classified by the speaker ID by the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculates the speaker recognition loss; judges whether there is noise in the first speaker sensitive feature by the adversarial training discriminator of the initial speech enhancement joint training system, and adjusts the adversarial loss according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss to obtain the trained target speech enhancement joint training system. It can be seen that by jointly training the speech enhancement module, shared encoder and speaker recognition module, the negative impact of acoustic distortion introduced by the speech enhancement module in the traditional serial system on the downstream speaker recognition task is avoided; the adversarial training mechanism further constrains the output of the speech enhancement module to ensure that the enhanced speech retains the speaker's sensitive characteristics and reduces error propagation. Furthermore, through adversarial training and multi-task loss design, the goal conflict between the two tasks (speech enhancement task and speaker recognition task) is coordinated to ensure that the enhancement module retains the speaker's identity information while reducing noise.
[0113] Furthermore, the present application also discloses an electronic device. Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0114] Figure 4 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the speech enhancement training method based on speaker perception disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0115] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0116] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0117] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0118] Among them, the operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, so as to realize the operation and processing of the processor 21 on the massive data 223 in the memory 22, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the speech enhancement training method based on speaker perception performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks. In addition to data transmitted from an external device received by the electronic device, the data 223 can also include data collected by its own input and output interface 25.
[0119] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disclosed method for speech enhancement training based on speaker perception is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0120] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0121] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly with hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory RAM (Random Access Memory), memory, read-only memory ROM (Read Only Memory), electrically programmable EPROM (Electrically Programmable Read Only Memory), electrically erasable programmable EEPROM (ElectricErasable Programmable Read Only Memory), register, hard disk, removable disk, CD-ROM (CompactDisc-Read Only Memory), or any other form of storage medium known in the technical field.
[0122] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0123] The scheme provided by the present invention is introduced in detail above. Specific examples are used in this article to illustrate the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A speech enhancement training method based on speaker perception, characterized in that: include: Inputting a first noisy speech sample into an initial speech enhancement joint training system so that a pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, outputs a first enhanced speech sample to determine a speech enhancement loss, and inputs the first enhanced speech sample into a pre-trained shared encoder of the initial speech enhancement joint training system; Extracting features of a speaker sensitive area of the first enhanced speech sample by using the pre-trained shared encoder to obtain a first speaker sensitive feature, and inputting the first speaker sensitive feature into a pre-trained speaker recognition module, so as to perform speaker ID classification on the first speaker sensitive feature by using the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculate a speaker recognition loss; The adversarial training discriminator of the initial speech enhancement joint training system is used to judge whether the first speaker sensitive feature has noise, and the adversarial loss is adjusted according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss and the adversarial loss to obtain the trained target speech enhancement joint training system.
2. The method for speech enhancement training based on speaker perception according to claim 1, characterized in that: The method of updating the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss, and the adversarial loss to obtain a trained target speech enhancement joint training system includes: Constructing a multi-task loss function based on the speech enhancement loss, the speaker recognition loss, the adversarial loss and corresponding weight coefficients; Based on the multi-task loss function, a step-by-step update strategy is performed on the system parameters of the initial speech enhancement joint training system to balance the speech enhancement task objectives and the speaker recognition task, thereby obtaining a trained target speech enhancement joint training system.
3. The method for speech enhancement training based on speaker perception according to claim 1, characterized in that: The step of adjusting the adversarial loss according to the error between the noise judgment result and the noise sample label of the first noisy speech sample comprises: If the noise determination result of the first noisy speech sample is a noisy speech determination result, adding adversarial loss based on gradient reversal processing; If the noise determination result of the first noisy speech sample is a noise-free speech determination result, the adversarial loss is reduced based on the gradient inversion process.
4. The method for speech enhancement training based on speaker perception according to claim 1, characterized in that: The first noisy speech sample is a speech sample carrying a real noise sample label and a real speaker ID label; Accordingly, the calculation of speaker recognition loss includes: A speaker recognition loss is determined based on the speaker ID classification prediction result and the speaker ID label.
5. The method for speech enhancement training based on speaker perception according to any one of claims 1 to 4, characterized in that: Also includes: Freezing the initial speaker recognition module, and performing basic noise reduction training on the initial speech enhancement module using the second noisy speech sample to obtain a pre-trained speech enhancement module; The pre-trained speech enhancement module is frozen, and the initial shared encoder and the initial speaker recognition module connected to the initial shared encoder are trained for speaker feature extraction using clean speech samples to obtain a pre-trained shared encoder and a pre-trained speaker recognition module.
6. The method for speech enhancement training based on speaker perception according to claim 5, characterized in that: The method of performing basic noise reduction training on the initial speech enhancement module using the second noisy speech sample to obtain a pre-trained speech enhancement module includes: Acquire a clean speech sample, and perform noise adding processing on the clean speech sample to obtain a second noisy speech sample; Inputting the second noisy speech sample into the SEGAN structure to perform a primary noise reduction process on the second noisy speech sample to obtain a process speech sample; The time domain loss value and the frequency domain loss value between the process speech sample and the clean speech sample are calculated so that the Adam optimizer reversely updates the network parameters of the SEGAN structure according to the time domain loss value, the frequency domain loss value, and the first preset learning rate to obtain a pre-trained speech enhancement module.
7. The method for speech enhancement training based on speaker perception according to claim 5, characterized in that: The method of using the clean speech sample to perform speaker feature extraction training on the initial shared encoder and the initial speaker recognition module connected to the initial shared encoder to obtain a pre-trained shared encoder and a pre-trained speaker recognition module includes: Inputting the clean speech sample into the initial shared encoder to train the initial shared encoder to learn to extract speaker-sensitive features, so as to obtain a pre-trained shared encoder, and output a second speaker-sensitive feature; Inputting the second speaker sensitive feature into a speaker recognition network to train the speaker recognition network to extract a speaker embedding vector from the second speaker sensitive feature; Calculating the classification loss value between the speaker embedding vector and the class center vector using the AM-Softmax loss function, wherein the class center vector is obtained by initializing the vector distribution information of the speaker embedding vector of each of the clean speech samples; The network parameters of the speaker recognition network are reversely updated according to the classification loss value and the second preset learning rate by an Adam optimizer to obtain a pre-trained speaker recognition module.
8. A speech enhancement training device based on speaker perception, characterized in that: include: a speech enhancement module, configured to input a first noisy speech sample into an initial speech enhancement joint training system, so that a pre-trained speech enhancement module of the initial speech enhancement joint training system performs denoising on the first noisy speech sample, output a first enhanced speech sample to determine a speech enhancement loss, and input the first enhanced speech sample into a pre-trained shared encoder of the initial speech enhancement joint training system; a speaker recognition module, configured to perform feature extraction of a speaker sensitive area on the first enhanced speech sample through the pre-trained shared encoder to obtain a first speaker sensitive feature, and input the first speaker sensitive feature into the pre-trained speaker recognition module, so as to perform speaker ID classification on the first speaker sensitive feature through the pre-trained speaker recognition module to obtain a speaker ID classification prediction result, and calculate a speaker recognition loss; The model training module is used to judge whether the first speaker sensitive feature has noise through the adversarial training discriminator of the initial speech enhancement joint training system, and adjust the adversarial loss according to the error between the noise judgment result and the noise sample label of the first noisy speech sample, so as to update the system parameters of the initial speech enhancement joint training system based on the speech enhancement loss, the speaker recognition loss and the adversarial loss to obtain the trained target speech enhancement joint training system.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the speech enhancement training method based on speaker perception as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the steps of the speech enhancement training method based on speaker perception as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Speech enhancement method based on voiceprint comparison and generative adversarial network
CN109326302A
End-to-end voice enhancement method based on generation of countermeasure network
CN110390950A