An intelligent communication device voice noise reduction method and system based on artificial intelligence
By deploying a noise reduction algorithm with a microphone array and a multi-task learning mechanism on an intelligent communication device, the problem of speech distortion and sound quality degradation in the speech noise reduction process of intelligent communication devices is solved, and effective processing of dialects and improving voice quality is achieved.
Patent Information
- Application Number
- CN202510388403.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Existing intelligent communication devices can easily lead to speech distortion and sound quality degradation during speech noise reduction, especially in low signal-to-noise ratio environments and when processing dialects, the algorithm may over-filter the frequency components in the voice signal, causing the voice to sound unnatural.
Design a speech noise reduction system for intelligent communication equipment based on artificial intelligence. By deploying a microphone array on intelligent communication equipment, it processes noise-free voice data in real time and uses a pre-trained noise reduction algorithm for processing. The algorithm includes a multi-task learning mechanism, which uses deep learning structures to combine with a generative adversarial network, expands the dialect voice database through the scene evolution layer, the tone change layer and the noise mixing layer, retains the dialect features, and optimizes the noise reduction effect through the signal-to-noise ratio estimator and cIRM model.
Effectively process dialects, retain dialect features, improve voice quality, reduce voice distortion and sound quality degradation, adapt to low signal-to-noise ratio environments, achieve real-time and high-quality voice enhancement, and improve user experience.
Smart Images

Figure CN119905101B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech noise reduction processing. More specifically, the present invention relates to a speech noise reduction method and system for intelligent communication devices based on artificial intelligence. Background Art
[0002] Speech distortion and reduced sound quality are common problems in the process of speech noise reduction for intelligent communication devices. Especially in the special scenario of combining dialects, the problem may be further amplified. During the noise reduction process, especially in a low signal-to-noise ratio environment, some algorithms may over-filter certain frequency components in the speech signal, resulting in unnatural speech (such as "robotic voice" or "hollowness"). When users use dialects, this problem may be more prominent because the speech characteristics of dialects (such as pitch, timbre, harmonic structure) are significantly different from those of Mandarin or other standard languages.
[0003] This is because most noise reduction algorithms are trained based on large-scale datasets of Mandarin, English, or other standard languages. These datasets often lack sufficient dialect samples, resulting in insufficient feature modeling of non-standard speech by the algorithms.
[0004] In addition, noise reduction algorithms usually distinguish speech and noise through spectral analysis. However, in a low signal-to-noise ratio environment, the harmonic or detailed components of the speech signal may overlap with the spectral characteristics of the noise. Some unique frequency components in dialects (such as high-frequency tones or low-frequency guttural sounds) are more likely to be misjudged as noise and filtered out, and may not be able to well adapt to this rapidly changing speech pattern, resulting in loss of details, sounding blurred or incomplete, and affecting the accurate transmission of semantics.
[0005] In view of this, the present invention proposes a speech noise reduction system for intelligent communication devices based on artificial intelligence to solve the above problems. Summary of the Invention
[0006] To overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A speech noise reduction system for intelligent communication devices based on artificial intelligence, including an intelligent communication device, on which a microphone array is deployed for collecting real-time noisy speech data of a user;
[0007] The collected noisy speech data is input into the edge layer, where a pre-trained noise reduction algorithm is deployed, and the noise-reduced speech data is obtained through real-time processing and output;
[0008] Among them, the hierarchical deployment of the noise reduction algorithm specifically includes: pre-building a dialect speech database, adding a scene evolution layer, a sound change layer, and a noise mixing layer to the conditional adversarial network cGAN to expand the dialect speech database, forming data samples, and performing three-dimensional annotation on each speech data in the data samples;
[0009] Design the input, output, and multi-task learning mechanism based on an end-to-end speech enhancement model, and use data samples for pre-training. The multi-task learning mechanism uses a combination of a deep learning structure and a generative adversarial network (GAN) for speech reconstruction. A shared encoder is set in the multi-task learning mechanism for joint tasks R1, R2, and R3. Task R1 designs an IRM estimator based on U-Net and uses the IRM estimator to train a cIRM model for noise suppression, and a lightweight signal-to-noise ratio estimator is added. Task R2 jointly performs retention supervision of dialect features based on a feature embedding layer, a tone protection module, a timbre protection module, and a semantic-assisted noise reduction module. Task R3 reconstructs speech through a generative adversarial network.
[0010] Preferably, the method of deploying a microphone array on an intelligent communication device for collecting user speech data includes:
[0011] Select a suitable array and spacing according to the size of the intelligent communication device and the computing resources of the edge layer. The microphone selects a high sampling rate and a wide frequency range, is equipped with an analog-to-digital converter with a high dynamic range, and is connected to the edge layer of the intelligent communication device using an interface.
[0012] Preferably, the method for collecting the data samples includes;
[0013] Record the user's dialect speech data in a low signal-to-noise ratio environment to construct a dialect speech database covering the low signal-to-noise ratio environment;
[0014] The architecture of the conditional generative adversarial network (cGAN) includes a generator and a discriminator. Input the speech data in the dialect speech database into the generator. Add a scene evolution layer, a sound variation layer, and a noise mixing layer in the generator. Adopt an alternating training method to train the generator and the discriminator in turn. After the training is completed, fix the discriminator and use the generator to generate new low signal-to-noise ratio dialect speech data. The generated data is used to expand the existing dialect speech database to form data samples;
[0015] Set a three-dimensional blank dataset. The three dimensions include multi-granularity acoustic features, noise features, and semantic features. The multi-granularity acoustic features include phoneme level, pronunciation manner, prosodic features, and timbre of the speech data. The noise features include noise events, noise components, and time-frequency distribution. The semantic features include word-level annotation, sentence-level annotation, and discourse-level annotation; Extract the three-dimensional data of each speech data in the data samples and fill the data into the blank dataset to form the three-dimensional annotation of each speech data.
[0016] Preferably, the method of adding a scene evolution layer, a sound variation layer, and a noise mixing layer in the generator includes:
[0017] The input of the scenario evolution layer is a preset scenario label. According to the scenario label, the corresponding noise segment is extracted from the dialect speech database, and the scenario speech is output through random mixing. Random mixing includes single-scenario isolation and multi-scenario superposition. Single-scenario isolation is to directly output the selected noise segment, and multi-scenario superposition is to linearly or non-linearly superpose the noise segments of multiple scenarios according to the preset weight distribution mechanism for each scenario;
[0018] The probability distribution mechanism means presetting a weight value range for each scenario, and the sum of the weights is less than or equal to 1 when multi-scenario superposition is performed;
[0019] The input of the phonetic variation layer is the output of the speech encoder and the phonetic variation control parameters. According to the phonetic characteristics of the dialect, phonetic variation rules are defined to form a rule-based phonetic variation model; using the dialect speech data with phonetic variation annotations, a sequence-to-sequence model is trained to predict and generate phonetic variations, forming a data-driven phonetic variation model; combining the rule-based phonetic variation model and the data-driven phonetic variation model, and adjusting the phonetic variation parameters to control the degree and type of phonetic variation, and outputting the phonetic feature sequence after phonetic variation, denoted as the phonetically varied speech;
[0020] The input of the noise mixing layer is the scenario speech and the phonetically varied speech, which are mixed according to the specified signal-to-noise ratio, and the speech feature sequence with noise and phonetic variation is output, denoted as the new low signal-to-noise ratio dialect speech data.
[0021] Preferably, the design of the noise reduction algorithm further includes:
[0022] Design a noise reduction algorithm based on an end-to-end speech enhancement model, including input, output, and multi-task learning mechanism. Set the input as the time-frequency spectrogram of the noisy speech. The multi-task learning mechanism uses a deep learning structure for time-frequency processing, combines the generative adversarial network GAN for speech reconstruction, and the output is the time-frequency spectrogram of the denoised clean speech;
[0023] The input noisy speech generates a time-frequency spectrogram through short-time Fourier transform STFT, and the output time-frequency spectrogram of the denoised clean speech reconstructs the time-domain signal through inverse STFT to form the denoised speech data;
[0024] Design the loss function as the weighted sum of the noise reduction loss, speech reconstruction loss, dialect feature loss, and naturalness loss. The noise reduction loss is used to measure the noise reduction effect in a low signal-to-noise ratio environment, the speech reconstruction loss is used to evaluate the perceptual quality of the denoised speech data, the dialect feature loss is used to measure the damage of the model to tones and timbres, and the naturalness loss is used to measure the naturalness of the denoised speech data. Add a loss priority rule, that is, preset two thresholds, satisfying that the speech reconstruction loss is greater than the threshold and the dialect feature loss is less than the threshold;
[0025] Pre-train the noise reduction algorithm using data samples until the value of the loss function meets the requirements or reaches the number of iterations, then stop to obtain the pre-trained noise reduction algorithm and deploy it on the edge layer for processing real-time speech data.
[0026] Preferably, when the multi-task learning mechanism uses a deep learning structure for time-frequency processing, the method of combining the generative adversarial network GAN for speech reconstruction includes:
[0027] Set up a shared encoder to extract deep feature representations of the spectrogram, use a deep learning structure, and output the encoded feature map.
[0028] The preset task-specific heads include task R1, task R2, and task R3.
[0029] Task R1 is used for noise suppression. A lightweight signal-to-noise ratio estimator is added after the shared encoder and before the IRM estimator in task R1 to evaluate the signal-to-noise ratio level of the input speech.
[0030] Task R2 is used for dialect feature retention. By introducing the supervision information of dialect features, key features are retained and not filtered out.
[0031] Task R3 is used for speech reconstruction. The details of the speech are reconstructed through the generative adversarial network, including using the generator of the generative adversarial network to reconstruct the harmonics, tones, and rapidly changing pronunciation details of the speech, and the discriminator enhances the naturalness of the speech; the dialect features output by task R2 are introduced as guidance.
[0032] Preferably, the method for task R1 to be used for noise suppression includes:
[0033] Task R1 designs an IRM estimator based on U-Net. Its input is the feature map output by the shared encoder. The structure includes an encoder, a decoder, and an output layer. The encoder is similar to or the same as the structure of the shared encoder and contains multiple downsampling modules. The decoder is symmetric to the encoder and contains multiple upsampling modules. The encoder and the decoder are jump-connected. The output layer is set as a 1x1 convolutional layer, and the Sigmoid activation function is used to map the feature map to the IRM mask. Among them, the size of the output IRM mask is the same as that of the input STFT spectrogram.
[0034] The loss function is the mean square error or binary cross entropy. The IRM model is trained using the noisy speech and clean speech data until convergence, and the output is the trained IRM model and the predicted IRM mask.
[0035] The output feature map of the shared encoder and the predicted IRM mask are used as the input of the cIRM model. The cIRM estimator is designed based on U-Net and has a similar structure to the IRM model. However, IRM fusion is added after the skip connection. The fusion method is to concatenate the predicted IRM with the output feature map of the encoder as the input of the cIRM encoder. In each upsampling module of the decoder, the predicted IRM is fused with the upsampled feature map. The cIRM contains a real part and an imaginary part. The loss function is set, and the output is the predicted cIRM mask;
[0036] Apply the predicted cIRM mask to the STFT spectrum of the noisy speech to obtain the STFT spectrum of the denoised speech.
[0037] Preferably, a lightweight signal-to-noise ratio estimator is added after the shared encoder and before the IRM estimator in the task R1 to evaluate the signal-to-noise ratio level of the input speech, including:
[0038] Set the input of the signal-to-noise ratio estimator to the feature map output by the shared encoder. The structure includes global average pooling, a fully connected layer, an activation function, and an output layer. The output layer uses a fully connected layer with a single output, and the output is the estimated SNR value;
[0039] Set the loss function as the mean square error to calculate the mean square error between the estimated SNR value and the true SNR value. Input the STFT spectrogram of the noisy speech into the shared encoder, input the output feature map of the shared encoder into the signal-to-noise ratio estimator, calculate the MSE loss between the estimated SNR value and the true SNR value, and use the backpropagation algorithm to update the weights of the signal-to-noise ratio estimator;
[0040] In the IRM estimator, the estimated SNR value is used as an additional input and input into the IRM estimator together with the original input.
[0041] Preferably, the task R2 is used for dialect feature retention. By introducing the supervision information of dialect features, the method for retaining key features without being filtered includes:
[0042] Including a feature embedding layer, a tone protection module, a timbre protection module, and a semantic-assisted noise reduction module;
[0043] The input of the feature embedding layer is the noisy speech data + dialect label; the dialect label is converted into a vector with a fixed dimension, denoted as the embedding vector of the dialect; this embedding vector of the dialect is concatenated with the speech feature vector or fused through an attention mechanism to form a new vector, and the new vector will be used as the input of the subsequent noise reduction algorithm;
[0044] The input of the tone protection module is noisy speech data. Using a pre-trained dialect tone classifier, the output is the tone category probability distribution and tone contour features for each time frame;
[0045] During the noise reduction process, the tone contour features of the output speech are constrained by a loss function to make their similarity with the tone contour features of the tone protection module meet the requirements;
[0046] The input of the timbre protection module is noisy speech data. Using a pre-trained timbre feature extractor, the output is a timbre feature vector with a fixed dimension. In the loss function of the noise reduction process, a term is added to penalize the difference in timbre feature vectors before and after noise reduction;
[0047] The input of semantic-assisted noise reduction is noisy speech data. Using a pre-trained dialect speech recognition model, the output is the text transcription of the speech, phoneme / syllable-level information. According to the speech recognition result, the tone changes, syllable boundaries, and key phonemes are extracted; and protection is carried out respectively according to the recognition results.
[0048] A method for speech noise reduction of an intelligent communication device based on artificial intelligence, including:
[0049] Step S1: Deploy a microphone array on the intelligent communication device to collect the user's real-time noisy speech data;
[0050] Step S2: Design the input, output, and multi-task learning mechanism of the noise reduction algorithm based on an end-to-end speech enhancement model;
[0051] Step S3: Use data samples for pre-training to obtain a trained noise reduction algorithm and deploy it to the edge layer;
[0052] Step S4: Input the collected noisy speech data into the edge layer and output the noise-reduced speech data through real-time processing.
[0053] The technical effects and advantages of the speech noise reduction system of an intelligent communication device based on artificial intelligence according to the present invention:
[0054] 1. By pre-building a dialect speech database and using a conditional adversarial network (cGAN) to augment the data, the system can effectively process various dialects, retain dialect features, and can adapt to complex noise environments of single-scene and multi-scene superposition through the scene evolution layer, which is more in line with the actual usage scenarios.
[0055] 2. Data collection and algorithm training specifically designed for low signal-to-noise ratio environments enable the system to maintain good noise reduction effects in noisy environments. Through three-dimensional annotation (acoustic features, noise features, and semantic features) and multi-task learning mechanisms, the system can simultaneously achieve noise reduction, retain dialect features, and reconstruct high-quality speech. Deploying the pre-trained noise reduction algorithm on the edge layer enables real-time processing, reduces latency, and improves the user experience.
[0056] 3. The phonetic variation layer and dialect feature retention mechanism ensure that the key features of speech, especially tone, timbre, and semantic information, are not damaged during the noise reduction process. Through the signal-to-noise ratio estimator, the noise reduction strategy can be automatically adjusted according to different signal-to-noise ratio levels to avoid speech distortion caused by over-noise reduction. Through processing with the cIRM model (Complex Ideal Ratio Mask), both amplitude and phase information can be processed simultaneously, which is superior to traditional methods that only process amplitude. The three tasks (R1 noise suppression, R2 dialect feature retention, R3 speech reconstruction) work together, and through the shared encoder and specific task head design, a balance between noise reduction and feature retention is achieved. Using semantic information to guide the noise reduction process can more accurately protect important speech components and reduce the loss of semantic information.
[0057] In summary, while effectively reducing noise while retaining dialect characteristics, the problem that traditional noise reduction systems are prone to filtering out key dialect features when processing dialects is solved. At the same time, real-time and high-quality speech enhancement is achieved through multi-task learning and edge computing, improving the usage experience of intelligent communication devices in complex noise environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic structural diagram of a voice noise reduction system for an intelligent communication device based on artificial intelligence according to the present invention;
[0059] Figure 2 It is a schematic hierarchical diagram of a noise reduction algorithm in a voice noise reduction system for an intelligent communication device based on artificial intelligence according to the present invention;
[0060] Figure 3 It is a schematic step diagram of a voice noise reduction method for an intelligent communication device based on artificial intelligence according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] Example 1, please refer to Figure 1 andFigure 2 As shown, the main design content of the voice noise reduction system for an intelligent communication device based on artificial intelligence in this embodiment is as follows:
[0063] Voice distortion and reduced sound quality are common problems in the voice noise reduction process of intelligent communication devices. Especially in the special scenario of combining dialects, the problem may be further amplified. Specifically:
[0064] During the noise reduction process, especially in a low signal-to-noise ratio environment, some algorithms may over-filter certain frequency components in the voice signal, resulting in unnatural-sounding speech (such as "robotic voice" or "hollow feeling"). When users use dialects, this problem may be more prominent because the voice characteristics of dialects (such as pitch, timbre, harmonic structure) are significantly different from those of Mandarin or other standard languages.
[0065] For example:
[0066] Pitch loss: Many dialects (such as Cantonese, Minnan dialect) have complex tone systems. The noise reduction algorithm may mistakenly filter out some high-frequency or low-frequency tone components as noise, resulting in flat-sounding speech or loss of semantic information.
[0067] Timbre change: The unique timbre in dialects (such as nasal sounds, guttural sounds) may be misidentified as noise or non-speech components by the algorithm, resulting in unnatural-sounding speech or even making people feel like a "robotic voice".
[0068] Detail loss: Some rapidly changing pronunciation details in dialects (such as entering tones, light tones) may be smoothed by the algorithm, causing speech to be blurred or semantic confusion.
[0069] Hollow feeling: In a low signal-to-noise ratio (SNR) environment, the noise reduction algorithm may over-suppress background noise, resulting in "holes" in the spectrum of the voice signal. This phenomenon is particularly obvious in dialects because the spectrum distribution of dialects may be quite different from the standard speech in the algorithm training data.
[0070] This is because most noise reduction algorithms (such as deep learning-based voice enhancement models) are trained based on large-scale datasets of Mandarin, English, or other standard languages. These datasets often lack sufficient dialect samples, resulting in insufficient modeling of the characteristics of non-standard speech (such as the tones and harmonics of dialects) by the algorithm.
[0071] Noise reduction algorithms usually distinguish speech and noise through spectrum analysis. However, in a low signal-to-noise ratio environment, the harmonic or detail components of the voice signal may overlap with the spectral characteristics of the noise. Some unique frequency components in dialects (such as high-frequency tones or low-frequency guttural sounds) are more likely to be misjudged as noise and filtered out.
[0072] To suppress noise, many algorithms adopt smoothing techniques (such as spectral subtraction, Wiener filtering, or masking in deep learning), which may weaken the fast-changing features in the speech signal, and the fast-changing pronunciation details such as the entering tone and the light tone in dialects are particularly vulnerable to being affected.
[0073] Many dialects (for example, Cantonese has 9 tones while Mandarin has only 4) rely on high-frequency or low-frequency tone information to convey semantics, and noise reduction algorithms may not be able to accurately retain this information. For example, the common contrast between "voiceless sounds" and "voiced sounds" in Wu dialects (such as Shanghai dialect), the "entering tone" syllables in Cantonese, and the "r-ending" in Sichuan dialect may all be misrecognized as noise or non-speech components by the algorithms. The speech rate and rhythm of dialects may be different from those of Mandarin. For example, the speech rate of some dialects (such as Northeast dialect) is relatively fast, and noise reduction algorithms may not be able to well adapt to this rapidly changing speech pattern, resulting in the loss of details.
[0074] In a low signal-to-noise ratio environment, the energy of the speech signal is severely masked by noise. In order to improve the signal-to-noise ratio, noise reduction algorithms often adopt more aggressive filtering strategies. This strategy is more likely to cause speech distortion when dealing with dialects because the speech features of dialects may not match the "speech templates" of the algorithms.
[0075] The speech after noise reduction sounds blurred or incomplete. For example, the loss of tones in Cantonese may cause words such as "chicken" and "machine" to be indistinguishable, affecting the accurate transmission of semantics.
[0076] In scenarios such as meetings, phone calls, or voice assistants, the speech of dialect users may be over-processed, making it difficult for the call recipients or speech recognition systems to understand.
[0077] The timbre and tones of dialects are important components of their cultural and emotional expressions. Excessive filtering by noise reduction algorithms may make the speech sound mechanical or lose its emotion. For example, after the "r-ending" in Sichuan dialect is filtered out, the speech may lose its sense of intimacy.
[0078] Users may feel that the speech is "robotized", reducing the comfort of phone calls or voice interactions.
[0079] For users who rely on dialects for daily communication, speech distortion may directly affect communication efficiency, especially in noisy environments (such as vegetable markets, subway stations).
[0080] In voice assistants or smart devices (such as smart speakers), improper processing of dialects by noise reduction algorithms may lead to a decrease in speech recognition rate, reducing the convenience of use.
[0081] Dialects are not only communication tools but also carriers of cultural inheritance and emotional expressions. The distortion of noise reduction algorithms may weaken the cultural characteristics of dialects, affecting users' cultural identity and usage experience.
[0082] Based on this, a voice noise reduction method for intelligent communication devices based on artificial intelligence is designed, including an intelligent communication device, on which a microphone array is deployed to collect the voice data of users;
[0083] The collected noisy voice data is input into the edge layer, and the noise-reduced voice data is obtained through real-time processing and output;
[0084] Among them, the edge layer is deployed on the intelligent communication device in an end-to-end generation architecture, which can be installed through a chip. The purpose of the end-to-end generation architecture is to directly generate clean speech from noisy speech, avoiding the destruction of dialect features by intermediate processing steps. The hierarchical deployment of the noise reduction algorithm specifically includes: pre-building a dialect speech database, adding a scene evolution layer, a phonetic variation layer, and a noise mixing layer to the conditional adversarial network cGAN to expand the dialect speech database, forming data samples, and performing three-dimensional annotation on each voice data in the data samples;
[0085] Based on the end-to-end speech enhancement model, an input, output, and multi-task learning mechanism is designed, and pre-training is performed using the data samples. The multi-task learning mechanism combines a deep learning structure with the generative adversarial network GAN for speech reconstruction. A shared encoder is set in the multi-task learning mechanism for joint tasks R1, R2, and R3. Task R1 designs an IRM estimator based on U-Net, and uses the IRM estimator to train the cIRM model for noise suppression, and a lightweight signal-to-noise ratio estimator is added; Task R2 jointly performs retention supervision of dialect features based on a feature embedding layer, a tone protection module, a timbre protection module, and a semantic-assisted noise reduction module; Task R3 reconstructs speech through a generative adversarial network.
[0086] The method of deploying a microphone array on an intelligent communication device to collect the voice data of users includes:
[0087] Select a suitable array and spacing according to the size of the intelligent communication device and the computing resources of the edge layer. The arrays include linear arrays, circular arrays, planar arrays, and three-dimensional arrays, etc. The spacing generally includes narrow spacing, wide spacing, and mixed spacing. Generally speaking, the larger the array, the better the performance. The microphone selects a high sampling rate (such as 48 kHz or higher) and a wide frequency range to ensure that it can comprehensively capture the complete spectral information of the voice signal including dialects, covering high-frequency tones, low-frequency gutturals, and other key audio features. To further improve the sound pickup effect, source localization and beamforming are achieved through the spatial distribution of multiple microphones, thereby enhancing the extraction ability of the target voice signal. Equip with an analog-to-digital converter (ADC) with a high dynamic range and connect it to the edge layer of the intelligent communication device using an interface (such as I2S, SPI, or USB) to ensure high fidelity during the digitization of the audio signal. At the same time, dynamically adjust the gain of the microphone according to the real-time monitored environmental noise level to ensure effective improvement of the signal quality in a low signal-to-noise ratio environment and avoid signal saturation or loss of details.
[0088] The methods for collecting data samples include;
[0089] The low signal-to-noise ratio environment is the core challenge of the noise reduction algorithm. Therefore, it is necessary to construct training data specifically for low signal-to-noise ratio scenarios while ensuring coverage of the diversity of dialects.
[0090] Record the dialect voice data of users in a low signal-to-noise ratio environment. For example, record the voice data of dialect users in noisy environments (such as streets, factories, subway stations), ensuring that the data contains different types of noise (stationary noise, non-stationary noise, burst noise). Collect the voice data of the main dialects (such as Cantonese, Wu dialect, Minnan dialect, Sichuan dialect, Northeast dialect, etc.). After recording, the voice data is voluntarily uploaded by the users of the intelligent device (subject to privacy protection regulations) to enrich the dialect samples in a low signal-to-noise ratio environment and construct a dialect voice database covering a low signal-to-noise ratio environment (-10 dB to 5 dB);
[0091] The architecture of the conditional generative adversarial network cGAN includes a generator and a discriminator. Input the voice data in the dialect voice database into the generator. Add a scene evolution layer, a phonetic variation layer, and a noise mixing layer in the generator. Adopt an alternating training method to train the generator and the discriminator alternately. After the training is completed, fix the discriminator and use the generator to generate new low signal-to-noise ratio dialect voice data. The generated data is used to expand the existing dialect voice database to form data samples;
[0092] The goal of the above processing is to construct a generation model that can generate low signal-to-noise ratio dialect speech data with specific scene noise and phonetic variation characteristics, so as to expand the existing dialect speech dataset and improve the performance of tasks such as speech recognition and speech enhancement in low signal-to-noise ratio environments. Conditional generative adversarial network (cGAN), cGAN is a special GAN that adds conditional information to the inputs of both the generator and the discriminator. In this solution, the conditional information includes scene labels and phonetic variation parameters. Deep learning models such as convolutional neural network (CNN), recurrent neural network (RNN, such as LSTM, GRU), or Transformer can be used to construct the generator and the discriminator.
[0093] When the speech data is input into the generator, the generator is designed with an input layer for receiving pure dialect speech signals (which can be represented as STFT spectrograms, MFCC features, or other acoustic features), and a speech encoder for encoding the input speech signal into a low-dimensional latent space representation, and structures such as CNN, RNN, or Transformer can be used.
[0094] In the model training stage, an alternating training method is adopted to train the generator and the discriminator in turn. For the generator loss, adversarial loss is adopted to encourage the generator to generate more realistic samples, making it difficult for the discriminator to distinguish. Wasserstein distance, least squares loss, etc. can be used. Other loss functions can also be introduced, for example: cycle consistency loss (ensuring that the speech generated by the generator can reconstruct the original pure speech after passing through the discriminator and the generator in a cycle), perceptual loss (using a pre-trained speech recognition model to extract features and calculating the distance between the generated speech and the real speech in the feature space). The discriminator loss is used to encourage the discriminator to more accurately distinguish real samples and generated samples. Adam, SGD, or other common optimization algorithms are used. A validation set can be preset and used to adjust the hyperparameters of the model, such as learning rate, batch size, signal-to-noise ratio range, phonetic variation parameters, etc.
[0095] Set a three-dimensional blank dataset. The three dimensions include multi-granularity acoustic features, noise features, and semantic features. The multi-granularity acoustic features include phoneme level, pronunciation mode, prosodic features, and timbre of the speech data. The noise features include noise events, noise components, and time-frequency distribution. The semantic features include word-level annotation, sentence-level annotation, and discourse-level annotation; extract the three-dimensional data of each speech data in the data sample and fill the data into the blank dataset to form the three-dimensional annotation of each speech data.
[0096] Among them, phoneme-level annotation not only annotates tones but also the initials and finals of each syllable (as well as the onset, nucleus, and coda of the finals, if distinguishable in the dialect). This is crucial for capturing the subtle pronunciation differences in dialects. Articulation manner annotation is used to annotate the articulation manner of phonemes (such as aspirated / unaspirated, voiceless / voiced, stop / fricative / nasal, etc.). This is particularly useful for distinguishing special sound changes in dialects (such as the voiceless / voiced contrast in Wu dialect). Prosodic feature annotation is used to annotate prosodic information such as stress, intonation, and pause. These information are more vulnerable to interference under low signal-to-noise ratio, but are important for semantic understanding and speech naturalness. Timbre annotation refinement can use more professional phonetic terms to describe timbre, such as "bright", "deep", "hoarse", "nasal", etc., and try to maintain the consistency of annotation.
[0097] For sudden noises (such as car horns, knocking sounds), noise event annotation labels their start time, end time, type (such as "short honk", "continuous honk"), source (such as car), etc. For steady-state noises, in addition to labeling the type (such as "white noise", "pink noise"), noise component analysis also labels its main components (such as "dominated by human voices", "dominated by traffic noise"). Time-frequency distribution refinement can use more refined time-frequency analysis tools (such as wavelet transform) to describe the time-frequency distribution of noise.
[0098] Word-level annotation is the basic annotation of speech recognition results. Sentence-level annotation is to annotate the intention, emotion, etc. of sentences. Discourse-level annotation can annotate the topic, turn, etc. of conversations.
[0099] Adding a scene evolution layer, a sound change layer, and a noise mixing layer in the generator, the methods include:
[0100] The input of the scene evolution layer is the preset scene label. The scene label is processed using one-hot vectors or embedding vectors. The scene label can be manually marked without repetition according to different scenes. According to the scene label, the corresponding noise segments are extracted from the dialect speech database, and the scene speech is output through random mixing. Random mixing includes single-scene isolation and multi-scene superposition. Single-scene isolation is to directly output the selected noise segment. Multi-scene superposition is to linearly or non-linearly superpose the noise segments of multiple scenes according to the preset weight distribution mechanism for each scene. The mixed scene speech has the same dimension and length as the input speech data.
[0101] The probability distribution mechanism means that a weight value range is preset for each scene, and the sum of weights is less than or equal to 1 when multi-scene superposition is performed;
[0102] Specifically, we build a dataset containing noises from various scenes, such as street noise (cars, crowds, horns), restaurant noise (conversation sounds, clashes of cutlery), and office noise (keyboard sounds, air conditioner sounds, phone rings). We randomly select one or more scenes, assign a weight to each scene (to simulate the intensity of different scenes), and add the noises from multiple scenes according to the weights. We can also randomly crop, change the speed, and change the pitch of the noise samples to increase the diversity of the noise.
[0103] The input of the sound change layer is the output of the speech encoder and the sound change control parameters. According to the phonetic characteristics of the dialect, the sound change rules are defined, such as consonant weakening, vowel nasalization, tone drift, etc. Specifically, for slurred speech, the simulated consonant pronunciation is unclear, for example, / t / is pronounced as / d / , / k / Pronounce as / g / , for swallowing, simulate the omission of certain syllables or phonemes, for example, pronounce "we" as "I", for connected reading, simulate the unclear transition between syllables, and for tone change, simulate the slight shift or instability of dialect tones, so as to form a rule-based sound change model; use dialect speech data with sound change annotations to train a sequence-to-sequence model (for example, an encoder-decoder model based on an attention mechanism) for predicting and generating sound changes, so as to form a data-driven sound change model; combine the rule-based sound change model with the data-driven sound change model, control the degree and type of sound change by adjusting sound change parameters (for example, sound change probability, sound change intensity, etc.), and output a sequence of speech features that have undergone sound change, recorded as sound-changed speech;
[0104] The input of the noise mixing layer is the scene speech and the sound-changed speech, which are mixed according to the specified signal-to-noise ratio. Additive mixing, convolution mixing or other mixing methods can be used to output a speech feature sequence with noise and sound changes, which is recorded as the new low signal-to-noise ratio dialect speech data.
[0105] Design a noise reduction algorithm specifically for low signal-to-noise ratio environments, focusing on solving the problems of speech distortion, sound quality degradation, and misprocessing of dialect features. The algorithm uses a deep learning framework combined with adaptive strategies and multi-task learning.
[0106] A pre-built dialect speech database solves the problem of insufficient dialect processing in traditional noise reduction systems. The three-dimensional annotation system (acoustic, noise, semantic) provides comprehensive feature protection guidance. The feature embedding layer integrates dialect label information into the noise reduction process to enhance the ability to retain dialect features. The scene evolution layer, sound change layer and noise mixing layer in the cGAN architecture form a complete data enhancement pipeline. The single-scene isolation and multi-scene superposition mechanism simulates the real usage environment. The combination of rule-based and data-driven sound change models comprehensively covers the dialect speech change characteristics.
[0107] The design of the noise reduction algorithm also includes:
[0108] Design a noise reduction algorithm based on an end-to-end speech enhancement model, including input, output, and multi-task learning mechanisms. Set the input as the time-frequency spectrogram of noisy speech. The multi-task learning mechanism uses deep learning structures (such as Deep ComplexU-Net, Transformer, Conformer, CRNN, etc.) for time-frequency processing, and combines with a generative adversarial network GAN (such as SEGAN, MetricGAN, etc., which are GAN structures specifically designed for speech enhancement) for speech reconstruction. The output is the time-frequency spectrogram of the denoised clean speech;
[0109] The input noisy speech generates a time-frequency spectrogram through the short-time Fourier transform STFT, and the time-frequency spectrogram of the denoised clean speech output reconstructs the time-domain signal through the inverse STFT to form the denoised speech data;
[0110] Using an end-to-end model can directly map from the time-frequency spectrogram of noisy speech to the time-frequency spectrogram of clean speech, avoiding error accumulation between multiple modules in traditional methods. Using deep learning for time-frequency processing and combining with GAN for speech reconstruction can effectively learn complex noise patterns and generate high-quality speech. Adding a lightweight SNR estimator helps the model adaptively process speech with different signal-to-noise ratios. Introducing auxiliary noise reduction technology to protect dialect features can improve the performance of the model.
[0111] Design the loss function as the weighted sum of noise reduction loss, speech reconstruction loss, dialect feature loss, and naturalness loss. The noise reduction loss is used to measure the noise reduction effect in a low signal-to-noise ratio environment, and the mean square error (MSE) or signal log-spectrum distance (LSD) can be used as the loss function for the noise reduction task. The speech reconstruction loss is used to evaluate the perceptual quality of the denoised speech data. Specifically, introduce a perceptual loss and use a pre-trained speech recognition model (such as HuBERT or Wav2Vec) to evaluate the perceptual quality of the denoised speech data. The dialect feature loss is used to measure the damage of the model to tones and timbres. Specifically, design special tone loss and timbre loss. For example, use the output of the tone classifier to calculate the tone consistency loss. The naturalness loss is used to measure the naturalness of the denoised speech data. Specifically, use the discriminator of the generative adversarial network to optimize the naturalness of the enhanced speech and avoid "robotic voice" or "hollowness"; add a loss priority rule, that is, preset two thresholds, satisfying that the speech reconstruction loss is greater than the threshold and the dialect feature loss is less than the threshold; the goal is to increase the weights of the speech reconstruction loss and the dialect feature loss in a low signal-to-noise ratio environment to ensure the priority retention of the speech signal.
[0112] The above loss calculation can be based on the content marked in three dimensions. For example, the noise reduction loss is used to measure the noise reduction effect in a low signal-to-noise ratio environment. Then, the noise events, noise components, and time-frequency distribution in the noise features can be used as comprehensive consideration factors. Also, for example, the phoneme level, pronunciation manner, prosody features, and timbre in the multi-granularity acoustic features can be used to determine whether the denoised speech data is still the same as the input, and whether the dialect features have not changed, etc. The loss calculation based on the three-dimensional annotation (noise features, multi-granularity acoustic features) can more finely evaluate the model performance.
[0113] Use data samples to pre-train the noise reduction algorithm until the value of the loss function meets the requirements or reaches the number of iterations, then stop to obtain the pre-trained noise reduction algorithm, and deploy it on the edge layer for processing real-time speech data.
[0114] Adopt an innovative architecture that shares an encoder for three tasks (R1, R2, R3), achieving a balance between noise reduction and feature retention. Task R1 effectively suppresses noise through the IRM and cIRM estimators of the U-Net structure. Task R2 focuses on retaining dialect features to avoid key speech features being filtered. Task R3 uses GAN technology to reconstruct speech details and improve naturalness.
[0115] The multi-task learning mechanism uses a deep learning structure for time-frequency processing. The method of combining the generative adversarial network GAN for speech reconstruction includes:
[0116] Set up a shared encoder to extract the deep feature representation of the spectrogram. These features should be able to capture the common information of speech and noise, providing a processing basis for subsequent tasks R1, R2, and R3. Use a deep learning structure, such as a multi-layer convolutional neural network (CNN), recurrent neural network (RNN), Transformer layer, or their combination (such as Conformer). Output the encoded feature map;
[0117] Preset task-specific heads include task R1, task R2, and task R3.
[0118] Task R1 is used for noise suppression. It removes background noise by learning the time-frequency mask, and focuses on optimizing the mask accuracy in a low signal-to-noise ratio environment. Add a lightweight signal-to-noise ratio estimator (such as a model based on CNN or RNN) after the shared encoder and before the IRM estimator in task R1 to evaluate the signal-to-noise ratio level of the input speech. Specifically, it includes:
[0119] Task R1 designs an IRM estimator based on U-Net. Its input is the feature map output by the shared encoder. The structure includes an encoder, a decoder, and an output layer. The encoder is similar to or the same as the shared encoder structure (the number of layers can be reduced), and it contains multiple downsampling modules (for example, a convolutional layer + a pooling layer), gradually reducing the spatial resolution of the feature map while increasing the number of channels. The decoder is symmetric to the encoder and contains multiple upsampling modules (for example, a transposed convolutional layer or an upsampling layer + a convolutional layer), gradually restoring the spatial resolution of the feature map while reducing the number of channels. The encoder and the decoder are connected by skip connections. The skip connection is to concatenate or add the output feature map of each downsampling module in the encoder with the input feature map of the corresponding upsampling module in the decoder. This helps to transfer the low-level detailed information to the decoder and improve the mask accuracy. The output layer is set as a 1x1 convolutional layer, and the Sigmoid activation function is used to map the feature map to the IRM mask. Among them, the size of the output IRM mask is the same as that of the input STFT spectrogram;
[0120] The loss function is mean square error or binary cross entropy. The IRM model is trained with noisy speech and clean speech data until convergence. The output is the trained IRM model (denoted as IRM_Model) and the predicted IRM mask.
[0121] The output feature map of the shared encoder and the predicted IRM mask are used as the input of the cIRM model. A cIRM estimator is designed based on U-Net, and its structure is similar to that of the IRM model. Specifically, the encoder and the decoder are similar to the IRM model, but the encoder may need some adjustments (for example, increasing the number of channels) to adapt to the complex characteristics of cIRM. The output layer of the decoder needs to predict the real part and the imaginary part of the complex mask. The encoder and the decoder are connected by skip connections. The structure of the output layer is similar to that of the IRM model, that is, usually a 1x1 convolutional layer followed by a Sigmoid activation function to map the feature map to the IRM mask. The size of the output IRM mask is the same as that of the input STFT spectrogram. However, IRM fusion is added after the skip connection. The fusion method is to concatenate the predicted IRM with the output feature map of the encoder as the input of the cIRM encoder. In each upsampling module of the decoder, the predicted IRM is fused with the upsampled feature map (for example, the fusion can be an attention mechanism, weighted average). The function is to predict the complex ideal ratio mask (cIRM), and cIRM contains a real part and an imaginary part: cIRM = (real part) cIRM_real + j (Imaginary part) cIRM_imag, where j represents a number. Set the loss function. The loss function can use the MSE loss in the complex domain, or calculate the MSE losses of the real and imaginary parts separately and then sum them with weights. During training, the weights of the trained IRM model can be used to initialize the weights of the cIRM model (except for the output layer), and then the cIRM model is fine-tuned using noisy speech and clean speech data. The output is the predicted cIRM mask (including real and imaginary parts).
[0122] Apply the predicted cIRM mask (complex number) to the STFT spectrum of the noisy speech to obtain the STFT spectrum of the denoised speech; the cIRM model processes both amplitude and phase information simultaneously, which is better than traditional methods that only process amplitude.
[0123] Add a lightweight SNR estimator (such as a model based on CNN or RNN) after the shared encoder and before the IRM estimator in Task R1 to evaluate the SNR level of the input speech, which mainly solves three key problems of traditional speech denoising systems in dealing with low SNR environments:
[0124] 1. Traditional denoising systems mostly focus on the amplitude spectrum and ignore the phase information, resulting in a decline in the quality of the reconstructed speech.
[0125] 2. In high-noise environments, the accuracy of traditional IRM masks is insufficient, making it difficult to accurately distinguish noise and speech.
[0126] 3. Lack of an evaluation mechanism for the noise level of the input signal, and the denoising strategy is not flexible enough.
[0127] The reason why it is the best position after the shared encoder and before the IRM estimator is as follows:
[0128] The shared encoder has already extracted the deep features of the input speech, and these features are more suitable for SNR estimation than the original STFT spectrogram because they have undergone a certain degree of abstraction and denoising processing.
[0129] The estimated SNR can be used as prior information to guide the IRM estimator to work better. For example, for speech with low SNR, the IRM estimator can suppress noise more aggressively; for speech with high SNR, the IRM estimator can suppress noise more conservatively to avoid damaging the speech.
[0130] Adding the SNR estimator before the IRM estimator can avoid estimating the SNR for each time-frequency unit, thus improving the computational efficiency.
[0131] It includes: setting the input of the SNR estimator to the feature map output by the shared encoder. The structure includes global average pooling, fully connected layers, activation functions, and an output layer. The output layer uses a fully connected layer with a single output, and the output is the estimated SNR value (in decibels).
[0132] Perform global average pooling on the feature map output by the shared encoder in the time and frequency dimensions, converting the feature map into a vector. Each element of this vector represents the average activation value of a channel. Its role is to compress the spatial information into a global feature vector, and this vector is somewhat representative of the SNR.
[0133] Use one or two fully connected layers (also known as dense layers) to process the vector output by global average pooling. The first fully connected layer can map the feature vector to a lower-dimensional space (e.g., 64 or 128 units). If two fully connected layers are used, the second fully connected layer can further map the features to an even lower-dimensional space (e.g., 16 or 32 units). The role of the fully connected layers is to learn the non-linear mapping from the global average pooling features to the SNR value.
[0134] Use activation functions such as ReLU (Rectified Linear Unit) or LeakyReLU after each fully connected layer. The role of the activation function is to introduce non-linearity and enhance the expressive power of the model.
[0135] Set the loss function to mean squared error, which is used to calculate the mean squared error between the estimated SNR value and the true SNR value. Input the STFT spectrogram of the noisy speech into the shared encoder, input the output feature map of the shared encoder into the SNR estimator, calculate the MSE loss between the estimated SNR value and the true SNR value, and use the backpropagation algorithm to update the weights of the SNR estimator.
[0136] In the IRM estimator, use the estimated SNR value as an additional input and input it into the IRM estimator together with the original input.
[0137] Design a two-layer mask architecture. First, generate a real-valued mask through the IRM estimator to provide basic information for the subsequent cIRM. Then, generate a complex mask through the cIRM estimator to process both amplitude and phase information simultaneously. The cascaded processing of the two layers of masks improves the noise reduction accuracy and speech quality. Fuse the predicted IRM mask with the encoder feature map to provide prior knowledge for the cIRM. Integrate the IRM information at each layer of the decoder to guide the generation process of the cIRM, and achieve fine integration through the attention mechanism or weighted average to improve the mask quality.
[0138] The encoder-decoder symmetric structure effectively extracts and reconstructs time-frequency features. The skip connection mechanism preserves low-level detail information and avoids the loss of important features. The lightweight SNR estimator design reduces the computational overhead, and the shared encoder architecture saves model parameters and computational resources.
[0139] Task R2 is used for dialect feature retention. By introducing the supervision information of dialect features, key features (such as tones, timbres, etc.) are retained and not filtered out.
[0140] It includes a feature embedding layer, a tone protection module, a timbre protection module, and a semantic-assisted noise reduction module;
[0141] The input of the feature embedding layer is (original) noisy speech data + dialect label (e.g., "Cantonese" or an integer ID representing Cantonese). The dialect label is converted into a vector with a fixed dimension (e.g., using one-hot encoding or learning an embedding matrix), denoted as the embedding vector of the dialect. This embedding vector of the dialect is concatenated with the speech feature vector (e.g., MFCC, Mel spectrogram) or fused through an attention mechanism to form a new vector, which will be used as the input of the subsequent noise reduction model. The purpose is to make the model aware that the current processing is a specific dialect (e.g., Cantonese), so as to adjust its noise reduction strategy and avoid misprocessing of the unique features of Cantonese.
[0142] The input of the tone protection module is (original) noisy speech data. Using a pre-trained dialect tone classifier (e.g., trained on a large amount of Cantonese data based on an LSTM or Transformer model), the output is the tone category probability distribution for each time frame (e.g., the probabilities of 9 Cantonese tones) and the tone contour features (e.g., the fundamental frequency F0 and its variations).
[0143] During the noise reduction process, through a loss function (e.g., mean squared error (MSE) can be used), the tone contour features of the output speech are constrained to reach a required similarity with the tone contour features of the tone protection module (e.g., a similarity of 95% is required to ensure as consistent performance as possible). The purpose is to ensure that the spectral features and change patterns of the 9 tones of the dialect (Cantonese) are retained after noise reduction.
[0144] The input of the timbre protection module is (original) noisy speech data. A pre-trained timbre feature extractor (e.g., an autoencoder based on WaveNet, or a speaker recognition model pre-trained on a large amount of speech data) is used to output a timbre feature vector of fixed dimension (e.g., i-vector, x-vector). An item is added to the loss function of the denoising process to penalize the difference between the timbre feature vectors before and after denoising. For example, cosine similarity or Euclidean distance can be used as part of the loss function. A similarity requirement can also be set. If the timbre difference exceeds a threshold, the penalty is increased. This prevents the denoising process from changing the timbre characteristics of the speech and maintains the speaker's personalized characteristics.
[0145] The input of semantic-assisted denoising is (raw) noisy speech data. Using a pre-trained dialect speech recognition model (e.g., Wav2Vec 2.0, Conformer), the output is a text transcription of the speech (possibly with a timestamp), phoneme / syllable-level information (e.g., syllable boundaries, tone information). Based on the speech recognition results, the tone changes (identifying which syllables have tone changes (e.g., "chicken" and "machine")), syllable boundaries (determining the start and end time of each syllable) and key phonemes (identifying phonemes that are important for semantic understanding (e.g., stops and fricatives in Cantonese)) are extracted. Based on the recognition results, they are protected separately. For example, for tone change protection, the role of the tone protection module can be enhanced near the syllable boundaries of the tone changes to ensure that the tone differences are not smoothed out. For syllable boundary protection, the denoising intensity can be reduced near the syllable boundaries to prevent the transition between syllables from being destroyed. For key phoneme protection, the retention of these phoneme features can be enhanced in the time-frequency region corresponding to the key phonemes.
[0146] By using semantic information, the noise reduction process can be more finely controlled to ensure that speech components important for semantic expression are not filtered out. Through multi-level and multi-angle constraints, it is ensured that the intonation, timbre and semantic characteristics of Cantonese are retained to the greatest extent while noise reduction is being performed. It can effectively improve the quality and intelligibility of dialect (Cantonese) speech noise reduction in low signal-to-noise ratio environments.
[0147] Task R3 is used for speech reconstruction, which reconstructs speech details (such as intonation and harmonics) by generating adversarial networks to avoid the "empty feeling".
[0148] The generator of the generative adversarial network is used to reconstruct the harmonics, intonation and fast-changing pronunciation details of the speech, and the discriminator is used to enhance the naturalness of the speech;
[0149] Using the dialect features output by task R2 as a guide, ensure that the enhanced speech is as close as possible to the sound quality of the real dialect. For example, use a fundamental frequency (F0) estimator to extract the fundamental frequency information of the speech, and force the fundamental frequency to be retained during the reconstruction process to avoid tone loss. Use formant analysis to enhance the timbre features of the speech and avoid timbre changes.
[0150] Implement in a low signal-to-noise ratio environment to preferentially enhance the fundamental frequency and formant features of the speech, avoid the "hollow feeling", and at the same time reduce the excessive filtering of high-frequency noise.
[0151] Example 2, please refer to Figure 1 and Figure 3 As shown, for the parts not described in detail in this embodiment, refer to the description in Example 1. Provide an artificial intelligence-based intelligent communication device speech noise reduction method, including:
[0152] Step S1: Deploy a microphone array on the intelligent communication device to collect the noisy speech data of the user in real time;
[0153] Step S2: Design the input, output, and multi-task learning mechanism of the noise reduction algorithm based on the end-to-end speech enhancement model;
[0154] Step S3: Use data samples for pre-training to obtain a trained noise reduction algorithm and deploy it to the edge layer;
[0155] Step S4: Input the collected noisy speech data into the edge layer and output the noise-reduced speech data through real-time processing.
[0156] Through the innovative multi-task learning architecture and dialect feature protection mechanism, effectively solve the limitations of traditional noise reduction systems in dealing with dialects, and significantly improve the speech quality and user experience of intelligent communication devices in complex noise environments.
[0157] Example 3, this embodiment publicly provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the above-provided artificial intelligence-based intelligent communication device speech noise reduction system.
[0158] Since the electronic device introduced in this embodiment is the electronic device adopted by a voice noise reduction system for an intelligent communication device based on artificial intelligence in the embodiments of the present application, based on the voice noise reduction system for an intelligent communication device based on artificial intelligence introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted by a voice noise reduction system for an intelligent communication device based on artificial intelligence in the embodiments of the present application, it falls within the scope of protection of the present application.
[0159] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0160] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A voice noise reduction system for intelligent communication equipment based on artificial intelligence, characterized in that: Including intelligent communication equipment, on which a microphone array is deployed to collect real-time noisy voice data of users; The recorded noisy speech data is input into the edge layer, where a pre-trained noise reduction algorithm is deployed. The noise-reduced speech data is obtained through real-time processing and output. The hierarchical deployment of the noise reduction algorithm specifically includes: pre-building a dialect speech database, adding a scene evolution layer, a sound change layer, and a noise mixing layer in the conditional adversarial network cGAN to expand the dialect speech database, form data samples, and perform three-dimensional annotation on each speech data in the data sample; Based on the end-to-end speech enhancement model, the input, output and multi-task learning mechanisms are designed, and data samples are used for pre-training. The multi-task learning mechanism combines deep learning structure with generative adversarial network GAN for speech reconstruction. A shared encoder is set in the multi-task learning mechanism for joint tasks R1, R2 and R3. Task R1 designs an IRM estimator based on U-Net, and uses the IRM estimator to train the cIRM model for noise suppression, and adds a lightweight signal-to-noise ratio estimator; Task R2 jointly supervises the retention of dialect features based on the feature embedding layer, tone protection module, timbre protection module and semantic-assisted denoising module; Task R3 reconstructs speech through a generative adversarial network.
2. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 1, characterized in that: The method of deploying a microphone array on the intelligent communication device to collect real-time noisy voice data of the user includes: According to the size of the intelligent communication device and the computing resources of the edge layer, a suitable array and spacing are selected. The microphone has a high sampling rate and a wide frequency range, is equipped with an analog-to-digital converter with a high dynamic range, and is connected to the edge layer of the intelligent communication device using an interface.
3. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 2, characterized in that: The data sample collection method includes: Record the user's dialect voice data in a low signal-to-noise ratio environment and build a dialect voice database covering the low signal-to-noise ratio environment; The architecture of the conditional adversarial network cGAN includes a generator and a discriminator. The speech data in the dialect speech database is input into the generator. The scene evolution layer, sound change layer and noise mixing layer are added to the generator. The generator and the discriminator are trained in turn by alternating training. After the training is completed, the discriminator is fixed and the generator is used to generate new low signal-to-noise ratio dialect speech data. The generated data is used to expand the existing dialect speech database to form data samples. A three-dimensional blank dataset is set up, where the three dimensions include multi-granularity acoustic features, noise features and semantic features. The multi-granularity acoustic features include the phoneme level, pronunciation method, rhythmic features and timbre of the speech data. The noise features include noise events, noise components and time-frequency distribution. The semantic features include word-level annotations, sentence-level annotations and paragraph-level annotations. The three-dimensional data of each speech data in the data sample is extracted, and the data is filled into the blank dataset to form a three-dimensional annotation for each speech data.
4. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 3 is characterized in that: The method of adding a scene evolution layer, a sound change layer and a noise mixing layer in the generator includes: The input of the scene evolution layer is the preset scene label. According to the scene label, the corresponding noise fragment is extracted from the dialect speech database, and the scene speech is output through random mixing. Random mixing includes single scene isolation and multi-scene superposition. Single scene isolation is to directly output the selected noise fragment, and multi-scene superposition is to linearly or nonlinearly superimpose the noise fragments of multiple scenes according to the preset weight distribution mechanism of each scene. The probability distribution mechanism means that a weight value range is preset for each scenario, and the sum of the weights is less than or equal to 1 when multiple scenarios are superimposed; The input of the sound change layer is the output of the speech encoder and the sound change control parameters. According to the phonetic characteristics of the dialect, the sound change rules are defined to form a rule-based sound change model. Using the dialect speech data with sound change annotations, a sequence-to-sequence model is trained to predict and generate sound changes, forming a data-driven sound change model. The rule-based sound change model and the data-driven sound change model are combined to control the degree and type of sound change by adjusting the sound change parameters, and the sound-changed speech feature sequence is output, which is recorded as the sound-changed speech. The input of the noise mixing layer is the scene speech and the sound-changed speech, which are mixed according to the specified signal-to-noise ratio, and the speech feature sequence with noise and sound changes is output, which is recorded as the new low signal-to-noise ratio dialect speech data.
5. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 4, characterized in that: The design of the noise reduction algorithm also includes: Design a noise reduction algorithm based on an end-to-end speech enhancement model, including input, output, and multi-task learning mechanism. Set the input as the time-frequency spectrum of noisy speech. The multi-task learning mechanism uses a deep learning structure for time-frequency processing and combines it with a generative adversarial network (GAN) for speech reconstruction. The output is a time-frequency spectrum of pure speech after noise reduction. The input noisy speech is transformed into a time-frequency spectrum through short-time Fourier transform (STFT), and the output clean speech time-frequency spectrum after noise reduction is reconstructed into a time domain signal through inverse STFT to form the noise-reduced speech data; The loss function is designed to be the weighted sum of noise reduction loss, speech reconstruction loss, dialect feature loss and naturalness loss. The noise reduction loss is used to measure the noise reduction effect in a low signal-to-noise ratio environment. The speech reconstruction loss is used to evaluate the perceptual quality of the speech data after noise reduction. The dialect feature loss is used to measure the damage of the model to the tone and timbre. The naturalness loss is used to measure the naturalness of the speech data after noise reduction. A loss priority rule is added, that is, two thresholds are preset to meet the requirement that the speech reconstruction loss is greater than the threshold and the dialect feature loss is less than the threshold. Use data samples to pre-train the denoising algorithm until the loss function value meets the requirements or the number of iterations is reached. Then, the pre-trained denoising algorithm is obtained and deployed at the edge layer to process real-time voice data.
6. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 5, characterized in that: The multi-task learning mechanism adopts a deep learning structure to perform time-frequency processing, and the method of combining a generative adversarial network (GAN) to perform speech reconstruction includes: Set up a shared encoder to extract deep feature representation of the spectrogram, use a deep learning structure, and output the encoded feature map; The preset task-specific headers include Task R1, Task R2, and Task R3; Task R1 is used for noise suppression. A lightweight signal-to-noise ratio estimator is added after the shared encoder and before the IRM estimator in Task R1 to evaluate the signal-to-noise ratio level of the input speech. Task R2 is used to preserve dialect features. By introducing the supervision information of dialect features, key features are retained and not filtered out. Task R3 is used for speech reconstruction. It reconstructs the details of speech through a generative adversarial network, including using the generator of the generative adversarial network to reconstruct the harmonics, tones, and rapidly changing pronunciation details of the speech, and the discriminator enhances the naturalness of the speech. The dialect features output by Task R2 are introduced as a guide.
7. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 6, characterized in that: The method for noise suppression in task R1 includes: Task R1 designs an IRM estimator based on U-Net, whose input is the feature map output by the shared encoder. The structure includes an encoder, a decoder, and an output layer. The encoder contains multiple downsampling modules. The decoder is symmetrical with the encoder and contains multiple upsampling modules. The encoder and the decoder are jump-connected. The output layer is set to a 1x1 convolutional layer. The Sigmoid activation function is used to map the feature map to the IRM mask. The size of the output IRM mask is the same as the input STFT spectrum map. The loss function is mean square error or binary cross entropy. The IRM model is trained using noisy speech and clean speech data until convergence. The output is the trained IRM model and the predicted IRM mask. The output feature map of the shared encoder and the predicted IRM mask are used as the input of the cIRM model. The cIRM estimator is designed based on U-Net, and IRM fusion is added after the jump connection. The fusion method is to splice the predicted IRM with the output feature map of the encoder as the input of the cIRM encoder. In each upsampling module of the decoder, the predicted IRM is fused with the upsampled feature map. The cIRM contains real and imaginary parts. The loss function is set and the output is the predicted cIRM mask. The predicted cIRM mask is applied to the STFT spectrum of the noisy speech to obtain the STFT spectrum of the denoised speech.
8. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 7, characterized in that: A lightweight signal-to-noise ratio estimator is added after the shared encoder in the task R1 and before the IRM estimator to evaluate the signal-to-noise ratio level of the input speech, including: The input of the signal-to-noise ratio estimator is set to the feature map output by the shared encoder. The structure includes global average pooling, a fully connected layer, an activation function, and an output layer. The output layer uses a single-output fully connected layer, and the output is the estimated SNR value. The loss function is set to mean square error to calculate the mean square error between the estimated SNR value and the true SNR value. The STFT spectrogram of the noisy speech is input into the shared encoder, and the output feature map of the shared encoder is input into the signal-to-noise ratio estimator to calculate the MSE loss between the estimated SNR value and the true SNR value. The weight of the signal-to-noise ratio estimator is updated using the back propagation algorithm. In the IRM estimator, the estimated SNR value is input as an additional input into the IRM estimator together with the original input.
9. The voice noise reduction system for intelligent communication equipment based on artificial intelligence according to claim 8, characterized in that: The task R2 is used to retain dialect features, by introducing the supervision information of dialect features, retaining the key features from being filtered out. include: It includes feature embedding layer, tone protection module, timbre protection module and semantic-assisted noise reduction module; The input of the feature embedding layer is noisy speech data + dialect label; the dialect label is converted into a vector of fixed dimension, recorded as the dialect embedding vector; this dialect embedding vector is concatenated with the speech feature vector or fused through the attention mechanism to form a new vector, which will be used as the input of the subsequent noise reduction algorithm; The input of the tone protection module is noisy speech data. It uses a pre-trained dialect tone classifier and outputs the tone category probability distribution and tone contour features for each time frame. In the process of noise reduction, the tone contour features of the output speech are constrained by the loss function so that the similarity between the tone contour features of the tone protection module and the tone contour features reaches the required level. The input of the timbre protection module is noisy speech data. The pre-trained timbre feature extractor is used to output a timbre feature vector of fixed dimension. A term is added to the loss function of the denoising process to penalize the difference between the timbre feature vectors before and after denoising. The input of semantic-assisted denoising is noisy speech data. Using a pre-trained dialect speech recognition model, the output is text transcription of the speech and information at the phoneme / syllable level. Based on the speech recognition results, tone changes, syllable boundaries and key phonemes are extracted; and protection is performed separately based on the recognition results.
10. A method for voice noise reduction of intelligent communication equipment based on artificial intelligence, applied to the voice noise reduction system of intelligent communication equipment based on artificial intelligence according to any one of claims 1 to 9, characterized in that: The method for reducing voice noise of intelligent communication equipment based on artificial intelligence includes: Step S1: deploying a microphone array on the intelligent communication device to collect the user's real-time noisy voice data; Step S2: designing the input, output and multi-task learning mechanism of the noise reduction algorithm based on the end-to-end speech enhancement model; Step S3: Use data samples for pre-training, obtain the trained denoising algorithm, and deploy it to the edge layer; Step S4: input the collected noisy speech data into the edge layer, and output the noise-reduced speech data through real-time processing.
Citation Information
Patent Citations
Voiceprint recognition model training method based on multi-task learning and adversarial training
CN114171031A
Time domain voice separation method based on full convolutional neural network multi-task learning
CN117912482A