Speech enhancement method and device based on multi-scale feature learning, equipment and medium

Through technical means such as multi-scale feature learning and deep residual networks, the problem of insufficient speech enhancement capabilities of existing neural vocoders in complex noise environments is solved, and a higher quality speech enhancement effect is achieved.

CN120220712APending Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510417793.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing neural vocoders have insufficient voice enhancement capabilities in complex noise environments, making it difficult to effectively remove background interference, affecting the quality of voice interaction.

Method used

The speech enhancement method based on multi-scale feature learning is adopted, and the frequency domain multi-scale features of Mel spectrum features are extracted through multi-scale convolutional neural networks. The deep residual network performs noise suppression, and the non-autoregressive generation model optimizes feature conversion, and the speech waveform is reconstructed by the generation of adversarial networks.

Benefits of technology

It improves the purity and clarity of speech signals, enhances the modeling efficiency of speech enhancement and the naturalness and clarity of speech generation, adapts to complex noise environments, and improves the quality of speech interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220712A_ABST
    Figure CN120220712A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech processing, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a speech enhancement method based on multi-scale feature learning, which comprises the following steps: framing an input audio signal, extracting a Mel-frequency spectrum feature, extracting a frequency domain feature by using a multi-scale convolutional neural network, and carrying out multi-scale feature learning on the frequency domain feature; carrying out coding dimension reduction on the image; noise is suppressed through a deep residual network, enhanced audio features are generated, a non-autoregression generative model is adopted for feature conversion, and finally a generative adversarial network is used for reconstructing a target voice waveform. Voice frequency domain features are extracted through the multi-scale convolutional neural network, and the feature expression ability of different frequency bands is improved; noise suppression is carried out through a deep residual network, and the purity of the voice signals is enhanced; feature conversion is optimized through a non-autoregression generation model, and the modeling efficiency of speech enhancement is improved; the target voice waveform is reconstructed through the generative adversarial network, and the naturalness and definition of the generated voice are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a speech enhancement method, device, equipment and storage medium based on multi-scale feature learning. Background Art

[0002] In the field of speech signal processing, as an advanced speech generation technology, neural vocoders have been widely studied and applied in application scenarios such as speech synthesis and voice conversion. However, in speech enhancement tasks, existing neural vocoders still have some technical bottlenecks, which limit their application effects in complex environments.

[0003] In the field of medical and health services, applications such as remote consultation, intelligent voice assistants, and voice medical record keeping require real-time processing of the speech of doctors and patients to ensure high-quality voice interaction. However, the enhancement ability of existing neural vocoders in complex noise environments is still insufficient. For example, in a hospital environment, medical equipment noise, ward background noise, etc. may affect the speech quality of patients. When facing these complex noises, current speech enhancement technologies often cannot effectively remove background interference, which may affect doctors' remote diagnosis and patients' speech understanding. In addition, speech in medical scenarios often involves important features such as speech rate, intonation, and timbre. Existing neural vocoders mainly focus on the generation of speech waveforms and are difficult to accurately restore these key information, which may lead to a decrease in the accuracy of speech records and affect the integrity of medical records.

[0004] In the field of fintech services, application scenarios such as intelligent customer service, identity authentication, and voice payment have extremely high requirements for the clarity and real-time performance of speech. However, existing neural vocoders usually consume a large amount of computing resources and have a slow inference speed, making it difficult to meet the low-latency requirements. For example, in a bank's remote customer service system, customer speech may contain factors such as environmental noise and signal interference, resulting in difficulty for the speech enhancement model to effectively remove noise, affecting the accuracy of speech recognition, which may lead to failed customer identity verification and even affect the security of financial transactions. In addition, due to the large differences in speech data among different banks, securities, and insurance businesses, the generalization ability of existing neural vocoders across business scenarios is weak, which may lead to unstable speech enhancement effects and affect the service quality of financial institutions.

[0005] In the field of speech enhancement technology, existing neural vocoders usually rely on clean training data and can generate relatively clear speech signals in a controlled environment. However, in a noisy environment or a scenario with complex echoes, their enhancement effect may be unstable. In addition, current neural vocoders are mainly trained based on specific datasets and lack the ability to adapt to unseen noise types and speech scenarios, resulting in poor generalization performance. For example, in applications such as far-field speech recognition and speech translation, if the speech enhancement model cannot effectively handle different noise environments and device conditions, it may reduce the overall performance of the system and affect the user experience and the stability of speech interaction. Summary of the Invention

[0006] The main object of the present invention is to provide a speech enhancement method, device, equipment and storage medium based on multi-scale feature learning, aiming to solve the technical problems that existing neural vocoders consume a large amount of computing resources and have poor adaptability to complex noise environments in speech enhancement tasks, which limit their application in real-time speech processing scenarios.

[0007] To achieve the above object, the present invention provides a speech enhancement method based on multi-scale feature learning, including:

[0008] Obtain an input audio signal and perform frame splitting on the input audio signal to generate audio frames;

[0009] Perform time-frequency conversion on the audio frames to generate Mel spectrogram features;

[0010] Extract the frequency-domain multi-scale features of the Mel spectrogram features through a multi-scale convolutional neural network;

[0011] Encode the frequency-domain multi-scale features into a dimensionality-reduced latent representation;

[0012] Perform noise suppression processing on the dimensionality-reduced latent representation through a deep residual network to generate enhanced audio features;

[0013] Convert the enhanced audio features into an enhanced Mel spectrogram through a non-autoregressive generation model;

[0014] Reconstruct the enhanced Mel spectrogram into a target speech waveform through a generative adversarial network.

[0015] Furthermore, to achieve the above object, the present invention provides a speech enhancement device based on multi-scale feature learning, including:

[0016] An audio preprocessing module for obtaining an input audio signal and performing frame splitting on the input audio signal to generate audio frames;

[0017] A time-frequency transformation module for performing time-frequency conversion on the audio frames to generate Mel spectrogram features;

[0018] A feature extraction module, configured to extract frequency-domain multi-scale features of the Mel spectrogram features through a multi-scale convolutional neural network;

[0019] An encoding processing module, configured to encode the frequency-domain multi-scale features into a dimension-reduced latent representation;

[0020] A noise suppression module, configured to perform noise suppression processing on the dimension-reduced latent representation through a deep residual network to generate enhanced audio features;

[0021] A feature conversion module, configured to convert the enhanced audio features into enhanced Mel spectrograms through a non-autoregressive generative model;

[0022] A speech generation module, configured to reconstruct the enhanced Mel spectrogram into a target speech waveform through a generative adversarial network.

[0023] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a speech enhancement program based on multi-scale feature learning stored in the memory and executable on the processor. When the speech enhancement program based on multi-scale feature learning is executed by the processor, the steps of the speech enhancement method based on multi-scale feature learning as described above are implemented.

[0024] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a speech enhancement program based on multi-scale feature learning is stored. When the speech enhancement program based on multi-scale feature learning is executed by a processor, the steps of the speech enhancement method based on multi-scale feature learning as described above are implemented.

[0025] Beneficial effects: The present invention relates to the technical field of speech processing and can be applied to business scenarios such as medical health and fintech. It discloses a speech enhancement method based on multi-scale feature learning, including: acquiring an input audio signal and performing frame division processing to generate audio frames; performing time-frequency conversion on the audio frames to extract Mel spectrogram features; extracting frequency-domain multi-scale features through a multi-scale convolutional neural network; encoding the extracted frequency-domain multi-scale features to obtain a dimension-reduced latent representation; using a deep residual network to perform noise suppression on the dimension-reduced latent representation to obtain enhanced audio features; converting the enhanced audio features into enhanced Mel spectrograms through a non-autoregressive generative model; using a generative adversarial network to reconstruct the enhanced Mel spectrogram into a target speech waveform. The present invention extracts speech frequency-domain features through a multi-scale convolutional neural network to improve the feature expression ability of different frequency bands; performs noise suppression through a deep residual network to enhance the purity of the speech signal; optimizes feature conversion through a non-autoregressive generative model to improve the modeling efficiency of speech enhancement; and reconstructs the target speech waveform through a generative adversarial network to improve the naturalness and clarity of the generated speech. Brief Description of the Drawings

[0026] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0027] Figure 1 is a schematic diagram of an application environment of a speech enhancement method based on multi-scale feature learning in an embodiment of the present invention;

[0028] Figure 2 is a schematic flowchart of an embodiment of the speech enhancement method based on multi-scale feature learning of the present invention;

[0029] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of a speech enhancement device based on multi-scale feature learning of the present invention;

[0030] Figure 4 is a schematic structural diagram of a computer device in an embodiment of the present invention;

[0031] Figure 5 is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed Description of the Embodiments

[0032] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0033] The speech enhancement method based on multi-scale feature learning provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server through a network. The server can obtain the input audio signal through the client and perform frame splitting to generate audio frames; perform time-frequency conversion on the audio frames to extract Mel spectrum features; extract frequency-domain multi-scale features through a multi-scale convolutional neural network; encode the extracted frequency-domain multi-scale features to obtain a dimension-reduced latent representation; use a deep residual network to suppress noise on the dimension-reduced latent representation to obtain enhanced audio features; convert the enhanced audio features into enhanced Mel spectra through a non-autoregressive generation model; and reconstruct the enhanced Mel spectra into target speech waveforms through a generative adversarial network. The present invention extracts speech frequency-domain features through a multi-scale convolutional neural network to improve the feature expression ability of different frequency bands; suppresses noise through a deep residual network to enhance the purity of the speech signal; optimizes feature conversion through a non-autoregressive generation model to improve the modeling efficiency of speech enhancement; and reconstructs target speech waveforms through a generative adversarial network to improve the naturalness and clarity of the generated speech. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0034] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the speech enhancement method based on multi-scale feature learning provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0035] As Figure 2 shown, the speech enhancement method based on multi-scale feature learning proposed by the present invention includes the following steps:

[0036] S10, obtain an input audio signal, and perform frame splitting processing on the input audio signal to generate audio frames;

[0037] In this embodiment, the acquisition and frame splitting processing of audio signals are the basis of speech signal processing, involving multiple links such as signal acquisition, digitization, and preprocessing. Obtaining audio signals usually depends on microphone arrays, mobile devices, or professional audio acquisition systems. These devices receive external sounds through acoustic sensors, convert analog signals into electrical signals, and then perform analog-to-digital conversion to form discrete digital signals. The performance, sensitivity, frequency response range, etc. of different acquisition devices will all affect subsequent processing. Therefore, in different scenarios, the selection of acquisition devices needs to be optimized for the target application. For example, in an embedded environment, a low-power microphone array is usually used, while in a professional speech enhancement scenario, an audio acquisition device with a high sampling rate and high signal-to-noise ratio may be required.

[0038] Frame splitting processing is to slice continuous audio data to make it suitable for subsequent short-time analysis. Since speech signals have time-varying characteristics, using a fixed-length window to split the signal into frames can improve the stability of time-frequency analysis and reduce cross-frame information loss. Frame splitting usually uses a fixed-length time window, such as 20 ms or 25 ms, and at the same time maintains frame-to-frame continuity through overlapping to prevent information loss. The setting of the frame shift generally depends on the specific application. For example, in real-time speech enhancement, a shorter frame shift is required to reduce latency, while for offline processing, a larger frame shift can be selected to reduce computational overhead. Windowing processing is an essential step in the frame splitting process. Common windowing methods include Hamming window, Hanning window, and rectangular window, etc. Different window functions have different frequency domain characteristics. The Hamming window can reduce spectral leakage and improve the resolution of the signal in the frequency domain, while the Hanning window performs better in terms of signal smoothness and is suitable for application scenarios that require more stable characteristics.

[0039] Short-time energy detection is one of the important processing steps after frame segmentation. Its purpose is to remove low-energy silent frames to improve the efficiency of subsequent speech processing. Short-time energy calculation is usually based on the sum-of-squares energy function, and the threshold can be set using a static threshold or a dynamic adaptive threshold. The static threshold method is simple and efficient and is suitable for scenarios with a stable noise environment, while the dynamic threshold method can be adjusted dynamically according to the change of the environmental noise level and is suitable for complex noise environments.

[0040] In different application environments, the acquisition method of audio signals can be different. For example, in mobile devices, audio signals are usually acquired through built-in microphones, while in remote conferencing systems, directional speech acquisition may be achieved through external microphone arrays. In certain specific scenarios, bone conduction microphones can also be used to reduce the interference of environmental noise and improve the quality of speech signals.

[0041] The implementation method of frame segmentation can be adjusted according to the computing resources and real-time requirements. For example, on embedded devices with limited computing resources, a frame segmentation strategy with a fixed frame length and frame shift can be adopted, while in a high-performance server environment, an adaptive frame segmentation method can be used to dynamically adjust the frame length and frame shift according to the characteristics of the input signal to optimize the performance of speech enhancement. The choice of windowing processing can be adjusted according to application requirements. For example, in a high-noise environment, a smoother window function can be adopted to reduce the impact of boundary effects on signal processing.

[0042] The implementation of short-time energy detection can be combined with environmental noise analysis to improve robustness. In a low-noise environment, a fixed-threshold method can be used, while in a complex noise environment, the threshold can be dynamically adjusted by combining a noise estimation algorithm to more accurately remove silent frames and improve the efficiency of speech processing. In some applications sensitive to energy changes, other voice activity detection (VAD) methods, such as the detection method based on the zero-crossing rate, can also be combined to further improve the detection accuracy.

[0043] Example illustration: In the field of medical and health services, when an intelligent voice assistant is used to assist doctors in recording medical records, it is necessary to accurately extract the doctor's speech information in a noisy hospital environment. By performing frame segmentation on the input audio signal and combining short-time energy detection to remove background noise, the accuracy of speech recognition can be effectively improved, ensuring the integrity and reliability of medical record keeping. In addition, during remote medical consultations, the patient's speech may be affected by factors such as network fluctuations and environmental noise. Adopting appropriate frame segmentation and windowing strategies can improve the robustness of speech enhancement and ensure that doctors can accurately receive the patient's speech information.

[0044] In the field of fintech business, the intelligent customer service system of banks needs to stably identify user voices in different call environments to complete tasks such as identity verification and business consultation. When users are on the phone, factors such as environmental noise and echo may cause the quality of the voice signal to decline. Through frame processing, short-time energy detection, and window optimization, the clarity of the voice signal can be improved, thereby enhancing the accuracy of the voice recognition system, reducing recognition errors caused by voice quality problems, and improving the service efficiency and security of financial services.

[0045] In the field of voice enhancement technology, intelligent conference systems need to enhance voice signals in real time in scenarios of multi-person voice interaction to ensure the intelligibility of voice content. In a conference environment, different speakers may take turns speaking, and background noise may be present. By performing efficient frame processing on the input audio signal, the system can accurately capture the voices of each speaker and remove background noise in non-voice regions, improving the real-time performance and stability of the voice enhancement system.

[0046] By performing frame processing on the audio signal, the voice signal can be analyzed more effectively, reducing cross-frame information loss and improving the time consistency of the signal. By adopting appropriate frame length and frame shift strategies, the utilization rate of computing resources can be optimized while ensuring voice quality. Through window processing, spectral leakage can be reduced, the frequency domain resolution of the signal can be improved, and the clarity of the voice signal can be enhanced. Short-time energy detection can effectively remove useless silent frames, reduce the computational burden, and avoid the influence of background noise on subsequent processing, improving the overall performance of voice enhancement.

[0047] S20, perform time-frequency conversion on the audio frame to generate Mel spectrum features;

[0048] In this embodiment, the time-frequency conversion of the audio frame is a key step in voice signal processing, and its purpose is to convert the time-domain signal into a frequency-domain representation for subsequent feature extraction and analysis. Since the voice signal is a non-stationary signal that changes over time, directly performing spectral analysis on the entire signal may not accurately reflect the time-varying characteristics of the signal. Therefore, the method of short-time Fourier transform (STFT) is often used to divide the audio frame into short time periods so that it can be considered a steady-state signal within each short time period, enabling more effective spectral analysis.

[0049] The calculation of the short-time Fourier transform involves applying a window function to the audio frame to reduce the impact of spectral leakage. Common window functions include the Hamming window, Hanning window, and Gaussian window, etc. Different window functions have different characteristics in terms of time-domain smoothness and frequency-domain resolution. For example, the Hamming window can reduce spectral leakage and improve frequency resolution, and is suitable for precise spectral analysis, while the Hanning window performs better in terms of smoothness and is more suitable for the processing of speech signals. The audio frame after applying the window function is calculated by the fast Fourier transform (FFT) to obtain the complex spectrum, which consists of the magnitude spectrum and the phase spectrum. The magnitude spectrum represents the energy distribution of the signal and reflects the intensity of different frequency components, while the phase spectrum describes the phase relationship between different frequency components.

[0050] To reduce the computational complexity and improve the stability of features, usually only the magnitude spectrum is retained and further mapped to the Mel spectrum. The Mel spectrum is a spectral representation that conforms to the auditory characteristics of the human ear. Its frequency resolution decreases with the increase of frequency and can better simulate the human ear's perception ability of different frequencies. The calculation process of the Mel spectrum includes weighted summation of the magnitude spectrum through a set of Mel filter banks. Each filter corresponds to a specific frequency range and adopts a triangular filter shape to ensure smooth transition between adjacent filters. Finally, after logarithmic compression and normalization processing, the Mel spectrum features for subsequent speech enhancement are obtained.

[0051] In different application scenarios, the parameters of the time-frequency conversion can be optimized according to the computing resources, real-time requirements, and signal characteristics. For example, on devices with limited computing resources (such as embedded systems, mobile devices), a shorter window length (such as 20 ms) and a larger frame shift (such as 10 ms) can be adopted to reduce the computational burden, while in a high-performance server or cloud computing environment, a longer window length (such as 25 ms) and a finer frequency resolution can be used to obtain higher-quality Mel spectrum features.

[0052] The selection of the window function for the short-time Fourier transform can be adjusted according to the target application. For example, in a high-noise environment (such as industrial scenarios, remote conferences), a longer window length can be adopted to improve the robustness of the signal, while in low-latency speech processing (such as real-time communication, speech recognition), a short window length and a fast calculation method (such as sliding FFT) can be used to reduce the processing delay. During the calculation process of the Mel spectrum, the number and distribution of Mel filters can also be adjusted. For example, usually 40 filters are adopted in speech recognition tasks, while 80 filters may be used in high-quality speech synthesis or enhancement tasks to improve the sensitivity to subtle frequency changes.

[0053] For the normalization of Mel spectrograms, different methods can be selected, such as Mean-Variance Normalization (MVN), logarithmic dynamic range normalization, etc., to improve stability in different speech environments. For example, in the scenario of telephone speech processing, since different telephone channels may cause large variations in speech amplitude and spectral range, the mean-variance normalization method can be used to make the audio recorded by different devices have a consistent feature distribution and improve the generalization ability of the speech enhancement system.

[0054] By performing time-frequency conversion on audio frames, the frequency components of speech signals can be analyzed more effectively, improving the processing ability of speech enhancement algorithms for different frequency information. Using the short-time Fourier transform can reduce the computational complexity while ensuring time-frequency resolution, improving the real-time performance of speech enhancement. Using the Mel filter bank for spectral transformation makes the features more in line with human auditory perception, improving the naturalness and intelligibility of speech enhancement. Normalization processing can reduce the influence of different devices and environments on speech features, improving the robustness and generalization ability of the system in different scenarios.

[0055] S30, extracting the frequency-domain multi-scale features of the Mel spectrogram features through a multi-scale convolutional neural network;

[0056] In this embodiment, extracting the frequency-domain multi-scale features of Mel spectrogram features through a multi-scale convolutional neural network is a crucial step in the speech enhancement process, aiming to fully utilize the information in different frequency ranges and improve the model's ability to capture the details of speech signals. The features of speech signals in different frequency ranges have different physical meanings. The low-frequency part mainly contains the pitch and formant information of speech, while the high-frequency part contains rich speech texture and clarity information. Therefore, relying solely on a single-scale convolutional kernel for feature extraction may result in the loss of some important frequency-domain information and affect the effect of speech enhancement.

[0057] The core idea of a multi-scale convolutional neural network is to extract frequency-domain features in parallel at different scales. By combining convolutional kernels of different sizes, it captures the feature information in different frequency ranges. In implementation, a multi-branch convolutional structure is usually adopted, and each branch is responsible for processing information at different scales. For example, a narrow-band convolutional kernel (such as a 3×3 convolution) can be used to extract local detail information, suitable for capturing high-frequency features; a medium-band convolutional kernel (such as a 5×5 convolution) can be used to extract features in the medium-frequency range to balance global information and local details; a wide-band convolutional kernel (such as a 7×7 or 9×9 convolution) can be used to extract global frequency information and enhance the model's ability to capture low-frequency components.

[0058] After multi-scale feature extraction, it is necessary to fuse features of different scales to form a more complete frequency-domain representation. The fusion strategies usually include two methods: channel concatenation and weighted summation. The channel concatenation method combines the features extracted by different convolutional branches in the channel dimension, enabling the model to comprehensively utilize information of different scales. The weighted summation method, on the other hand, performs weighted fusion of information of different scales by learning the importance weights of features of different scales, allowing the model to automatically focus on the most discriminative features. In addition, to enhance the expressive power of the model, a cross-channel attention mechanism is usually introduced. By calculating the correlation between channels, the weights of different feature channels are dynamically adjusted to improve the adaptability of feature representation.

[0059] In practical applications, the normalization process of convolutional neural networks is also essential. Since the numerical distributions of convolutional features of different scales may vary, direct fusion may lead to uneven feature distributions, affecting the stability of the model. Therefore, before multi-scale feature fusion, batch normalization (BN) or layer normalization (LN) is usually adopted to make the feature distributions of different scales more consistent and improve the stability of training. At the same time, to introduce non-linear mapping capabilities, after normalization, a non-linear activation function (such as ReLU or Swish) is usually used to increase the expressive power of the network and improve the modeling ability for complex speech features.

[0060] Extracting the frequency-domain multi-scale features of speech through a multi-scale convolutional neural network can enhance the model's ability to perceive information in different frequency ranges and improve the effect of speech enhancement. Using convolutional kernels of different scales can ensure that the model can capture high-frequency details while retaining low-frequency structural information, improving the clarity and naturalness of speech. Multi-scale feature fusion can make full use of features in different frequency ranges, improve the generalization ability of the model, and enable it to adapt to different noise environments. Introducing a cross-channel attention mechanism can dynamically adjust the weights of feature channels, increase the attention to key information of speech signals, and further enhance the effect of speech enhancement. In addition, normalization and non-linear mapping can improve the training stability and expressive power of the model, making it more adaptable in different application scenarios.

[0061] S40, encoding the frequency-domain multi-scale features into a dimension-reduced latent representation;

[0062] In this embodiment, encoding the frequency-domain multi-scale features into a dimension-reduced latent representation is a crucial step in the speech enhancement process. The purpose is to reduce the data dimension, remove redundant information, and at the same time retain the key speech features to improve the computational efficiency and enhancement effect of subsequent processing. In speech signal processing, the information in different frequency ranges may have different importance. Directly using the complete frequency-domain multi-scale features for subsequent calculations may lead to too high computational complexity and increase the sensitivity of the model to redundant information. Therefore, by adopting a dimension-reducing encoding method, feature representations with higher information content can be extracted, improving the performance and robustness of the system.

[0063] The encoding process usually adopts technical means such as convolutional neural networks (CNNs), variational autoencoders (VAEs), or self-attention mechanisms (Transformers) to perform dimension reduction on the multi-scale features. To ensure that the information in different frequency ranges is reasonably represented, the encoder usually adopts a hierarchical structure to gradually extract high-level speech features. In implementation, the first layer of the encoder is usually a local feature extraction layer, using small-sized convolutional kernels (such as 3×3) to capture local frequency-domain information and extract short-term local patterns; the second layer is an intermediate feature extraction layer, using medium-sized convolutional kernels (such as 5×5) to extract patterns within a longer time window to improve the robustness of the features; the third layer is a global feature extraction layer, using large-sized convolutional kernels (such as 7×7) to model the information in the entire frequency-domain range to obtain the overall representation of the speech signal.

[0064] In the encoding process, feature fusion is an indispensable step. Since the features at different levels have different physical meanings, they need to be fused to form a more discriminative latent representation. Common fusion methods include cross-layer connections, channel concatenation, and attention mechanisms. For example, cross-layer connections (such as the ResNet structure) can be used to combine the features of the shallow and deep layers, retaining both the detailed information of the low layer and integrating the global information of the high layer to improve the expression ability of the model. In addition, a channel attention mechanism can be introduced. By calculating the weights of different channels, the model can adaptively focus on the most informative features, thereby improving the robustness of the encoding.

[0065] To further optimize the dimension reduction process, normalization is usually required to ensure the consistency of the distributions of different feature components and improve the stability of training. Common normalization methods include batch normalization (BN), layer normalization (LN), and instance normalization (IN), and different normalization methods are suitable for different scenarios. For example, BN is suitable for small-batch training to improve the convergence speed of the model, while LN is suitable for self-attention mechanisms to improve the sequence modeling ability. In addition, to enhance the non-linear expression ability, activation functions (such as ReLU, Leaky ReLU, or Swish) are usually applied after normalization to increase the model's fitting ability for complex patterns.

[0066] The final step of dimensionality reduction usually employs a linear projection layer to map high-dimensional features to a low-dimensional latent space through a fully connected layer or 1×1 convolution, generating the final dimensionality-reduced latent representation. The role of linear projection is to further reduce the redundancy of features, enabling the model to efficiently process key speech features, while reducing computational complexity and improving the real-time performance of the system.

[0067] By encoding and reducing the dimensionality of multi-scale features in the frequency domain, redundant information can be reduced, the computational efficiency of the model can be improved, and at the same time, the discriminability of speech features can be enhanced. Using a hierarchical encoder structure can ensure that information in different frequency ranges is fully expressed, improving the effect of speech enhancement. Feature fusion strategies (such as cross-layer connections and channel attention) can further optimize the dimensionality-reduced representation, enabling the model to focus on the most important speech features and improving the robustness of speech enhancement. Normalization processing can improve the stability of training, making the model more adaptable in different environments. In addition, the linear projection layer can reduce the feature dimension, improve the inference speed, and enable the model to run stably in real-time applications.

[0068] S50, performing noise suppression processing on the dimensionality-reduced latent representation through a deep residual network to generate enhanced audio features;

[0069] In this embodiment, the deep residual network is used to perform noise suppression processing on the dimensionality-reduced latent representation, aiming to reduce the background noise in the speech signal and improve the clarity and intelligibility of the enhanced audio features. The deep residual network is a neural network architecture widely used in signal processing and computer vision. The core idea is to reduce the vanishing gradient problem in deep networks through residual connections while enhancing the modeling ability. In speech enhancement tasks, the deep residual network can learn complex noise patterns and extract a cleaner speech signal based on the dimensionality-reduced latent representation while minimizing the damage to the speech content.

[0070] Taking the dimensionality-reduced latent representation as the input, it first passes through an initial convolutional layer, which is used to extract local features and expand the receptive field, improving the sensitivity to noise in different frequency ranges. Subsequently, the input passes through multiple residual blocks, each of which includes two or more convolutional layers and skip connections, enabling the network to directly learn the residual between the input and output. The advantage of residual learning is that the model can focus on learning the noise components rather than the complete speech signal, thus retaining more original speech features during the enhancement process and improving the quality of the enhanced audio.

[0071] The convolutional layers of deep residual networks usually adopt 1D or 2D convolutions. Among them, 1D convolution is suitable for processing time-series signals, and 2D convolution is more suitable for spectral-domain feature extraction. To adapt to different noise environments, the network can adopt dynamic normalization (BatchNormalization, Instance Normalization) methods to ensure the model stability under different input conditions. In addition, non-linear activation functions (such as ReLU, Leaky ReLU, Swish) are used to enhance the model's expressive ability, enabling it to more effectively distinguish speech signals and noise components.

[0072] Based on residual learning, the final output of noise suppression needs to be optimized through a post-processing layer. Common post-processing techniques include the gated mechanism, which is used to further enhance the noise separation ability, and the attention mechanism, which is used to dynamically adjust the weights of different frequency bands to improve the adaptability of noise suppression. Finally, the generated enhanced audio features exhibit a better signal-to-noise ratio in the frequency domain or time domain, enabling subsequent speech synthesis or speech recognition tasks to obtain clearer inputs.

[0073] Example illustration: In the field of medical and health services, when doctors conduct remote consultations, the speech signal may be interfered by hospital equipment noise, patient conversations, etc. By using a deep residual network to perform noise suppression on the dimensionality-reduced latent representation, these background noises can be effectively removed, making the doctor's speech clearer and improving the communication quality of remote consultations. In addition, in medical voice recordings, the doctor's speech may be accompanied by echoes or environmental noise. The noise suppression module can optimize the voice quality, ensure the accuracy of medical records, and improve the reliability of medical data.

[0074] In the field of fintech services, a bank's voice customer service system needs to process users' voice requests in various environments. For example, in noisy environments such as airports and subways, the speech recognition system may have difficulty accurately recognizing the customer's speech. By using a deep residual network for noise suppression, background noise can be removed, improving the accuracy of speech recognition and enabling the customer service system to handle business more smoothly. In addition, during financial transactions, some voice authentication technologies rely on high-quality inputs of voice features. The noise suppression ability of the deep residual network can improve the security of voice identity recognition and prevent recognition errors caused by environmental noise.

[0075] In the field of speech enhancement technology, a remote conferencing system needs to ensure that the voices of all speakers are clear and intelligible. However, in a multi-person meeting, a microphone may pick up ambient noise, keyboard tapping sounds, or other interfering sounds. The noise suppression process of a deep residual network can enhance the intelligibility of meeting speech, making the spoken content clearer and improving the accuracy of meeting records and speech transcription. In addition, in speech translation applications, clear speech input can significantly improve the accuracy of the translation system, making cross-language communication smoother.

[0076] By performing noise suppression processing through a deep residual network, background noise in the speech signal can be effectively reduced, the clarity of the speech signal can be improved, and at the same time, damage to the speech content can be avoided. Residual learning enables the model to focus on learning the noise components, thereby more accurately retaining the target speech information during the speech enhancement process. The adoption of multiple residual blocks and skip connections can improve the training efficiency of the network, reduce the vanishing gradient problem, and enable the model to learn deeper speech features. In addition, through the cooperation of dynamic normalization and non-linear activation functions, the stability and adaptability of the model can be improved, enabling it to maintain a good enhancement effect in different noise environments.

[0077] S60, converting the enhanced audio features into an enhanced Mel spectrogram through a non-autoregressive generation model;

[0078] In this embodiment, converting the enhanced audio features into an enhanced Mel spectrogram through a non-autoregressive generation model aims to improve the efficiency and quality of speech enhancement, enabling the model to quickly generate high-quality speech features with low latency. Although the autoregressive method can generate high-quality speech features, due to its dependence on the output of the previous moment and the serial dependence in the calculation, the inference speed is slow, making it difficult to meet the requirements of real-time speech enhancement. Therefore, the non-autoregressive (Non-Autoregressive, NAR) method is introduced to reduce the calculation time while maintaining the generated speech quality.

[0079] A non-autoregressive generation model usually consists of multiple core modules, including a multi-scale temporal convolution module, a time-frequency transformation module, a fundamental frequency prediction module, a frequency band compensation module, a multi-head self-attention mechanism, an adversarial training discriminator, etc. First, the enhanced audio features are input into the multi-scale temporal convolution module, which extracts local and global temporal context features through convolutional kernels of different scales, thereby enhancing the ability to model the temporal dynamic changes of the speech signal. Multi-scale temporal convolution can ensure that the model can capture both short-term dependencies and model long-term temporal relationships, improving the stability of speech enhancement.

[0080] In order to further optimize the generation effect, the non-autoregressive model introduces a time-frequency transformation module, which performs time-frequency conversion on the time series features to obtain frequency domain information, and combines it with the fundamental frequency prediction module to extract the fundamental frequency features. The fundamental frequency prediction module is used to analyze the fundamental frequency and harmonic peak of the speech, and generate a harmonic distribution mask to guide subsequent frequency adjustment to make the generated Mel spectrum more natural. Based on the harmonic distribution mask, the system uses a gated recurrent unit (GRU) to calculate the frequency band adaptive compensation weight matrix, and multiplies the matrix with the time series context features band by band to compensate for the energy changes in different frequency bands and improve the accuracy of spectrum conversion.

[0081] The non-autoregressive generative model also integrates a multi-head self-attention mechanism that can model the global dependencies between different frequency components, so that the enhanced Mel spectrum can retain more speech features and reduce artifacts and sound quality loss. During the model training process, an adversarial training discriminator is used to optimize the generated results. The discriminator calculates the difference between the generated Mel spectrum and the real speech data through the multi-scale short-time Fourier transform (STFT) consistency loss, and optimizes the parameters of the non-autoregressive generative model through gradient backpropagation to improve the quality of speech enhancement.

[0082] Finally, after power spectrum normalization and Mel spectrum mapping, the optimized features are converted to Mel band representation. In the Mel spectrum mapping process, the normalized optimized features are mapped to the Mel band dimension using a linear projection layer, and the spectrum is reconstructed through a Mel filter bank to generate the final enhanced Mel spectrum. In order to ensure the stability of the conversion process, a pseudo-inverse transform (such as singular value decomposition SVD) can also be used to optimize the inverse operation of the Mel filter matrix, reduce the conversion error, and improve the fidelity of the Mel spectrum.

[0083] Mel-spectrogram conversion through non-autoregressive generative models can improve the computational efficiency of speech enhancement and reduce computational overhead, so that the enhancement process can be completed under low latency conditions. The use of multi-scale time convolution and self-attention mechanisms can enhance the model's ability to model the time-frequency structure of speech and improve the clarity and naturalness of speech. Fundamental frequency prediction and frequency band compensation methods can optimize harmonic features, making the generated Mel-spectrogram closer to real speech and improving speech intelligibility and sound quality. Adversarial training mechanisms can reduce speech artifacts and improve the overall stability of speech enhancement.

[0084] S70, reconstructing the enhanced Mel spectrum into a target speech waveform through a generative adversarial network.

[0085] In this embodiment, reconstructing the enhanced Mel spectrogram into the target speech waveform through a generative adversarial network is a crucial audio signal reconstruction step in the speech enhancement system. Since the Mel spectrogram only contains the amplitude information of the speech signal, and the complete representation of speech also requires phase information, when converting the enhanced Mel spectrogram into a speech waveform, additional phase estimation and waveform reconstruction techniques need to be introduced to ensure the clarity, naturalness, and intelligibility of the speech. Traditional methods such as the Griffin-Lim algorithm rely on iterative optimization to recover the phase, but due to the high computational complexity and the fact that the recovered phase may deviate significantly from the true speech phase, the reconstructed speech is distorted. To solve these problems, a generative adversarial network (GAN) is used for speech waveform reconstruction, leveraging the collaborative optimization of the generator and discriminator to make the synthesized speech more realistic, clear, and reduce artifacts.

[0086] This process first performs phase prediction processing on the enhanced Mel spectrogram to recover the initial phase information. Phase prediction can use a phase reconstruction network to learn the phase characteristics of the speech signal based on a deep neural network, or a data-driven phase compensation method to estimate the phase through known speech data. After the initial phase is generated, a complex convolutional network is further used to correct the phase. The complex convolutional network can model both the amplitude information and the phase information simultaneously, making the phase prediction more accurate and thus reducing the phase distortion during waveform reconstruction.

[0087] Subsequently, the optimized phase spectrum is combined with the amplitude information of the enhanced Mel spectrogram to generate a complex spectrum, which contains the complete time-frequency information of the speech signal, making the subsequent waveform reconstruction more accurate. This complex spectrum is transformed back to the time domain through the inverse short-time Fourier transform (ISTFT) to generate a preliminary time-domain waveform. During the calculation of the inverse short-time Fourier transform, the overlap-add strategy is used to ensure the continuity of the splicing of different time windows, reduce the artifacts in the overlapping part of the speech waveform, and improve the fluency of the speech.

[0088] Based on the generated time-domain waveform, a generative adversarial network is introduced for optimization. The generator is used to generate a more natural speech waveform, and the discriminator is used to distinguish the difference between the generated speech and the real speech and provide feedback to optimize the generator. The discriminator adopts a multi-scale short-time Fourier transform (STFT) consistency loss to calculate the error of the speech signal at different time scales, thereby optimizing the generator to enable it to generate a more realistic speech waveform. In addition, to further improve the reconstruction quality, a perceptual loss can be introduced to make the generated speech more perceptually close to the real speech while reducing computational artifacts and improving the naturalness of the speech.

[0089] The conversion from Mel spectrogram to speech waveform through a generative adversarial network can effectively improve the quality of speech enhancement, making the generated speech waveform closer to real speech.

[0090] The present invention relates to the technical field of speech processing and can be applied to business scenarios such as medical health and fintech. It discloses a speech enhancement method based on multi-scale feature learning, including: obtaining an input audio signal and performing frame division processing to generate audio frames; performing time-frequency conversion on the audio frames to extract Mel spectrogram features; extracting frequency-domain multi-scale features through a multi-scale convolutional neural network; encoding the extracted frequency-domain multi-scale features to obtain a dimension-reduced latent representation; using a deep residual network to suppress noise in the dimension-reduced latent representation to obtain enhanced audio features; converting the enhanced audio features into enhanced Mel spectrograms through a non-autoregressive generative model; and reconstructing the enhanced Mel spectrograms into target speech waveforms using a generative adversarial network. The present invention extracts speech frequency-domain features through a multi-scale convolutional neural network to improve the feature expression ability of different frequency bands; suppresses noise through a deep residual network to enhance the purity of the speech signal; optimizes feature conversion through a non-autoregressive generative model to improve the modeling efficiency of speech enhancement; and reconstructs target speech waveforms through a generative adversarial network to improve the naturalness and clarity of the generated speech.

[0091] In one embodiment, the above step S10 includes:

[0092] S101, performing pre-emphasis filtering on the input audio signal to generate a pre-emphasized audio signal;

[0093] S102, dividing the pre-emphasized audio signal into multiple pre-emphasized audio frames according to a preset frame length and a preset frame shift;

[0094] S103, performing windowing on each pre-emphasized audio frame through a Hamming window processing module to generate windowed audio frames;

[0095] S104, detecting the short-time energy of the windowed audio frames and filtering out the silent frames in the windowed audio frames according to a preset energy threshold to generate the audio frames after frame division processing.

[0096] In this embodiment, obtaining the input audio signal and performing frame division processing is a basic step in speech signal processing. Its purpose is to convert the continuous audio signal into fixed-length segments suitable for subsequent processing for time-frequency analysis and feature extraction. Since the speech signal is a continuous waveform that changes over time, directly processing the entire audio segment may lead to too high computational complexity and it is difficult to capture the dynamic changes in speech. Therefore, by using the frame division processing method, the long-time signal can be split into small windows, which can improve the processing efficiency while ensuring the temporal integrity of the signal.

[0097] First, perform pre-emphasis filtering on the input audio signal to compensate for high-frequency attenuation and improve the high-frequency resolution of the signal. The pre-emphasis filter usually adopts a first-order high-pass filter, and its mathematical expression is:

[0098] y(n) = x(n) - αx(n - 1)

[0099] where x(n) is the original audio signal, y(n) is the pre-emphasized signal, and α is the filtering coefficient, usually taking values between 0.95 and 0.98. The role of pre-emphasis filtering is to reduce the low-frequency components of the speech signal, make the high-frequency signal more prominent, thereby improving the feature discrimination ability of speech. Especially in the subsequent spectrum analysis and speech enhancement processes, it can better extract key features.

[0100] Subsequently, the pre-emphasized audio signal is segmented according to the preset frame length and frame shift to generate multiple pre-emphasized audio frames. The frame length determines the duration of each speech frame, and the frame shift determines the overlap degree between adjacent frames. Generally, the frame length is selected from 20 ms to 30 ms to ensure that each frame contains sufficient information while avoiding a decrease in time-domain resolution; the frame shift is usually set to about 10 ms to ensure sufficient overlap between adjacent frames, thereby retaining the continuity of speech.

[0101] After the frame segmentation process, each audio frame will undergo Hamming Window processing to reduce the spectral leakage problem caused by frame-by-frame cutting. Finally, short-time energy detection is performed on the windowed audio frames, and silent frames are filtered according to the preset energy threshold. By setting a reasonable energy threshold, meaningless silent frames can be removed, thereby reducing redundant data and improving the efficiency of subsequent speech processing.

[0102] In this embodiment, through the frame segmentation process of the audio signal, the processing efficiency of the speech signal can be effectively improved, while retaining the timing information of the speech, making the subsequent feature extraction and enhancement processes more stable. Pre-emphasis filtering can enhance the high-frequency components of speech, improve speech clarity, and enable the speech enhancement algorithm to better model speech features. The windowing process based on the Hamming window can reduce spectral leakage and improve the accuracy of frequency-domain analysis, enabling more important information to be retained during the speech enhancement process. Short-time energy detection and silent frame filtering can reduce the computational overhead and improve the efficiency of speech processing, enabling the system to still operate efficiently in a low-computing-resource environment.

[0103] In one embodiment, the above step S20 includes:

[0104] S201, perform a short-time Fourier transform on the audio frame to generate a linear spectrum;

[0105] S202, separate the magnitude spectrum and phase spectrum of the linear spectrum, and retain the magnitude spectrum;

[0106] S203. Map the amplitude spectrum to a Mel spectrum through a Mel filter bank;

[0107] S204. Perform logarithmic compression processing on the Mel spectrum to generate a log Mel spectrum;

[0108] S205. Perform mean - variance normalization processing on the log Mel spectrum to generate the Mel spectrum features.

[0109] In this embodiment, performing time - frequency conversion on the audio frame and generating Mel spectrum features is a core step in speech processing. The purpose is to convert the time - domain signal into a more discriminative frequency - domain representation, making the subsequent speech enhancement, recognition, or synthesis process more efficient and stable. The Mel spectrum is a frequency representation method based on the human auditory perception. It can more accurately capture the timbre and pitch characteristics of speech and is applicable to various speech enhancement applications.

[0110] First, perform a short - time Fourier transform (STFT) on the audio frame to convert the time - domain signal into a frequency - domain signal to obtain the linear spectrum. The short - time Fourier transform uses a fixed - length time window, calculates the frequency components of the audio signal within each time window, and extracts data of consecutive frames by sliding the window, enabling the frequency - domain representation to reflect the dynamic change characteristics of speech. This step can map the speech data from the time domain to the frequency domain without losing the timing information, providing a basis for subsequent speech enhancement.

[0111] Subsequently, separate the amplitude spectrum and the phase spectrum in the linear spectrum, and retain the amplitude spectrum for feature extraction. The amplitude spectrum reflects the energy distribution of each frequency component in the audio signal, while the phase spectrum is mainly used to restore the time - domain waveform. In speech enhancement tasks, usually only the amplitude spectrum is used for processing, and the phase information is compensated during speech reconstruction.

[0112] Next, use a Mel filter bank to map the amplitude spectrum to generate a Mel spectrum. The Mel filter bank consists of a series of band - pass filters, which can simulate the human ear's perception of different frequencies, resulting in higher resolution in the low - frequency part and lower resolution in the high - frequency part, thus better conforming to the auditory characteristics of speech signals. The design of the Mel filter bank makes the extracted features closer to the actual perception of speech by the human ear, improving the effectiveness of the speech enhancement system.

[0113] To further compress the dynamic range of the data, the Mel spectrum undergoes logarithmic compression processing to generate a log Mel spectrum. The logarithmic transformation can reduce the amplitude variation of the speech signal energy, making larger energy changes more prominent and smaller energy changes having less impact on the features, thereby improving the stability of speech enhancement and enhancing the model's ability to capture speech details.

[0114] Finally, perform mean-variance normalization on the logarithmic Mel spectrogram to generate the final Mel spectrogram features. The purpose of mean-variance normalization is to adjust the distribution of the features so that their mean is zero and variance is one, thereby reducing the impact of different devices, environments, and recording conditions on speech features, improving the generalization ability of the speech enhancement algorithm, and enabling the system to more robustly adapt to different application scenarios.

[0115] In this embodiment, time-frequency conversion is performed through short-time Fourier transform, which can effectively convert the time-domain signal into more stable frequency-domain features, enabling the speech enhancement model to better extract speech information. The separation processing of the amplitude spectrum and phase spectrum reduces the interference of phase information on the enhancement process and improves the stability of the enhancement result. The application of the Mel filter bank can simulate the auditory characteristics of the human ear, making the enhanced speech more conform to the perception standard of human hearing and improving the speech quality. Logarithmic compression can reduce the dynamic range, improve the calculation efficiency, and enhance the robustness of the speech signal. Mean-variance normalization can reduce the impact of the environment and devices on speech features, improve the stability of the speech enhancement algorithm, enable the system to adapt to different application scenarios, and improve the intelligibility and naturalness of the enhanced speech signal.

[0116] In one embodiment, the above step S30 includes:

[0117] S301, input the Mel spectrogram features into a multi-scale convolutional neural network, where the multi-scale convolutional neural network includes a narrow-band convolution branch, a mid-band convolution branch, and a wide-band convolution branch;

[0118] S302, use a convolution kernel of a first preset size through the narrow-band convolution branch to extract the narrow-band frequency-domain features of the Mel spectrogram features;

[0119] S303, use a convolution kernel of a second preset size through the mid-band convolution branch to extract the mid-band frequency-domain features of the Mel spectrogram features;

[0120] S304, use a convolution kernel of a third preset size through the wide-band convolution branch to extract the wide-band frequency-domain features of the Mel spectrogram features;

[0121] S305, through the feature fusion module of the multi-scale convolutional neural network, splice the narrow-band frequency-domain features, mid-band frequency-domain features, and wide-band frequency-domain features in the channel dimension to generate multi-scale fusion features;

[0122] S306, through the cross-channel attention module of the multi-scale convolutional neural network, perform channel attention weight processing on the multi-scale fusion features to generate channel attention-weighted multi-scale features;

[0123] S307. Normalize the multi-scale features weighted by channel attention through the normalization module of the multi-scale convolutional neural network to generate normalized multi-scale features;

[0124] S308. Perform non-linear mapping on the normalized multi-scale features through the non-linear activation module of the multi-scale convolutional neural network to generate the frequency-domain multi-scale features.

[0125] In this embodiment, a multi-scale convolutional neural network is used to extract the frequency-domain multi-scale features of the Mel spectrogram features, aiming to enhance the expressive ability of the speech signal, enabling the model to capture speech information in different frequency ranges and improving the overall effect of speech enhancement. The Mel spectrogram is a feature representation method of the speech signal, which converts the speech signal into a spectrogram representation that conforms to the auditory characteristics of the human ear. However, due to the wide frequency range of the speech signal, the information in different frequency intervals may have different characteristics. Therefore, a multi-scale convolutional neural network (Multi-Scale Convolutional Neural Network, MSCNN) is needed for feature extraction to effectively capture the detailed information of the speech.

[0126] Features in different frequency ranges have different physical meanings and roles in speech perception. Therefore, when using a multi-scale convolutional neural network to extract Mel spectrogram features, the narrowband, midband, and wideband frequency ranges need to be considered separately to optimize the speech enhancement effect. Narrowband features mainly focus on a relatively small frequency range, generally between 300 Hz and 1 kHz. Its role is to extract the local detailed information of the speech, such as the transient features and fine structure information of consonant phonemes. These features have an important impact on the clarity and recognition of the speech. In scenarios such as telephone communication and low-bandwidth speech transmission, narrowband signals are mainly used to transmit basic speech information. However, due to their strong low-frequency components and lack of high-frequency information, the speech may sound monotonous and a lot of details may be lost. Therefore, when the narrowband convolution branch extracts features in this frequency range, a relatively small-sized convolutional kernel is usually used to ensure that fine speech features can be effectively captured and the accuracy of speech enhancement can be improved.

[0127] The mid-band features generally cover the frequency range from 1 kHz to 4 kHz, which contains the main energy distribution of speech, can reflect the overall timbre characteristics and clarity of speech, and plays a core role in the speech enhancement task. This frequency band not only includes the main formant information of speech, but also involves the clarity and naturalness of speech. For remote voice communication, speech recognition, and speech enhancement systems, the optimization of this frequency range is the key to improving speech intelligibility. Therefore, the mid-band convolution branch usually adopts a moderate convolution kernel size to ensure that it can capture the structural information in the mid-frequency range and reduce possible speech distortion during the feature extraction process. In addition, the features in this frequency band are particularly important for applications such as intelligent voice assistants and virtual customer service systems, because these systems need to accurately identify and process speech information in different environments.

[0128] The wide-band features generally cover the range from 4 kHz to 8 kHz or higher (such as 16 kHz or even 20 kHz), and are mainly used to capture the high-frequency information of speech, including the prosody, pitch variation, energy distribution, and ultra-high-frequency harmonics of speech. The wide-band features are crucial for high-quality speech enhancement, emotion analysis, speech synthesis, and high-definition voice call systems. The high-frequency components can enhance the naturalness of speech, making the speech sound fuller and more real. However, the high-frequency features are usually vulnerable to environmental noise interference. Therefore, in the wide-band convolution branch, a larger convolution kernel size is usually adopted to ensure that global features can be effectively extracted and the stability of speech enhancement can be improved. In addition, in order to adapt to different application scenarios, such as remote meetings and automatic caption generation, the optimized processing of wide-band information can enhance the fidelity of speech, making the enhanced speech clearer and more natural.

[0129] This process first inputs the Mel spectrum features into a multi-scale convolutional neural network, and performs feature extraction through the narrow-band convolution branch, mid-band convolution branch, and wide-band convolution branch respectively:

[0130] The narrow-band convolution branch uses a convolution kernel of the first preset size to perform convolution operations on the Mel spectrum with a smaller receptive field, so as to extract local high-precision frequency domain information. The small-scale convolution kernel helps to capture the minute changes in the speech signal, such as the detailed features at the phoneme level.

[0131] The mid-band convolution branch uses a convolution kernel of the second preset size to perform convolution on the Mel spectrum within a larger receptive field range to extract mid-frequency features. This convolution kernel can better capture the overall timbre information of speech and improve the quality of speech enhancement.

[0132] The wide-band convolution branch uses a convolution kernel of the third preset size to perform convolution on the features of the entire frequency range, extracts global feature information, enables the model to learn the overall structure of the speech signal, and improves the naturalness of the enhanced speech.

[0133] After the feature extraction of different branches is completed, the feature fusion module stitches the frequency-domain features of the narrowband, midband, and wideband in the channel dimension to generate multi-scale fusion features. The purpose of feature fusion is to synthesize information in different frequency ranges, enabling the enhancement model to make full use of local and global information in the spectrum and improving the stability of the enhancement effect.

[0134] Next, the cross-channel attention module processes the channel attention weights of the multi-scale fusion features to enhance the expression ability of important features. The cross-channel attention mechanism can automatically identify the most important frequency-domain components in the speech signal and assign higher weights, thereby improving the model's perception ability of key speech features.

[0135] Then, the normalization module normalizes the multi-scale features weighted by channel attention to generate normalized multi-scale features. The purpose of normalization is to adjust the distribution of feature data so that its mean and variance remain within a stable range, thereby improving the training stability of the model and reducing feature deviations in different speech environments.

[0136] Finally, the non-linear activation module performs non-linear mapping on the normalized multi-scale features to generate the final frequency-domain multi-scale features. The non-linear activation function can enhance the non-linear expression ability of speech features, improve the model's modeling ability for complex speech signals, and thus enhance the effect of speech enhancement.

[0137] In this embodiment, the multi-scale convolutional neural network extracts the frequency-domain multi-scale features of the Mel spectrogram, which can effectively improve the effect of speech enhancement, enabling the model to capture local details and global features of the speech signal more comprehensively. The combination of narrowband, midband, and wideband convolutions enables the model to learn the features of the speech signal at different frequency scales, thereby improving the robustness and generalization ability of speech enhancement. The feature fusion module can make full use of information in different frequency ranges and improve the model's comprehensive feature extraction ability. The cross-channel attention mechanism can optimize the weight allocation of important features, enabling the model to pay more attention to key speech information and improving the effect of speech enhancement. Normalization processing and non-linear mapping can enhance the stability of the model, enabling it to adapt to different speech signal environments and improving the robustness of speech enhancement.

[0138] In one embodiment, the above step S40 includes:

[0139] S401, input the frequency-domain multi-scale features into the first convolutional layer of the encoder to generate first-level local frequency-domain features;

[0140] S402, input the first-level local frequency-domain features into the second convolutional layer of the encoder to generate second-level intermediate frequency-domain features;

[0141] S403. Input the intermediate frequency domain features of the second level into the third convolutional layer of the encoder to generate the global frequency domain features of the third level;

[0142] S404. Concatenate the local frequency domain features of the first level with the global frequency domain features of the third level through a cross-layer connection module to generate the fused frequency domain features;

[0143] S405. Perform normalization processing on the fused frequency domain features to generate the normalized frequency domain features;

[0144] S406. Perform a non-linear transformation on the normalized frequency domain features through a non-linear activation function to generate the activated frequency domain features;

[0145] S407. Input the activated frequency domain features into a linear projection layer for dimensionality reduction processing to generate the dimensionality-reduced latent representation.

[0146] In this embodiment, the frequency domain multi-scale features contain rich spectral information. However, due to the high dimensionality of the original features, direct processing may lead to excessive consumption of computing resources and increase the complexity of model training. Therefore, an encoder is needed to perform feature dimensionality reduction, mapping the high-dimensional features to a compact latent representation to improve the computational efficiency of the model and retain the most critical information. In this method, the frequency domain multi-scale features undergo multi-layer convolutional encoding and feature fusion, gradually extracting the frequency domain features of different levels, and finally generating the dimensionality-reduced latent representation.

[0147] First, input the frequency domain multi-scale features into the first convolutional layer of the encoder to generate the local frequency domain features of the first level. The convolutional kernel size of this layer is small, and its main function is to extract the local information of the frequency domain features, such as the detailed features of certain specific frequency components. The local frequency domain features of the first level can help the model identify the subtle changes in the speech signal, such as the clarity of consonants, the boundaries of phonemes, and other key information. The convolutional operation of this layer can improve the smoothness of the features and reduce the influence of noise, making the feature extraction of subsequent layers more stable.

[0148] Subsequently, the local frequency domain features of the first level are input into the second convolutional layer of the encoder to generate the intermediate frequency domain features of the second level. The convolutional kernel size of the second layer is relatively large, which can perform feature extraction in a wider frequency domain range, thereby capturing the global information of the speech. The intermediate frequency domain features are mainly used to represent the timbre features, formant information, and spectral energy distribution of the speech, and these information are crucial for the naturalness and enhancement effect of the speech. In addition, during the convolutional process of this layer, the local features are gradually integrated with the spectral information in a larger range, improving the model's ability to understand complex speech structures.

[0149] Next, the intermediate frequency domain features at the second level are input into the third convolutional layer of the encoder to generate the global frequency domain features at the third level. The convolutional kernel size of the third layer is larger and the receptive field is wider, which can capture the feature information within the entire frequency spectrum range. The features at this level can describe the overall prosody, energy distribution, and global patterns of the speech, ensuring that the overall structure is not overlooked during the speech enhancement process. During the modeling process of the global frequency domain features, this layer can effectively remove local instability factors and optimize the overall fluency of the speech.

[0150] Since the features at different levels contain different frequency domain information, cross-layer feature fusion is required to make full use of the feature representations at different levels. First, the local frequency domain features at the first level and the global frequency domain features at the third level are concatenated in channels through a cross-layer connection module to generate fused frequency domain features. This fusion method can combine the fine features at the low level with the global features at the high level, improve the robustness of the model, and enhance the utilization efficiency of information at different levels during the speech enhancement process. The role of the cross-layer connection module is to balance the importance of features at different levels, so that the finally generated features have both local fineness and global consistency.

[0151] The fused frequency domain features may contain feature distributions at different scales, so normalization processing is required to generate normalized frequency domain features. The role of normalization is to adjust the mean and variance of the feature data, making them have the same numerical range, improving the stability of model training, and reducing the feature deviation caused by different recording devices or environments. In addition, normalization can prevent gradient vanishing or gradient explosion and improve the generalization ability of the model under different data conditions.

[0152] To further enhance the expressive ability of the features, a non-linear activation function is used to perform non-linear transformation on the normalized frequency domain features to generate activated frequency domain features. Non-linear activation functions (such as ReLU, Leaky ReLU, or SiLU) can improve the model's ability to model complex features, enabling it to learn non-linear frequency domain change patterns and enhance the key information in the speech signal. During the speech enhancement process, the role of non-linear transformation is to strengthen the information of the speech structure while suppressing irrelevant or redundant noise information.

[0153] Finally, the activated frequency domain features are input into a linear projection layer for dimensionality compression processing to generate a reduced-dimensional latent representation. The linear projection layer can be implemented through a fully connected layer or a 1×1 convolutional layer. Its main role is to reduce the dimension of the features while retaining the effective information of the features, so that the final latent representation is both compact and can retain sufficient speech features. The reduced-dimensional latent representation can reduce the computational complexity, improve the model inference efficiency, and provide a stable input for subsequent speech enhancement steps.

[0154] In this embodiment, by encoding the frequency-domain multi-scale features into a dimensionality-reduced latent representation, the feature dimension can be effectively reduced, the computational efficiency can be improved, and the integrity of the speech features can be maintained at the same time. The hierarchical convolutional structure can fully extract the frequency-domain information at different levels, enabling the speech enhancement process to take into account both local details and global structures. Cross-layer feature fusion can integrate information at different scales, improve the robustness of the model, and reduce the information loss caused by single-scale features. The normalization process and the non-linear activation transformation can improve the stability of the model, making it adaptable to different speech data and improving the quality of the enhanced speech. In addition, the generation of the dimensionality-reduced latent representation can reduce the computational amount, make the model more efficient, and be applicable to real-time speech enhancement systems, improving the response speed and accuracy of speech processing.

[0155] In one embodiment, step S60 above includes:

[0156] S601, input the enhanced audio features into the multi-scale temporal convolutional module of the non-autoregressive generation model to generate multi-scale temporal context features;

[0157] S602, perform time-frequency transformation on the multi-scale temporal context features to generate a time-frequency representation, and perform fundamental frequency prediction analysis and harmonic peak detection based on the time-frequency representation to generate a harmonic distribution mask;

[0158] S603, according to the harmonic distribution mask, generate a frequency-band adaptive compensation weight matrix through a gated recurrent unit;

[0159] S604, multiply the frequency-band adaptive compensation weight matrix and the multi-scale temporal context features band by band to generate harmonic compensation temporal features;

[0160] S605, perform global dependency modeling on the harmonic compensation temporal features through a multi-head self-attention mechanism to generate attention-weighted temporal features;

[0161] S606, perform time-step alignment processing on the attention-weighted temporal features through linear interpolation to generate aligned temporal features;

[0162] S607, input the aligned temporal features into the adversarial training discriminator, and generate the parameter gradient of the non-autoregressive generation model through the multi-scale short-time Fourier transform spectral consistency loss analysis module of the discriminator;

[0163] S608, update the parameters of the non-autoregressive generation model based on the parameter gradient, and input the enhanced audio features into the updated non-autoregressive generation model to generate discriminator-optimized features;

[0164] S609, perform power spectrum normalization processing on the discriminator-optimized features to generate normalized optimized features;

[0165] S610. Input the normalized and optimized features into the Mel spectrogram mapping module, and map the normalized and optimized features to the Mel band dimension through the linear projection layer of the Mel spectrogram mapping module to generate Mel band features.

[0166] S611. Perform an inverse pseudo-transformation operation on the Mel band features through the Mel filter bank of the Mel spectrogram mapping module to generate the enhanced Mel spectrogram.

[0167] In this embodiment, the enhanced audio features need to be converted into an enhanced Mel spectrogram through a non-autoregressive generative model to further optimize the intelligibility and naturalness of speech. Compared with traditional autoregressive methods, the non-autoregressive generative model can predict the entire sequence in parallel, improve the inference efficiency, and reduce the cumulative error. This process includes multi-scale temporal feature extraction, harmonic information modeling, time alignment, adversarial training optimization, and Mel spectrogram mapping and reconstruction.

[0168] First, the enhanced audio features are input into the multi-scale temporal convolution module of the non-autoregressive generative model to generate multi-scale temporal context features. The multi-scale temporal convolution module consists of multiple convolutional kernels with different receptive fields, which are used to extract speech features at different time scales. For example, convolutional kernels with smaller receptive fields can capture the short-term dynamic changes of speech, while convolutional kernels with larger receptive fields can extract long-term dependencies. The role of this module is to construct a complete temporal feature representation to make subsequent harmonic analysis and compensation more accurate.

[0169] Then, perform a time-frequency transformation on the multi-scale temporal context features to generate a time-frequency representation. The time-frequency representation is calculated based on the short-time Fourier transform (STFT) or other transformation methods, which describes the frequency distribution of the speech signal at different time points. Based on this time-frequency representation, fundamental frequency prediction analysis and harmonic peak detection can be performed to generate a harmonic distribution mask. Fundamental frequency prediction is used to estimate the fundamental frequency trajectory of speech, while harmonic peak detection is used to determine the harmonic structure of the speech signal. These information are crucial for speech enhancement because the harmonic components directly affect the timbre, clarity, and naturalness of speech.

[0170] According to the generated harmonic distribution mask, use a gated recurrent unit (GRU) to generate a frequency band adaptive compensation weight matrix. The gated recurrent unit is a recurrent neural network structure that can capture temporal dependencies. Its role is to dynamically adjust the weights of different frequency bands according to the harmonic mask to compensate for information loss in the spectrum. This weight matrix can allocate different compensation coefficients in different frequency ranges to ensure that there is no spectral distortion during the speech enhancement process.

[0171] Subsequently, the band adaptive compensation weight matrix is multiplied by the multi-scale temporal context features band by band to generate harmonic compensation temporal features. The purpose of this operation is to enhance the harmonic information of the speech, so that the enhanced speech maintains the original pitch and timbre, and minimizes the deterioration of the speech quality caused by noise or distortion as much as possible.

[0172] After obtaining the harmonic compensation temporal features, the multi-head self-attention mechanism is used to model the global dependency relationship to generate the attention-weighted temporal features. The multi-head self-attention mechanism can calculate the dependency relationship between different frequency bands, thereby improving the comprehensive processing ability of the speech enhancement system for different frequency components. The introduction of this mechanism can optimize the overall coherence of the speech, making the enhanced speech more fluent and natural.

[0173] Next, the attention-weighted temporal features are processed by linear interpolation for time step alignment to generate aligned temporal features. Time step alignment is a key step in speech enhancement, especially in non-autoregressive generation models. Due to the uneven change speed of the speech signal, there may be alignment deviations in the features of different time steps. Through linear interpolation, the consistency of the temporal features with the original speech signal on the time axis can be ensured, reducing speech distortion caused by alignment errors.

[0174] Then, the aligned temporal features are input into the adversarial training discriminator, and through the multi-scale short-time Fourier transform spectral consistency loss analysis module of the discriminator, the parameter gradients of the non-autoregressive generation model are generated. The discriminator is used to evaluate whether the generated speech is real and natural, and optimizes the parameters of the non-autoregressive generation model through the loss function, making the generated enhanced Mel spectrogram closer to the Mel spectrogram of the real speech.

[0175] Based on the calculated parameter gradients, the parameters of the non-autoregressive generation model are updated, and the enhanced audio features are input into the updated non-autoregressive generation model to generate discriminator-optimized features. In this process, the non-autoregressive model will be continuously iteratively optimized to reduce the artifacts and distortions generated during the speech enhancement process.

[0176] Power spectrum normalization is performed on the generated discriminator-optimized features to generate normalized optimized features. Power spectrum normalization can balance the energy distribution of different frequency components, improve the stability of the enhanced speech, and reduce the influence of background noise.

[0177] Then, the normalized optimized features are input into the Mel spectrogram mapping module, and through the linear projection layer of the Mel spectrogram mapping module, the normalized optimized features are mapped to the Mel frequency band dimension to generate Mel frequency band features. This mapping process corresponds the enhanced speech features to the Mel spectrogram through a linear transformation, enabling further speech reconstruction.

[0178] Finally, use the Mel filter bank of the Mel spectrum mapping module to perform a pseudo-inverse transformation operation on the Mel band features to generate an enhanced Mel spectrum. Pseudo-inverse transformation refers to recovering a representation as close as possible to the original data from the dimensionality-reduced features through a certain mathematical method. In this method, the pseudo-inverse calculation of the Mel filter bank can be achieved through singular value decomposition (SVD) or the least squares method, effectively recovering the spectral information of the speech and ensuring that the final enhanced Mel spectrum has high fidelity.

[0179] In this embodiment, the enhanced audio features are converted into an enhanced Mel spectrum through a non-autoregressive generation model, which can improve the computational efficiency of speech enhancement while ensuring speech quality. The multi-scale temporal convolutional module can extract rich temporal features, making the enhanced speech smoother. The fundamental frequency prediction analysis and harmonic compensation mechanism can optimize the pitch and timbre of the speech, improving the naturalness of the speech. The introduction of the adversarial training discriminator makes the generated speech closer to real speech, improving the stability of speech enhancement. Through the pseudo-inverse transformation of the Mel filter bank, the spectral information of the speech can be effectively recovered, making the finally generated enhanced Mel spectrum more realistic and improving the quality of subsequent speech reconstruction.

[0180] In one embodiment, the above step S70 includes:

[0181] S701, input the enhanced Mel spectrum into the waveform generator of the generative adversarial network, and generate an initial phase spectrum through the phase prediction module of the waveform generator;

[0182] S702, correct the phase error of the initial phase spectrum through the complex convolutional network of the waveform generator to generate an optimized phase spectrum;

[0183] S703, merge the optimized phase spectrum and the amplitude spectrum of the enhanced Mel spectrum into a complex spectrum;

[0184] S704, perform an inverse short-time Fourier transform on the complex spectrum to generate a time-domain waveform segment;

[0185] S705, perform an overlap-and-add synthesis process on the time-domain waveform segment to generate an initial speech waveform;

[0186] S706, add noise to the initial speech waveform to generate a noisy speech waveform;

[0187] S707, input the noisy speech waveform and the initial speech waveform into the discriminator of the generative adversarial network;

[0188] S708, analyze the distribution difference between the noisy speech and the initial speech through the multi-scale spectral consistency loss module and the adversarial loss module of the discriminator, and generate the discriminator parameter gradient;

[0189] S709, update the parameters of the discriminator based on the discriminator parameter gradient;

[0190] S710, fix the parameters of the discriminator, input the enhanced Mel spectrogram into the waveform generator, and generate a new speech waveform;

[0191] S711, calculate the adversarial loss of the new speech waveform through the discriminator to generate the generator parameter gradient;

[0192] S712, update the parameters of the waveform generator based on the generator parameter gradient;

[0193] S713, input the enhanced Mel spectrogram into the updated waveform generator to generate an optimized target speech waveform.

[0194] In this embodiment, the Mel spectrogram features need to be converted into speech waveforms to ensure the coherence and naturalness of the finally output speech in the time domain. Since traditional direct inverse transformation methods (such as the Griffin-Lim algorithm) may cause phase distortion, thus affecting the speech quality, a generative adversarial network (GAN) is adopted to optimize the speech waveform generation process. This process makes the waveform generator (Generator) and the discriminator (Discriminator) confront each other to improve the quality of speech enhancement and ensure that the generated speech waveform is close to real speech.

[0195] First, input the enhanced Mel spectrogram into the waveform generator, and generate an initial phase spectrum through the phase prediction module of the waveform generator. In traditional Mel spectrogram conversion, since the Mel spectrogram only contains amplitude information and lacks phase information, it is necessary to predict an appropriate phase spectrum. The phase prediction module estimates the corresponding phase information based on the assumption of spectral consistency by learning the time-frequency features of real speech, ensuring that the finally synthesized speech waveform is more natural.

[0196] Subsequently, correct the phase error of the initial phase spectrum through the complex convolutional network of the waveform generator to generate an optimized phase spectrum. The complex convolutional network is a neural network structure that can directly process complex data. It performs feature transformation through complex weights, effectively reducing phase distortion and improving the clarity and naturalness of speech. This process ensures that the finally generated speech waveform is more stable in terms of phase alignment, avoiding the degradation of speech quality caused by inaccurate phase information.

[0197] Next, merge the optimized phase spectrum with the amplitude spectrum of the enhanced Mel spectrogram to generate a complex spectrum. This operation makes the generated spectrum contain not only the amplitude information of the speech signal but also the optimized phase information, providing complete spectral data for subsequent waveform generation.

[0198] Then, perform the inverse short-time Fourier transform (ISTFT) on the complex spectrum to generate time-domain waveform segments. The role of the inverse short-time Fourier transform is to restore the frequency-domain information to a time-domain signal for synthesizing the final speech waveform. During this process, after the inverse transform of the frequency-domain data, there may still be certain phase errors or waveform smoothness issues, so further processing is required.

[0199] To obtain a complete speech signal, perform overlap-and-add synthesis on the time-domain waveform segments to generate an initial speech waveform. This process can smooth the transition between adjacent frames, reduce the problem of sound quality breaks caused by frame-level processing, and improve the coherence and naturalness of the speech.

[0200] During the adversarial training process, in order to optimize the stability of the waveform generator, noise addition processing needs to be introduced. Therefore, perform noise addition processing on the initial speech waveform to generate a noisy speech waveform. The purpose of noise addition processing is to simulate the noise interference in the real environment, enabling the generative adversarial network to learn more robust speech features and improve the generalization ability of the enhanced speech.

[0201] Next, input the noisy speech waveform and the initial speech waveform into the discriminator of the generative adversarial network. Through the multi-scale spectral consistency loss module and the adversarial loss module of the discriminator, analyze the distribution differences between the noisy speech and the initial speech, and generate the discriminator parameter gradients. Among them, the multi-scale spectral consistency loss module is used to calculate the speech consistency within different frequency ranges, while the adversarial loss module improves the authenticity of the generated speech through the adversarial training strategy of the GAN.

[0202] Then, update the parameters of the discriminator based on the calculated discriminator parameter gradients, enabling it to more accurately distinguish real speech from generated speech and improve the convergence effect of the adversarial training. The updated discriminator can more effectively evaluate the quality of the speech and guide the waveform generator to generate more realistic speech signals.

[0203] In the next stage of the adversarial training, fix the parameters of the discriminator, and then input the enhanced Mel spectrum into the waveform generator to generate a new speech waveform. Fixing the parameters of the discriminator means that it can effectively distinguish real speech and generated speech in the current state. The focus of the subsequent training is to optimize the performance of the waveform generator.

[0204] Then, calculate the adversarial loss for the new speech waveform through the discriminator to generate the generator parameter gradients. The purpose of calculating the adversarial loss is to measure the gap between the quality of the speech output by the current waveform generator and real speech, and optimize the parameters of the generator accordingly.

[0205] Update the parameters of the waveform generator based on the generator parameter gradients, enabling the waveform generator to gradually generate more natural and realistic speech waveforms and improve the overall quality of speech enhancement.

[0206] Finally, the enhanced Mel spectrogram is input into the updated waveform generator to generate an optimized target speech waveform. This speech waveform is more natural than the initially generated speech waveform, has higher sound quality, and can adapt to different background environments and noise interferences.

[0207] In this embodiment, the speech waveform generation process is optimized through a generative adversarial network, effectively improving the clarity and naturalness of the speech. The phase prediction and correction mechanism can reduce phase distortion, making the generated speech closer to real speech. The inverse short-time Fourier transform combined with overlap-add processing improves the coherence of the speech and reduces the artifacts that may appear during the speech splicing process. The noise addition processing and adversarial training enhance the robustness of the speech enhancement system, enabling it to maintain good speech quality in different background noise environments.

[0208] In one embodiment, a speech enhancement device based on multi-scale feature learning is provided. The speech enhancement device based on multi-scale feature learning corresponds one-to-one with the speech enhancement method based on multi-scale feature learning in the above embodiment. Refer to Figure 3 , Figure 3 FIG. is a schematic diagram of the functional modules of a preferred embodiment of the speech enhancement device based on multi-scale feature learning of the present invention. An audio preprocessing module 10, a time-frequency transformation module 20, a feature extraction module 30, an encoding processing module 40, a noise suppression module 50, a feature conversion module 60, and a speech generation module 70. The detailed descriptions of each functional module are as follows:

[0209] The audio preprocessing module 10 is configured to obtain an input audio signal and perform frame division processing on the input audio signal to generate audio frames;

[0210] The time-frequency transformation module 20 is configured to perform time-frequency conversion on the audio frames to generate Mel spectrogram features;

[0211] The feature extraction module 30 is configured to extract frequency-domain multi-scale features of the Mel spectrogram features through a multi-scale convolutional neural network;

[0212] The encoding processing module 40 is configured to encode the frequency-domain multi-scale features into a dimension-reduced latent representation;

[0213] The noise suppression module 50 is configured to perform noise suppression processing on the dimension-reduced latent representation through a deep residual network to generate enhanced audio features;

[0214] The feature conversion module 60 is configured to convert the enhanced audio features into enhanced Mel spectrograms through a non-autoregressive generative model;

[0215] The speech generation module 70 is configured to reconstruct the enhanced Mel spectrogram into a target speech waveform through a generative adversarial network.

[0216] In one embodiment, the audio preprocessing module 10 is specifically configured to:

[0217] Perform pre-emphasis filtering on the input audio signal to generate a pre-emphasized audio signal;

[0218] Divide the pre-emphasized audio signal into multiple pre-emphasized audio frames according to a preset frame length and a preset frame shift;

[0219] Perform windowing on each pre-emphasized audio frame through a Hamming window processing module to generate windowed audio frames;

[0220] Detect the short-time energy of the windowed audio frames, and filter out the silent frames in the windowed audio frames according to a preset energy threshold to generate audio frames after frame processing.

[0221] In one embodiment, the time-frequency transformation module 20 is specifically configured to:

[0222] Perform short-time Fourier transform on the audio frames to generate a linear spectrum;

[0223] Separate the magnitude spectrum and the phase spectrum of the linear spectrum, and retain the magnitude spectrum;

[0224] Map the magnitude spectrum to a Mel spectrum through a Mel filter bank;

[0225] Perform logarithmic compression on the Mel spectrum to generate a logarithmic Mel spectrum;

[0226] Perform mean-variance normalization on the logarithmic Mel spectrum to generate the Mel spectrum features.

[0227] In one embodiment, the feature extraction module 30 is specifically configured to:

[0228] Input the Mel spectrum features into a multi-scale convolutional neural network, where the multi-scale convolutional neural network includes a narrow-band convolution branch, a medium-band convolution branch, and a wide-band convolution branch;

[0229] Extract narrow-band frequency domain features of the Mel spectrum features through the narrow-band convolution branch using a convolution kernel of a first preset size;

[0230] Extract medium-band frequency domain features of the Mel spectrum features through the medium-band convolution branch using a convolution kernel of a second preset size;

[0231] Extract wide-band frequency domain features of the Mel spectrum features through the wide-band convolution branch using a convolution kernel of a third preset size;

[0232] Through the feature fusion module of the multi-scale convolutional neural network, the narrowband frequency domain features, medium-band frequency domain features, and wideband frequency domain features are concatenated in the channel dimension to generate multi-scale fusion features;

[0233] Through the cross-channel attention module of the multi-scale convolutional neural network, channel attention weight processing is performed on the multi-scale fusion features to generate multi-scale features with channel attention weighting;

[0234] Through the normalization module of the multi-scale convolutional neural network, normalization processing is performed on the multi-scale features with channel attention weighting to generate normalized multi-scale features;

[0235] Through the non-linear activation module of the multi-scale convolutional neural network, non-linear mapping is performed on the normalized multi-scale features to generate the frequency domain multi-scale features.

[0236] In one embodiment, the encoding processing module 40 is specifically configured to:

[0237] Input the frequency domain multi-scale features into the first convolutional layer of the encoder to generate first-level local frequency domain features;

[0238] Input the first-level local frequency domain features into the second convolutional layer of the encoder to generate second-level intermediate frequency domain features;

[0239] Input the second-level intermediate frequency domain features into the third convolutional layer of the encoder to generate third-level global frequency domain features;

[0240] Concatenate the first-level local frequency domain features with the third-level global frequency domain features through the cross-layer connection module to generate fused frequency domain features;

[0241] Perform normalization processing on the fused frequency domain features to generate normalized frequency domain features;

[0242] Perform non-linear transformation on the normalized frequency domain features through a non-linear activation function to generate activated frequency domain features;

[0243] Input the activated frequency domain features into a linear projection layer for dimension compression processing to generate the dimension-reduced latent representation.

[0244] In one embodiment, the feature conversion module 60 is specifically configured to:

[0245] Input the enhanced audio features into the multi-scale temporal convolution module of the non-autoregressive generation model to generate multi-scale temporal context features;

[0246] Perform time-frequency transformation on the multi-scale temporal context features to generate a time-frequency representation, and perform fundamental frequency prediction analysis and harmonic peak detection based on the time-frequency representation to generate a harmonic distribution mask;

[0247] According to the harmonic distribution mask, generate a frequency band adaptive compensation weight matrix through a gated recurrent unit;

[0248] Multiply the frequency band adaptive compensation weight matrix and the multi-scale temporal context features band by band to generate harmonic compensation temporal features;

[0249] Perform global dependency modeling on the harmonic compensation temporal features through a multi-head self-attention mechanism to generate attention-weighted temporal features;

[0250] Perform time step alignment processing on the attention-weighted temporal features through linear interpolation to generate aligned temporal features;

[0251] Input the aligned temporal features into an adversarial training discriminator, and generate parameter gradients of the non-autoregressive generation model through the multi-scale short-time Fourier transform spectrum consistency loss analysis module of the discriminator;

[0252] Update the parameters of the non-autoregressive generation model based on the parameter gradients, and input the enhanced audio features into the updated non-autoregressive generation model to generate discriminator-optimized features;

[0253] Perform power spectrum normalization processing on the discriminator-optimized features to generate normalized optimized features;

[0254] Input the normalized optimized features into a Mel spectrum mapping module, and map the normalized optimized features to the Mel band dimension through the linear projection layer of the Mel spectrum mapping module to generate Mel band features;

[0255] Perform a pseudo-inverse transformation operation on the Mel band features through the Mel filter bank of the Mel spectrum mapping module to generate the enhanced Mel spectrum.

[0256] In one embodiment, the speech generation module 70 is specifically configured to:

[0257] Input the enhanced Mel spectrum into the waveform generator of the generative adversarial network, and generate an initial phase spectrum through the phase prediction module of the waveform generator;

[0258] Perform phase error correction on the initial phase spectrum through the complex convolutional network of the waveform generator to generate an optimized phase spectrum;

[0259] Merge the optimized phase spectrum and the amplitude spectrum of the enhanced Mel spectrum into a complex spectrum;

[0260] Perform inverse short-time Fourier transform processing on the complex frequency spectra to generate time-domain waveform segments;

[0261] Perform overlapping-add synthesis processing on the time-domain waveform segments to generate an initial speech waveform;

[0262] Perform noise addition processing on the initial speech waveform to generate a noisy speech waveform;

[0263] Input the noisy speech waveform and the initial speech waveform into the discriminator of the generative adversarial network;

[0264] Analyze the distribution differences between the noisy speech and the initial speech through the multi-scale spectral consistency loss module and the adversarial loss module of the discriminator, and generate discriminator parameter gradients;

[0265] Update the parameters of the discriminator based on the discriminator parameter gradients;

[0266] Fix the parameters of the discriminator, input the enhanced Mel spectrogram into the waveform generator, and generate a new speech waveform;

[0267] Calculate the adversarial loss of the new speech waveform through the discriminator to generate generator parameter gradients;

[0268] Update the parameters of the waveform generator based on the generator parameter gradients;

[0269] Input the enhanced Mel spectrogram into the updated waveform generator to generate an optimized target speech waveform.

[0270] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as shown in Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the server side of a speech enhancement method based on multi-scale feature learning.

[0271] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as shown in Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a speech enhancement method based on multi-scale feature learning

[0272] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:

[0273] Obtain an input audio signal, and perform frame splitting processing on the input audio signal to generate audio frames;

[0274] Perform time-frequency conversion on the audio frames to generate Mel spectrogram features;

[0275] Extract the frequency-domain multi-scale features of the Mel spectrogram features through a multi-scale convolutional neural network;

[0276] Encode the frequency-domain multi-scale features into a dimension-reduced latent representation;

[0277] Perform noise suppression processing on the dimension-reduced latent representation through a deep residual network to generate enhanced audio features;

[0278] Convert the enhanced audio features into an enhanced Mel spectrogram through a non-autoregressive generation model;

[0279] Reconstruct the enhanced Mel spectrogram into a target speech waveform through a generative adversarial network.

[0280] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are realized:

[0281] Obtain an input audio signal, and perform frame splitting processing on the input audio signal to generate audio frames;

[0282] Perform time-frequency conversion on the audio frames to generate Mel spectrogram features;

[0283] Extract the frequency-domain multi-scale features of the Mel spectrogram features through a multi-scale convolutional neural network;

[0284] Encode the frequency-domain multi-scale features into a dimension-reduced latent representation;

[0285] Noise suppression processing is performed on the dimensionality-reduced latent representation through a deep residual network to generate enhanced audio features;

[0286] The enhanced audio features are converted into enhanced Mel spectrograms through a non-autoregressive generation model;

[0287] The enhanced Mel spectrograms are reconstructed into target speech waveforms through a generative adversarial network.

[0288] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0289] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database or other media used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0290] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0291] It should be noted that in the embodiments of this application, if there are software tools or components that are not of our company, they are only used for illustrative introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A speech enhancement method based on multi-scale feature learning, characterized in that: The following steps are involved: Acquire an input audio signal, and perform frame processing on the input audio signal to generate audio frames; Performing time-frequency conversion on the audio frame to generate Mel-spectrogram features; Extracting frequency domain multi-scale features of the Mel spectrum features through a multi-scale convolutional neural network; encoding the frequency domain multi-scale features into a dimensionally reduced latent representation; performing noise suppression processing on the reduced-dimensional latent representation through a deep residual network to generate enhanced audio features; Converting the enhanced audio features into an enhanced Mel spectrum through a non-autoregressive generative model; The enhanced Mel spectrum is reconstructed into a target speech waveform through a generative adversarial network.

2. The speech enhancement method based on multi-scale feature learning according to claim 1, characterized in that: Acquiring an input audio signal and performing frame processing on the input audio signal to generate audio frames, including: Performing pre-emphasis filtering on the input audio signal to generate a pre-emphasis audio signal; Dividing the pre-emphasized audio signal into a plurality of pre-emphasized audio frames according to a preset frame length and a preset frame shift; Performing windowing processing on each pre-emphasized audio frame through a Hamming window processing module to generate a windowed audio frame; The short-time energy of the windowed audio frame is detected, and the silent frames in the windowed audio frame are filtered according to a preset energy threshold to generate audio frames after frame processing.

3. The speech enhancement method based on multi-scale feature learning according to claim 1, characterized in that: Performing time-frequency conversion on the audio frame to generate Mel-spectrogram features includes: Performing short-time Fourier transform on the audio frame to generate a linear spectrum; Separating the amplitude spectrum and the phase spectrum of the linear spectrum and retaining the amplitude spectrum; Mapping the amplitude spectrum into a Mel frequency spectrum through a Mel filter bank; Performing logarithmic compression processing on the Mel spectrum to generate a logarithmic Mel spectrum; The logarithmic Mel spectrum is subjected to mean-variance normalization processing to generate the Mel spectrum feature.

4. The speech enhancement method based on multi-scale feature learning according to claim 1, characterized in that: The frequency domain multi-scale features of the Mel spectrum features are extracted by a multi-scale convolutional neural network, including: Inputting the Mel spectrum feature into a multi-scale convolutional neural network, wherein the multi-scale convolutional neural network includes a narrowband convolution branch, a mid-band convolution branch, and a wideband convolution branch; Extracting narrowband frequency domain features of the Mel spectrum features by using a convolution kernel of a first preset size through the narrowband convolution branch; Extracting mid-band frequency domain features of the Mel spectrum features by using a convolution kernel of a second preset size through the mid-band convolution branch; Extracting the broadband frequency domain features of the Mel spectrum features by using the broadband convolution branch and a convolution kernel of a third preset size; The narrowband frequency domain features, mid-band frequency domain features and broadband frequency domain features are spliced ​​in channel dimension through the feature fusion module of the multi-scale convolutional neural network to generate multi-scale fusion features; Through the cross-channel attention module of the multi-scale convolutional neural network, the multi-scale fusion features are processed by channel attention weights to generate channel attention weighted multi-scale features; The multi-scale features weighted by the channel attention are normalized by a normalization module of the multi-scale convolutional neural network to generate normalized multi-scale features; The normalized multi-scale features are nonlinearly mapped through the nonlinear activation module of the multi-scale convolutional neural network to generate the frequency domain multi-scale features.

5. The method for speech enhancement based on multi-scale feature learning according to claim 1, characterized in that: Encoding the frequency domain multi-scale features into a reduced-dimensional latent representation includes: Inputting the frequency domain multi-scale features into the first convolutional layer of the encoder to generate first-level local frequency domain features; Inputting the first-level local frequency domain features into the second convolutional layer of the encoder to generate second-level intermediate frequency domain features; Inputting the second-level intermediate frequency domain features into the third convolutional layer of the encoder to generate third-level global frequency domain features; Channel-joining the first-level local frequency domain features with the third-level global frequency domain features through a cross-layer connection module to generate fused frequency domain features; Normalizing the fused frequency domain features to generate normalized frequency domain features; Performing a nonlinear transformation on the normalized frequency domain features through a nonlinear activation function to generate an activated frequency domain feature; The activated frequency domain features are input into a linear projection layer for dimensionality compression processing to generate the reduced-dimensional potential representation.

6. The speech enhancement method based on multi-scale feature learning according to claim 1, characterized in that: The enhanced audio features are converted into enhanced Mel spectrum through a non-autoregressive generative model, including: Inputting the enhanced audio features into a multi-scale temporal convolution module of a non-autoregressive generative model to generate multi-scale temporal context features; Performing time-frequency transformation on the multi-scale time series context features to generate a time-frequency representation, and performing fundamental frequency prediction analysis and harmonic peak detection based on the time-frequency representation to generate a harmonic distribution mask; According to the harmonic distribution mask, a frequency band adaptive compensation weight matrix is ​​generated by a gated cycle unit; Multiplying the frequency band adaptive compensation weight matrix by the multi-scale temporal context feature frequency band by frequency band to generate a harmonic compensation temporal feature; The harmonic compensation time series features are subjected to global dependency modeling through a multi-head self-attention mechanism to generate attention-weighted time series features; Performing time step alignment processing on the attention weighted temporal features through linear interpolation to generate aligned temporal features; Inputting the aligned time series features into an adversarial training discriminator, and generating parameter gradients of the non-autoregressive generative model through a multi-scale short-time Fourier transform spectrum consistency loss analysis module of the discriminator; updating the parameters of the non-autoregressive generative model based on the parameter gradient, and inputting the enhanced audio features into the updated non-autoregressive generative model to generate discriminator optimization features; Performing power spectrum normalization processing on the discriminator optimization features to generate normalized optimization features; Inputting the normalized optimized features into a Mel spectrum mapping module, mapping the normalized optimized features to a Mel frequency band dimension through a linear projection layer of the Mel frequency mapping module, and generating a Mel frequency band feature; The enhanced Mel spectrum is generated by performing a pseudo inverse transform operation on the Mel frequency band feature through the Mel filter bank of the Mel spectrum mapping module.

7. The method for speech enhancement based on multi-scale feature learning according to claim 1, characterized in that: Reconstructing the enhanced Mel spectrum into a target speech waveform through a generative adversarial network, including: Inputting the enhanced Mel spectrum into a waveform generator of a generative adversarial network, and generating an initial phase spectrum through a phase prediction module of the waveform generator; Performing phase error correction on the initial phase spectrum through the complex convolution network of the waveform generator to generate an optimized phase spectrum; Combining the optimized phase spectrum and the amplitude spectrum of the enhanced Mel spectrum into a complex spectrum; Performing short-time inverse Fourier transform processing on the complex frequency spectrum to generate a time-domain waveform segment; Performing overlap-addition synthesis processing on the time domain waveform segments to generate an initial speech waveform; Performing noise processing on the initial speech waveform to generate a noisy speech waveform; Inputting the noisy speech waveform and the initial speech waveform into a discriminator of a generative adversarial network; Analyzing the distribution difference between the noisy speech and the initial speech through the multi-scale spectral consistency loss module and the adversarial loss module of the discriminator, and generating the discriminator parameter gradient; Update the parameters of the discriminator based on the discriminator parameter gradient; Fixing the parameters of the discriminator, inputting the enhanced Mel spectrum into the waveform generator, and generating a new speech waveform; Performing adversarial loss calculation on the new speech waveform through the discriminator to generate a generator parameter gradient; updating parameters of the waveform generator based on the generator parameter gradient; The enhanced Mel spectrum is input into the updated waveform generator to generate an optimized target speech waveform.

8. A speech enhancement device based on multi-scale feature learning, characterized in that: The speech enhancement device based on multi-scale feature learning comprises: An audio preprocessing module, used to obtain an input audio signal and perform frame processing on the input audio signal to generate audio frames; A time-frequency conversion module, used for performing time-frequency conversion on the audio frame to generate Mel spectrum features; A feature extraction module, used for extracting frequency domain multi-scale features of the Mel spectrum features through a multi-scale convolutional neural network; An encoding processing module, used for encoding the frequency domain multi-scale features into a reduced-dimensional potential representation; A noise suppression module, configured to perform noise suppression processing on the reduced-dimensional potential representation through a deep residual network to generate enhanced audio features; A feature conversion module, used for converting the enhanced audio features into an enhanced Mel spectrum through a non-autoregressive generative model; The speech generation module is used to reconstruct the enhanced Mel spectrum into a target speech waveform through a generative adversarial network.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a speech enhancement program based on multi-scale feature learning stored in the memory and executable on the processor. When the speech enhancement program based on multi-scale feature learning is executed by the processor, the steps of the speech enhancement method based on multi-scale feature learning as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores a speech enhancement program based on multi-scale feature learning, and when the speech enhancement program based on multi-scale feature learning is executed by the processor, the steps of the speech enhancement method based on multi-scale feature learning as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Virtual platform-based lung cancer patient psychological health assessment intervention system

    CN120823969A

  • Speech enhancement method and device based on full-band and sub-band fusion, terminal and medium

    CN121054024A

  • Game sound effect data restoration method based on artificial intelligence

    CN121148403A

  • Intelligent water leakage detection method and system for water supply network

    CN121188581A

  • Rotary machinery air domain acoustic diagnosis method based on generative adversarial noise reduction

    CN121237122A