Speech enhancement method and system based on metricGAN model, and electronic equipment

By introducing an adversarial learning mechanism that combines adaptive spectral domain noise suppression and speech quality metrics into the metricGAN model, the problem of insufficient adaptability to noise environments in existing technologies is solved, achieving efficient speech enhancement under complex noise conditions and improving speech clarity and intelligibility.

CN121789706APending Publication Date: 2026-04-03MACAO POLYTECHNIC INST
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing speech enhancement technologies are not adaptable to noise changes in complex and non-stationary noise environments, resulting in residual noise or speech distortion in the enhanced speech. Furthermore, the collaborative utilization of different types of speech enhancement methods is not mature enough, limiting the overall performance improvement.

Method used

We employ a speech enhancement method based on the metricGAN model. By using an adaptive spectral domain noise suppression prior and an adversarial learning mechanism based on speech quality metrics, we construct generator and discriminator models, conduct joint training, optimize and update model parameters, and achieve end-to-end speech enhancement.

Benefits of technology

While reducing background noise, it effectively reduces music noise and speech distortion, improving the clarity and intelligibility of enhanced speech, and is especially suitable for speech enhancement scenarios under non-stationary noise and low signal-to-noise ratio conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789706A_ABST
    Figure CN121789706A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image reconstruction, and provides a speech enhancement method and system based on a metricGAN model, and electronic equipment. The method comprises the following steps: acquiring a to-be-processed noisy voice signal, and preprocessing the noisy voice signal to obtain a time-frequency feature; performing adaptive spectral domain noise suppression processing based on the time-frequency features to obtain initial enhanced speech features; constructing a metricGAN model comprising a generator and a discriminator; inputting the initial enhanced speech features into a generator to obtain second enhanced speech features, inputting the second enhanced speech features into a discriminator, and performing function fitting of speech quality evaluation indexes on the second enhanced speech features through the discriminator to obtain corresponding speech quality scores; performing joint training on the generator and the discriminator by adopting an adversarial learning mode based on the voice quality score; and processing the noisy voice signal by using the trained generator, and performing reconstruction according to the generated third enhanced voice feature to obtain an enhanced voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image reconstruction, and more specifically, to a speech enhancement method and system, and electronic device based on the metricGAN model. Background Technology

[0002] With the rapid development of voice interaction technology, intelligent voice communication, and voice recognition systems, the intelligibility and subjective quality of voice signals in complex noisy environments have become important factors restricting the performance of related applications. In practical application scenarios, voice signals are often inevitably affected by various factors such as environmental noise, equipment noise, and transmission interference, leading to a decline in voice quality, which in turn affects the accuracy of voice recognition, auditory experience, and the overall effect of subsequent voice processing tasks.

[0003] To improve the quality of noisy speech, various speech enhancement methods have been proposed in existing technologies. For example, speech enhancement algorithms based on traditional signal processing (such as spectral subtraction and Wiener filtering) are widely used due to their relatively simple implementation and low computational complexity. However, these methods usually rely on assumptions about the statistical characteristics of noise, and in application scenarios with low signal-to-noise ratios or large noise variations, they are prone to introducing distortion phenomena such as musical noise, resulting in limited improvement in speech quality. In addition, with the development of deep learning technology, neural network-based speech enhancement methods have gradually attracted attention. Related research attempts to use generative adversarial networks to model the mapping relationship between noisy speech and clean speech, and to improve the subjective quality of enhanced speech by introducing speech quality evaluation indicators as optimization criteria. In practical applications, existing speech enhancement technologies generally suffer from insufficient adaptability to noisy environments and inconsistent processing effects under different time or spectral characteristics, and distortion or residual noise may still exist in the enhanced speech. Furthermore, the synergistic utilization of different types of speech enhancement methods is still immature, which limits the overall performance improvement. Therefore, it is urgent to propose a technical solution to address at least one of the technical problems in the existing technologies. Summary of the Invention

[0004] In this context, the embodiments of this application aim to provide a speech enhancement method, system, and electronic device based on the metricGAN model. By employing an adaptive spectral domain noise suppression prior and an adversarial learning mechanism based on speech quality indicators in the metricGAN speech enhancement framework, the technical problems of insufficient adaptability to noise changes in complex and non-stationary noise environments, and residual noise or speech distortion in enhanced speech are solved in the prior art.

[0005] In a first aspect of this application, a speech enhancement method based on a metricGAN model is provided. The method includes: acquiring a noisy speech signal to be processed, and preprocessing the noisy speech signal, the preprocessing including at least time-frequency analysis to obtain time-frequency features characterizing the speech energy distribution; performing adaptive spectral domain noise suppression processing based on the time-frequency features to obtain initial enhanced speech features for subsequent modeling, wherein the adaptive spectral domain noise suppression processing is used to introduce noise suppression priors and reduce background noise interference, the adaptive spectral domain noise suppression processing including noise estimation and spectral subtraction operations on different frequency sub-bands respectively, and adaptively adjusting the noise suppression intensity according to the signal-to-noise ratio of each frequency sub-band; constructing a method including... The generator and discriminator are configured using a metricGAN model. The initial enhanced speech features are input into the generator to obtain second enhanced speech features, which are then input into the discriminator. The discriminator fits a speech quality evaluation metric function to the second enhanced speech features to obtain a corresponding speech quality score. Based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning. This joint training includes at least quality constraints on the second enhanced speech features at different time or spectral scales, and optimizing and updating the model parameters of the metricGAN model. The trained generator is then used to process the noisy speech signal, and the enhanced speech signal is reconstructed based on the generated third enhanced speech features.

[0006] In a second aspect of the embodiments of this application, a speech enhancement system based on a metricGAN model is provided. The system includes the following modules: a preprocessing module, used to acquire a noisy speech signal to be processed and preprocess the noisy speech signal, the preprocessing including at least time-frequency analysis to obtain time-frequency features characterizing the speech energy distribution; an enhancement module, used to perform adaptive spectral domain noise suppression processing based on the time-frequency features to obtain initial enhanced speech features for subsequent modeling, wherein the adaptive spectral domain noise suppression processing is used to introduce noise suppression priors and reduce background noise interference, the adaptive spectral domain noise suppression processing including noise estimation and spectral subtraction operations on different frequency sub-bands respectively, and adaptively adjusting the noise suppression intensity according to the signal-to-noise ratio of each frequency sub-band; and a construction module, using... The system constructs a metricGAN model including a generator and a discriminator; a training module is used to input the initial enhanced speech features into the generator to obtain second enhanced speech features, and input the second enhanced speech features into the discriminator. The discriminator fits a function of speech quality evaluation index to the second enhanced speech features to obtain the corresponding speech quality score; based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning, wherein the joint training includes at least quality constraints on the second enhanced speech features at different time scales or spectral scales, and optimizing and updating the model parameters of the metricGAN model; a reconstruction module is used to process the noisy speech signal using the trained generator, and reconstruct the enhanced speech signal based on the generated third enhanced speech features.

[0007] In a third aspect of the embodiments of this application, an electronic device is provided for implementing the speech enhancement method based on the metricGAN model described in the first aspect.

[0008] This application discloses a speech enhancement method, system, and electronic device based on the metricGAN model. The embodiment employs an adaptive spectral domain noise suppression prior and an adversarial learning mechanism based on speech quality metrics within the metricGAN speech enhancement framework. This effectively reduces background noise while minimizing musical noise and speech distortion, thereby improving the clarity, intelligibility, and overall perceived quality of the enhanced speech. It is particularly suitable for speech enhancement scenarios under non-stationary noise and low signal-to-noise ratio conditions. Attached Figure Description

[0009] Figure 1 This application illustrates a flowchart of a speech enhancement method based on the metricGAN model. Figure 2 This is a schematic diagram of the structure of a speech enhancement system based on the metricGAN model shown in this application. Detailed Implementation

[0010] To improve the quality of noisy speech, various speech enhancement methods have been proposed in the existing technology. Among them, speech enhancement algorithms based on traditional signal processing (such as spectral subtraction, Wiener filtering, etc.) are widely used due to their simplicity and low computational complexity. However, these methods usually rely on assumptions about the statistical characteristics of noise, and are prone to distortion problems such as musical noise in low signal-to-noise ratio or non-stationary noise environments, making it difficult to achieve an ideal balance between speech quality and noise suppression.

[0011] With the development of deep learning technology, speech enhancement methods based on neural networks have gradually become a research hotspot. Among them, Generative Adversarial Networks (GANs) have been introduced into the field of speech enhancement due to their advantages in modeling complex data distributions, and are used to learn the nonlinear mapping relationship from noisy speech to clean speech. However, existing technologies have difficulty effectively combining traditional speech enhancement methods with index-driven GANs, failing to fully leverage the prior advantages of traditional spectral domain processing methods in noise suppression, and failing to further improve the overall performance of speech enhancement through multi-scale, multi-objective optimization mechanisms. Therefore, how to introduce a generative adversarial learning mechanism based on speech quality indices while preserving the stability of traditional spectral subtraction methods, and improve the model's adaptability to complex noisy environments through multi-scale optimization strategies, remains a pressing technical problem to be solved in this field.

[0012] This application provides a speech enhancement method, system, and electronic device based on the metricGAN model. Compared to existing technologies, this application's implementation achieves the following technical effects by introducing an adaptive spectral domain noise suppression prior and an adversarial learning mechanism based on speech quality indicators into the metricGAN speech enhancement framework: First, by performing time-frequency analysis on the noisy speech signal and introducing adaptive spectral domain noise suppression processing, background noise is initially weakened before deep model modeling, making the input features closer to real clean speech in terms of energy distribution and signal-to-noise structure. On the one hand, noise estimation and suppression are performed separately for different frequency sub-bands, effectively avoiding speech distortion caused by uniform processing across the entire frequency range. On the other hand, the noise suppression intensity is adaptively adjusted according to the signal-to-noise ratio of each frequency sub-band, ensuring that speech details in high signal-to-noise regions are fully preserved and noise components in low signal-to-noise regions are significantly weakened, thereby introducing a stable and reliable noise suppression prior for subsequent model training. Second, by constructing a metricGAN model containing a generator and a discriminator, and having the discriminator fit a function of speech quality evaluation indicators to the enhanced speech features, the model optimization objective is transformed from the traditional signal reconstruction error to an indicator hypothesis directly oriented towards perceptual quality. During training, the generator no longer solely pursues numerical approximation of amplitude or spectrum, but is guided by speech quality scores fed back by the discriminator, thereby improving the overall consistency of enhanced speech in both subjective listening experience and objective evaluation metrics. Thirdly, quality constraints at different time or spectral scales are introduced during joint training, enabling the model to simultaneously consider short-term speech structure, long-term speech coherence, and the perceptual importance of different frequency bands. This multi-scale quality constraint mechanism effectively alleviates the local overfitting problem easily caused by single-scale optimization, improving the model's generalization ability and stability in complex noisy environments and various speech scenarios. Finally, the trained generator is used for end-to-end processing of noisy speech signals, achieving a synergistic enhancement effect from noise suppression prior guidance to perceptual quality-driven optimization.

[0013] This application's implementation method integrates traditional spectral domain noise suppression with the metricGAN model. This implementation method can effectively introduce noise suppression priors before deep model training, reduce the interference of background noise on the model learning process, and through adversarial learning driven by speech quality indicators, enable the generator to directly optimize perceptual quality. This improves the stability and consistency of the speech enhancement scheme under different noise conditions, different time scales, and spectral characteristics, and effectively improves the clarity, intelligibility, and overall subjective quality of the enhanced speech.

[0014] Figure 1 The flowchart of a speech enhancement method based on a metricGAN model provided in one embodiment of this application includes: Step S101: Obtain the noisy speech signal to be processed, and preprocess the noisy speech signal to obtain time-frequency features characterizing the speech energy distribution. Step S102: Based on the time-frequency features, perform adaptive spectral domain noise suppression processing to obtain initial enhanced speech features for subsequent modeling. The adaptive spectral domain noise suppression processing is used to introduce noise suppression priors and reduce background noise interference. The adaptive spectral domain noise suppression processing includes performing noise estimation and spectral subtraction operations on different frequency sub-bands respectively, and adaptively adjusting the noise suppression intensity according to the signal-to-noise ratio of each frequency sub-band. Step S103: Construct a metricGAN model including a generator and a discriminator; Step S104: Input the initial enhanced speech features into the generator to obtain the second enhanced speech features, and input the second enhanced speech features into the discriminator. The discriminator performs a function fitting of the speech quality evaluation index on the second enhanced speech features to obtain the corresponding speech quality score. Step S105: Based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning. The joint training includes at least quality constraints on the second enhanced speech features at different time scales or spectral scales, and optimizing and updating the model parameters of the metricGAN model. Step S106: The trained generator is used to process the noisy speech signal, and the enhanced speech signal is reconstructed based on the generated third enhanced speech features.

[0015] In this embodiment, the noisy speech signal to be processed can be understood as a speech signal collected or received in a real-world application scenario that is subject to one or more types of noise interference. The noise can originate from external environmental noise, equipment background noise, channel transmission noise, or a combination thereof, such as background human voices, traffic noise, mechanical operating noise, current noise, or random interference during communication. The noisy speech signal can be the original speech signal collected in real time through a microphone, speech acquisition terminal, or smart device, or historical speech data read from a storage medium, or speech data received through a communication network. In both the time and spectral domains, the noisy speech signal typically exhibits a superposition of speech and noise components, leading to disordered speech energy distribution and masking of speech features, thereby reducing speech clarity and intelligibility. This embodiment of the application achieves the highlighting of effective speech information and the suppression of noise interference by performing subsequent preprocessing, adaptive spectral domain noise suppression, and enhancement processing based on the metricGAN model on the aforementioned noisy speech signal.

[0016] In this embodiment, the speech quality evaluation index may include objective speech quality evaluation indices such as PESQ, which are introduced as metrics into the metricGAN model. Specifically, the discriminator is no longer used to distinguish between real and generated speech, but is configured to fit the functional relationship of the speech quality evaluation index to simulate the scoring process of the corresponding speech quality index. The discriminator receives the enhanced speech features output by the generator as input and outputs the quality score result corresponding to the speech quality evaluation index, thereby achieving an approximate evaluation of the perceived quality of speech.

[0017] During training, the generator uses the speech quality score output by the discriminator as optimization feedback. It generates enhanced speech features and feeds them into the discriminator to obtain the corresponding expected score. When the expected score gradually approaches the preset target score, the generator's output is considered to have achieved the desired effect in terms of speech quality metrics. In this way, the metricGAN model can maintain end-to-end trainability while directly aligning the generator's optimization objective with speech quality evaluation metrics, thereby guiding the generator to output enhanced speech with high quality in both objective evaluation metrics and subjective listening experience.

[0018] MetricGAN is a generative adversarial network framework optimized for evaluation metrics. It explicitly introduces speech quality metrics (such as PESQ and POLQA) into the adversarial structure, training the discriminator to become a learnable speech quality scorer. This score then directly guides the generator's optimization. In other words, while traditional GAN ​​discriminators primarily distinguish between real and fake samples, metricGAN's discriminator learns to map input speech features to specific quality scores. The generator's goal is not merely to confuse the discriminator with real or fake samples, but to make the discriminator's quality score as close as possible to the true metric, thus iteratively improving speech quality throughout the training process.

[0019] Structurally, the metricGAN model still consists of a generator and a discriminator. The generator typically takes the time-frequency features of noisy speech or initial enhanced speech as input, such as amplitude spectrum, sub-band energy spectrum, or receptive domain features, and outputs the enhanced speech features or reconstructed speech signal to achieve denoising, dereverberation, or overall quality improvement. The discriminator receives the generator's output and a reference clean speech (in the case of a reference), or receives the enhanced speech alone (in the case of no reference). It automatically extracts multi-level speech quality cues through a deep neural network, including residual noise distribution, spectral smoothness, speech harmonic structure integrity, and temporal continuity, and finally outputs a scalar score, which is designed to approximate the real speech quality evaluation index. By training on a large number of speech samples and their real quality scores, the discriminator gradually learns to approximate the calculation process of the evaluation index using an internal nonlinear mapping, thus becoming a differentiable quality estimation model that can be used for backpropagation.

[0020] In terms of training mechanism, metricGAN employs a combination of adversarial learning and metric fitting. Compared to traditional augmentation networks that rely solely on simple losses such as mean squared error and spectral distance, metricGAN softly embeds complex speech quality metrics into the neural network training objective through a discriminator. This allows the model to automatically balance multiple dimensions such as noise suppression, speech fidelity, and naturalness during the learning process, ultimately generating augmented speech that is closer to human auditory expectations in both objective metrics and subjective listening experience. In scenarios such as speech enhancement and speech restoration, the metricGAN model is an adversarial learning structure that performs end-to-end optimization centered on quality metrics.

[0021] As an optional embodiment, in step S101, the preprocessing of the noisy speech signal to obtain time-frequency features characterizing the speech energy distribution includes: performing frame segmentation and windowing processing on the noisy speech signal; performing short-time Fourier transform on the windowed speech frames to obtain the time spectrum of the speech signal; and performing frequency-weighted processing on the time spectrum in conjunction with auditory perception sensitivity to obtain the time-frequency features. The gain coefficients of the low-frequency and high-frequency ranges are adaptively adjusted through training data or psychoacoustic models so that the time-frequency spectrum features of the processed speech signal are more consistent with the auditory sensitivity characteristics of the human ear in terms of frequency distribution, thereby providing a basis for subsequent multi-band noise suppression.

[0022] In step S101, the noisy speech signal is first subjected to framing and windowing processing. Since the speech signal can be approximated as a stationary signal within a short time range, the continuous time-domain speech signal is segmented according to a preset frame length and frame shift, for example, dividing the speech signal into several consecutive short-time speech frames. Each frame typically corresponds to a time length of tens of milliseconds, making the speech characteristics relatively stable within that time range. After framing, windowing processing is applied to each frame of speech signal to reduce the impact of discontinuities at frame boundaries on subsequent spectral analysis. For example, by smoothing the transition of speech samples within each frame, the energy at the beginning and end of the frame gradually decreases, thereby reducing spectral leakage.

[0023] Secondly, a short-time Fourier transform is performed on the windowed speech frames to obtain the time-frequency spectrum of the speech signal. By performing frequency domain analysis on each frame of the speech signal, the speech signal, which originally varies with time, can be converted into a representation that includes both time and frequency information, thus intuitively reflecting the energy distribution of each frequency component at different time points. For example, for speech frames containing clear vowels, their time-frequency spectrum usually shows a strong energy concentration in the low-frequency region, while speech frames containing fricatives or background noise may show a dispersed or irregular energy distribution in the mid-to-high frequency region.

[0024] Understandably, performing a short-time Fourier transform on each windowed speech frame can be understood as decomposing the frequency components of each short-time speech signal to analyze the energy distribution of the speech at different frequencies within that time segment. Specifically, frequency domain analysis is performed on each frame sequentially according to the time order of the speech frames, and the analysis results are correlated with the time position of that frame, thus forming a spectrum sequence arranged continuously along the time axis. Through this processing, the speech signal, which originally only varied along the time axis, is transformed into a two-dimensional representation, where one dimension corresponds to the time frame sequence and the other dimension corresponds to the frequency interval. The value at each time-frequency position reflects the energy strength of that frequency component within the corresponding time period. For example, in actual speech, when a speaker produces stable vowel phonemes, the frequency structure within that speech frame is relatively stable, and its time spectrum often exhibits a continuous and concentrated energy distribution in the low-frequency range and its adjacent frequency bands; while when fricatives or transient consonants appear in the speech, the corresponding time spectrum may show a more dispersed energy distribution in the mid-to-high frequency range, and the changes over time are more drastic. In scenarios with background noise, the time-spectrum obtained from the short-time Fourier transform can intuitively reflect the differences between noise and speech. For example, stable environmental noise often exhibits a relatively uniform or continuous spectral distribution across multiple time frames, while speech components typically show significant energy spikes at specific time locations and frequency ranges. The time-spectrum obtained in this way not only preserves the temporal structure information of the speech but also clearly characterizes the differences in frequency distribution between speech and noise. Therefore, in the above embodiment, performing a short-time Fourier transform on the windowed speech frame can provide intuitive, fine-grained, and discriminative feature representations for subsequent multi-band noise estimation, spectral domain suppression, and deep modeling based on the metricGAN model, helping to improve the accuracy and stability of speech enhancement processing in complex noise environments.

[0025] Next, combining auditory perception sensitivity, frequency-weighted processing is performed on the time-frequency spectrum to obtain time-frequency characteristics that better match the characteristics of human hearing. In this process, the system adjusts the energy of each frequency range in the time-frequency spectrum according to the sensitivity of the human ear to different frequency ranges. Specifically, in frequency ranges where the human ear is more sensitive, the corresponding spectral components are appropriately enhanced. In frequency ranges where perception is relatively less sensitive, the spectral components are moderately attenuated. For example, spectral components carrying the main speech pitch information in the low-frequency range can be given higher gain to highlight the main structure of the speech. For spectral components in the high-frequency range that are mainly noise components, their weight can be reduced to weaken the impact of noise on subsequent processing.

[0026] Furthermore, the gain coefficients in the low-frequency and high-frequency ranges can be adaptively adjusted using training data or psychoacoustic models. In one implementation, the system can dynamically adjust the weighting of each frequency band by statistically analyzing the contribution of different frequency ranges to speech intelligibility and subjective quality based on a large number of speech samples. In another implementation, descriptions of auditory thresholds and frequency sensitivity from psychoacoustic models can be introduced, enabling the frequency weighting strategy to adaptively update with changes in speech content and noise environment. For example, in low signal-to-noise ratio environments, the weighting intensity of key speech bands can be appropriately increased to enhance speech intelligibility.

[0027] Through the above preprocessing steps, the obtained time-frequency features can not only accurately reflect the energy distribution of the speech signal in the time and frequency dimensions, but also better conform to the characteristics of human auditory perception in terms of frequency weight allocation. This provides a stable and perceptually meaningful feature foundation for subsequent multi-band adaptive noise suppression processing and metricGAN model modeling.

[0028] As an optional embodiment, after completing the inverse short-time Fourier transform and obtaining the time-domain enhanced speech signal, a set of coordinated time-domain post-processing procedures is introduced to further optimize the enhanced speech, forming a closed-loop enhancement mechanism that combines spectral domain optimization and time-domain optimization. Specifically, frame-level energy smoothing processing is first performed on the time-domain enhanced speech signal. Specifically, the time-domain enhanced speech signal is re-framed according to the frame structure corresponding to the aforementioned analysis stage, and the energy change trend of each frame is smoothed to make the energy transition between adjacent speech frames more continuous. Through this processing, transient amplitude jitter and abrupt changes introduced during spectral subtraction or phase optimization can be effectively suppressed, avoiding unnatural breaks or jumps in the enhanced speech, thereby improving the overall coherence and naturalness of the speech.

[0029] Furthermore, based on the energy-smoothed time-domain signal, a nonlinear gain adjustment strategy based on short-time envelope is adopted. Specifically, the short-time amplitude envelope of the time-domain enhanced speech signal is extracted, and adaptive nonlinear gain control is applied to the speech frames according to the envelope energy level. For speech segments with low energy but determined to contain effective speech information, a slight amplification gain is applied to enhance weak speech components and improve intelligibility; for speech segments with high energy, their gain is limited or compressed to prevent excessive peak amplitude from causing clipping distortion or auditory discomfort. This nonlinear gain adjustment can maintain a more balanced auditory performance of the enhanced speech in different loudness ranges without significantly amplifying background noise.

[0030] Next, the time-domain speech signal with gain adjustment is subjected to superposition reconstruction and de-overlap processing. Specifically, the speech signals of each frame are superimposed in the time domain according to a preset frame shift relationship, and a smooth transition is performed in the overlapping area to reconstruct a continuous time-domain enhanced speech signal. This step can eliminate the boundary effects caused by frame-level processing and ensure the continuity and integrity of the final output speech in the time dimension.

[0031] As an optional implementation, short-time speech activity detection (VAD) filtering can be further applied to the final enhanced speech signal. By distinguishing between speech segments and non-speech segments in the enhanced speech, residual noise components are further suppressed or attenuated within the time intervals determined to be non-speech activity, thereby reducing the impact of residual noise in silent or paused segments on the listening experience. This step further improves the cleanliness of the enhanced speech without compromising the effective speech content.

[0032] By integrating the aforementioned temporal smoothing, nonlinear gain adjustment based on short-time envelope, and optional VAD-assisted filtering into the speech reconstruction stage, this embodiment introduces a temporal-level adaptive adjustment mechanism based on the spectral domain optimization results. This allows spectral domain enhancement and temporal optimization to complement and synergize with each other, thereby effectively improving the naturalness, stability, and intelligibility of the enhanced speech. This demonstrates a significant creative effect compared to existing technologies.

[0033] In one specific embodiment of this application, the acquisition and application of human auditory sensitivity is not based on fixed parameters, but rather on comprehensive modeling and adaptive adjustment from multiple sources, so that the speech enhancement processing results are more in line with the auditory perception characteristics of different users and different usage scenarios.

[0034] As an optional implementation, in the above steps, firstly, the human ear's auditory sensitivity can be obtained through psychoacoustic models, experimental measurements, or training data. In one implementation, the system can incorporate an average hearing curve based on numerous auditory experiments to characterize the human ear's sensitivity to sound intensity across different frequency ranges. For example, the human ear is generally more sensitive to speech components in the low to mid-frequency range, while its perception of excessively low or high frequencies is relatively weak. In addition to the general average hearing curve, hearing threshold information for different age groups can be introduced. For instance, considering the significant hearing loss in the high-frequency region among the elderly, the perception weight of the high-frequency range can be appropriately increased. Furthermore, for specific language environments, such as language types where consonant intelligibility is crucial, a more language-specific perception sensitivity model can be formed by analyzing the contribution of different frequency bands to speech intelligibility in the training data.

[0035] Secondly, the gain coefficients in the low-frequency and high-frequency ranges are adaptively adjusted for different user groups or usage scenarios. In practice, the corresponding frequency weighting strategy can be dynamically selected or adjusted based on user profile information or usage scenario information. For example, when targeting children or adults with normal hearing, a more balanced frequency gain distribution can be used to maintain the naturalness of the speech. When targeting elderly users, the gain in the high-frequency range can be appropriately increased to compensate for their decreased ability to perceive high-frequency speech components. Regarding usage scenarios, when the user is detected in a quiet indoor environment, the overall frequency gain adjustment amplitude can be reduced to maintain the original timbre of the speech. In noisy outdoor environments, the gain weight of key audio segments can be increased to enhance the prominence of the speech in noise.

[0036] Furthermore, the adaptive adjustment of the gain coefficient can be dynamically updated based on historical speech sample statistics, online feedback, or user-defined preferences. In one implementation, speech samples generated during long-term user use can be statistically analyzed to assess the impact of enhancement on speech clarity and intelligibility in different frequency bands, thereby gradually adjusting the frequency gain parameters to better suit individual user auditory preferences. In another implementation, the system can also incorporate online feedback from users during actual use, such as subjective evaluations of speech clarity or adjustments to volume and timbre, to correct the frequency weighting strategy in real time. Additionally, users can actively select preferred modes through the settings interface, such as clarity priority, naturalness priority, or high-frequency enhancement modes, to configure and update the gain coefficients in the low-frequency and high-frequency ranges accordingly.

[0037] Through the aforementioned multi-source sensing information acquisition and adaptive adjustment mechanism, this embodiment enables frequency weighting processing to no longer rely on fixed rules, but to be dynamically optimized according to differences in the population, changes in the scene, and user preferences, thereby further improving the speech enhancement results in terms of subjective listening experience, intelligibility, and adaptability.

[0038] As an optional embodiment, step S102 involves performing noise estimation and spectral subtraction operations on different frequency sub-bands, and adaptively adjusting the noise suppression intensity based on the signal-to-noise ratio (SNR) of each frequency sub-band. This includes: dividing the time spectrum into multiple continuous and non-overlapping frequency sub-bands according to a preset frequency perception scale; within each frequency sub-band, independently estimating the background noise spectrum corresponding to each frequency sub-band based on statistical information from non-speech intervals or adjacent multi-frames to obtain the speech spectrum and noise spectrum within each frequency sub-band; calculating the posterior SNR of each frequency sub-band based on the speech spectrum and noise spectrum within each frequency sub-band; the posterior SNR being used to characterize the instantaneous ratio of speech signal to noise signal within each frequency sub-band; adaptively determining the spectral subtraction intensity parameter corresponding to each frequency sub-band based on the posterior SNR; wherein the spectral subtraction intensity parameter changes with the posterior SNR in a non-linear mapping manner; performing spectral subtraction operations on the speech spectrum within each frequency sub-band using the spectral subtraction intensity parameter, and fusing the spectral subtraction operation results within each frequency sub-band into the initial enhanced speech feature.

[0039] Specifically, in step S102, noise estimation and spectral subtraction are performed on different frequency sub-bands, and the noise suppression intensity is adaptively adjusted according to the signal-to-noise ratio of each frequency sub-band. Simultaneously, a sub-band weighted fusion and residual suppression mechanism is introduced. First, the time-spectrum of the speech signal is divided according to a preset frequency perception scale, dividing the complete spectrum into multiple continuous and non-overlapping frequency sub-bands. This frequency perception scale can be set based on the characteristics of human hearing or the distribution of the speech spectrum. For example, the low-frequency region can be divided into narrower sub-bands to finely characterize the speech pitch and formant information, while the mid-to-high frequency region can be divided into relatively wider sub-bands to cover fricative sounds and noise components. Through this method, the spectral characteristics within each frequency sub-band are relatively consistent, facilitating subsequent independent processing.

[0040] For example, to simulate the nonlinear perception of sound frequencies by the human ear, instead of directly processing all frequency points uniformly, a sub-band division operation can be performed. Based on preset frequency perception rules (such as Mel scale or Barker scale logic), multiple cutoff frequency points from low to high frequencies are determined. In the low-frequency region, because the human ear is sensitive to frequency changes and the main energy of speech (fundamental tone, formants) is concentrated here, the frequency intervals are set to be small, i.e., containing fewer frequency points, thus forming dense, narrow sub-bands to achieve high frequency resolution. In the high-frequency region, which mainly contains fricative sounds and ambient noise, the human ear's resolution decreases, so the frequency intervals are set to be larger, i.e., containing more frequency points, forming wide sub-bands. Finally, the original linear spectrum is reorganized into a set of sub-band vectors, each vector representing the energy distribution within a specific frequency range, and all subsequent processing is performed using these sub-bands as independent units.

[0041] Subsequently, noise estimation is performed separately within each frequency sub-band. In practice, the background noise spectrum within each frequency sub-band can be independently modeled using detected non-speech intervals or statistical information from multiple adjacent frames. For example, in speech pauses or weak speech segments, the main energy within that sub-band can be considered to originate from noise, thus updating the noise level of that frequency sub-band. In continuous speech frames, historical statistical information can be used to smooth and correct the noise estimation results. In this way, the corresponding speech and noise spectrum descriptions within each frequency sub-band can be obtained.

[0042] For example, within each predefined frequency sub-band, an independent background noise model is maintained. First, the noise floor of each sub-band is initialized using the silence segments at the beginning of the speech. Then, when processing each frame of speech, the energy characteristics of the current sub-band are calculated. The energy of the current frame's sub-band is compared with the historical minimum energy statistical value. If the current sub-band energy is close to the historical minimum, it is determined that the frequency band is mainly dominated by noise at that moment. A smooth update strategy is then adopted, incorporating a certain proportion of the current energy into the noise model to update the noise spectrum estimate. If the current sub-band energy is higher than the historical noise level, it is determined that speech activity exists at that moment. In this case, the noise model is stopped or updated only at a very low rate to prevent misclassifying speech components as noise, thereby ensuring the purity of the noise estimate.

[0043] Based on this, the posterior signal-to-noise ratio (SNR) of each frequency sub-band is evaluated according to the relationship between the speech spectrum and the noise spectrum. This SNR characterizes the instantaneous proportion of speech components relative to noise components within that frequency sub-band. For example, when the speech energy is significantly higher than the noise energy in a low-frequency sub-band, its posterior SNR is relatively high. Conversely, in high-frequency sub-bands dominated by noise, the posterior SNR is relatively low. This posterior SNR reflects the differences in the distribution of speech and noise across different frequency sub-bands.

[0044] For example, the posterior signal-to-noise ratio (SNR) of each subband is calculated using the observed spectrum (noisy speech) of the current frame and the real-time updated noise estimation spectrum. For each subband, the ratio of the currently observed spectral amplitude (or energy) to the estimated noise spectral amplitude (or energy) is calculated. This ratio directly reflects the signal strength relative to the noise signal strength within the current subband. A larger value indicates a purer speech component in that frequency band; a value closer to 1 (or 0 dB) indicates that the frequency band is almost completely submerged in noise. To avoid drastic fluctuations in SNR caused by instantaneous noise abrupt changes, a recursive smoothing mechanism is typically introduced. This mechanism combines the prior SNR result from the previous frame to correct the current posterior SNR, resulting in a more stable and smoother SNR index to guide subsequent suppression operations.

[0045] Next, based on the posterior signal-to-noise ratio (SNR), the spectral reduction intensity parameter corresponding to each frequency sub-band is adaptively determined. Specifically, the spectral reduction intensity parameter is not a fixed value, but is dynamically adjusted according to the change of the sub-band's posterior SNR through a nonlinear mapping method. For example, for frequency sub-bands with low posterior SNR and a large proportion of noise, the system will increase the spectral reduction intensity to more actively suppress noise components; while for sub-bands with high posterior SNR and dominant speech components, the spectral reduction intensity will be reduced to avoid excessive attenuation of effective speech information. Through this nonlinear adjustment method, differentiated and refined noise suppression control can be achieved between different frequency sub-bands.

[0046] For example, when a subband has an extremely low signal-to-noise ratio (high noise), the region is determined to be mainly interference, thus generating a large spectral subtraction intensity parameter. This means that in subsequent subtraction operations, not only the estimated noise amount must be subtracted, but an additional portion must be subtracted to suppress residual noise to the greatest extent. When a subband has a high signal-to-noise ratio (speech-dominated), a small spectral subtraction intensity parameter is generated, even approaching zero. This means performing as few subtraction operations as possible to preserve the original harmonic structure and subtle details of the speech to the greatest extent, avoiding speech distortion. Between extremely low and extremely high signal-to-noise ratios, the change of this parameter follows a non-linear curve (such as an S-curve), ensuring a smooth and natural transition from strong suppression to weak suppression.

[0047] After spectral subtraction of each frequency subband, subband weighted fusion and residual suppression are introduced. Specifically, the enhanced spectra of each frequency subband after spectral subtraction are weighted and fused according to the importance of the subband. The weights can be determined by the posterior signal-to-noise ratio or the perceived sensitivity of the human ear for the corresponding subband. For example, higher weights are assigned to mid-to-low frequency subbands that carry the main speech information and have higher perceived importance, while relatively lower weights are assigned to high frequency subbands with a larger proportion of noise, thereby highlighting key speech segments during the fusion process.

[0048] The spectral subtraction intensity parameter is a key control quantity used to regulate the noise reduction intensity during spectral subtraction denoising. It can be understood as an adjustable weight that modulates the amount of noise reduction in the spectral domain. Specifically, after obtaining the amplitude spectrum of the currently observed speech and the corresponding noise amplitude spectrum estimate, the noise spectrum is not simply subtracted from the observed spectrum. Instead, based on information such as the noise proportion and signal-to-noise ratio of the current frequency sub-band, a spectral subtraction intensity parameter is adaptively determined for that sub-band. This parameter is then used to scale the noise spectrum before the subtraction operation is performed. A larger value for the spectral subtraction intensity parameter indicates that more noise components are subtracted from the observed spectrum within that sub-band, resulting in stronger noise suppression. A smaller parameter value indicates that fewer noise components are subtracted, prioritizing the preservation of speech details.

[0049] Specifically, using the aforementioned adaptively determined spectral subtraction intensity parameters, amplitude-domain spectral subtraction is performed within each frequency sub-band to achieve targeted suppression of background noise. In practice, firstly, the currently observed amplitude spectrum (noisy speech) and the corresponding noise amplitude spectrum estimate are obtained for each frame and each sub-band. Then, the noise estimate is scaled proportionally according to the spectral subtraction intensity parameters for that sub-band. Subsequently, the scaled noise amplitude is subtracted from the observed amplitude to obtain the preliminary enhanced amplitude spectrum for that sub-band. The spectral subtraction intensity parameters are variable across the sub-band dimension, allowing for stronger subtraction in sub-bands with higher noise content and weaker subtraction in speech-dominated sub-bands, thus achieving a more suitable balance between noise suppression and preservation of speech details.

[0050] In practical applications, due to factors such as noise estimation errors, instantaneous non-stationary noise, and overlap between speech and noise spectra, the amplitude result after spectral subtraction may be over-subtracted, manifesting as outliers with local amplitudes that are negative or close to zero. Furthermore, to ensure the physical rationality of the enhanced spectrum and improve auditory stability, a spectral floor locking mechanism is introduced to constrain the spectral subtraction output. That is, when the spectral subtraction result at a certain frequency point (or a frequency point within a sub-band) falls below a preset lower limit, it is not directly set to zero, but rather clamped to a very small positive amplitude level. This spectral floor is usually set according to a fixed proportion of the noise spectrum estimation, or selected from local statistics of the noise spectrum, to ensure that a certain background noise component is retained even under strong suppression. This non-zero lower limit amplitude constraint avoids the sparse spectral structure caused by a large number of zero points in the spectrum, reducing the typical discrete narrowband residue (i.e., musical noise) phenomenon of spectral subtraction from the source. It also alleviates speech distortion and a dry or hollow auditory feel caused by over-suppression, making the enhanced speech more stable in terms of continuity, naturalness, and the usability of subsequent feature reconstruction.

[0051] Further optionally, in step S102, after performing spectral subtraction on the speech spectrum in each frequency sub-band using the spectral subtraction intensity parameter, the method further includes: setting a lower limit threshold for the spectral subtracted speech spectrum; when the amplitude of the spectral subtracted speech spectrum is lower than the lower limit threshold, the lower limit threshold or the minimum amplitude value of the corresponding frequency sub-band in the adjacent frame is used as a substitute to avoid the spectral components being directly set to zero.

[0052] Subsequently, the residual noise energy distribution in each frequency sub-band was further analyzed after weighted fusion. For sub-bands with high residual noise energy, the system can locally apply spectral enhancement or suppression strategies, such as further attenuating frequency ranges with prominent residual noise or appropriately compensating for over-suppressed speech segments, to reduce inconsistencies in energy and spectral shape between sub-bands. Through the above residual optimization processing, the spectral breakage or discontinuity problem caused by independent spectral subtraction of sub-bands can be effectively alleviated.

[0053] For example, weights are assigned to different sub-bands based on the masking effect and sensitivity curve of human hearing. For instance, higher retention weights are assigned to mid-to-low frequency sub-bands containing the fundamental frequency. For high-frequency sub-bands, if their signal-to-noise ratio remains low, lower weights are assigned for further attenuation. The processed spectral data from each sub-band are then spliced ​​or overlapped according to their original frequency index positions (depending on the partitioning method), and smooth transitions are performed at the boundaries to prevent spectral breaks, ultimately synthesizing a complete full-band enhanced amplitude spectrum.

[0054] Finally, the spectrum optimized by subband weighted fusion and residual suppression is used as the initial enhanced speech features and input into the subsequent metricGAN model for refined enhancement and reconstruction. For example, the system scans the entire frequency band to find isolated, high-energy narrowband peaks (i.e., energy points without neighborhood support on both the frequency and time axes) that exist independently on the time-frequency graph. These are often residual musical noise from spectral subtraction. Once such isolated points are detected, the system compares their energy differences with those of surrounding frequencies. If the difference is too large, the system smooths or replaces them using the energy average of adjacent frequencies or frames, thus smoothing them out. Through this optional embodiment, subband adaptive noise suppression and fusion optimization at the spectral domain level provides more consistent and stable input features for the deep model, thereby further improving the naturalness, continuity, and robustness of the overall speech enhancement effect.

[0055] Further optionally, in step S102, after performing spectral subtraction on the speech spectrum in each frequency sub-band using the spectral subtraction intensity parameter, the method further includes: adaptively adjusting the spectral subtraction result using a frequency-dependent nonlinear spectral subtraction function to obtain a nonlinearly adjusted speech spectrum; wherein, the adaptive adjustment includes: enhancing low amplitude values ​​in the spectral subtraction result, or conservatively suppressing high amplitude values ​​in the spectral subtraction result; in the noise spectrum estimation process, combining the statistical extreme values ​​of noise spectra from adjacent frames to obtain the changing peak values ​​of non-stationary noise, and filtering the estimated value of the noise spectrum based on the changing peak values ​​to obtain a smoothed noise spectrum; and fusing the nonlinearly adjusted speech spectrum with the smoothed noise spectrum to obtain the sub-band enhanced spectrum in each frequency sub-band.

[0056] In this embodiment, the frequency-dependent nonlinear spectral subtraction function can be understood as a transformation rule that operates simultaneously on both the frequency and amplitude axes. Depending on the frequency sub-band, the current spectral value, and the auditory importance and signal-to-noise ratio of that frequency band, the speech spectrum after spectral subtraction can be stretched or compressed in a differentiated and nonlinear manner. For example, in the mid-to-low frequency sub-bands carrying the pitch and formants, subtle signals that are still weak after spectral subtraction can be moderately stretched to compensate for the energy loss caused by the subtraction. In sub-bands with a high proportion of high-frequency noise, a more conservative suppression strategy is adopted for spectral peaks with high amplitudes to avoid over-amplifying occasional noise spikes. From an implementation perspective, this type of nonlinear function can set different gain curves in different frequency sub-bands. For example, in the low-frequency sub-band, a slightly higher gain is set for small amplitudes, while the gain for large amplitudes is close to one. In the high-frequency sub-band, a generally small gain is maintained, and a slight compression is applied to high-amplitude signals, thus reflecting the dual characteristics of frequency dependence and nonlinearity.

[0057] For example, a frequency-dependent nonlinear spectral subtraction function can be defined by establishing a nonlinear gain curve for the amplitude value after spectral subtraction on different frequency subbands. One possible approach is a piecewise gain curve. For each frequency subband, several amplitude ranges are first defined, such as low-amplitude, medium-amplitude, and high-amplitude ranges, and then different gain strategies are specified for different ranges. For instance, in the low-to-mid-frequency subband, a gain slightly greater than one can be given to the spectral subtraction result in the low-amplitude range to compensate for excessive attenuation caused by spectral subtraction, while a gain close to one is used for the spectral subtraction result in the high-amplitude range to keep strong speech components essentially unchanged. Conversely, in the high-frequency subband, a gain close to one or even slightly less than one can be used for the low-amplitude range to avoid amplifying high-frequency residual noise, while slight compression is used for the high-amplitude range to prevent high-spectral peaks from becoming excessively prominent. By segmenting by amplitude and setting thresholds and gains separately for each subband, a frequency-dependent piecewise nonlinear spectral subtraction function is formed.

[0058] The second option is a smooth, nonlinear curve with low-amplitude expansion and high-amplitude compression. This involves gently stretching or compressing the spectral subtraction result within each subband, rather than hard segmentation. Specifically, a curve can be designed that monotonically increases but gradually decreases: in the low-amplitude region, the output amplitude increases relatively faster, thus recovering weak speech details that have been excessively weakened by spectral subtraction; in the medium-amplitude region, the output is close to a linear relationship with the input, preserving the overall spectral structure as much as possible; in the high-amplitude region, the output amplitude gradually decreases or even slightly decreases, equivalent to slightly compressing the maximum spectral peaks to avoid excessive concentration of local energy. This curve can be set to encourage enhancement in the low-to-mid-frequency subbands to highlight speech harmonics and formants; while in the high-frequency subbands, the overall gain can be set lower, mainly to suppress abrupt spikes and musical noise.

[0059] The third option is a nonlinear function based on power-law amplitude transformation. For the same spectral subtraction result, different dynamic range adjustment methods can be used in different frequency subbands. For example, in the low-frequency and mid-frequency subbands, to improve the visibility of weak speech components, a dynamic range compression transformation can be applied to the amplitude, relatively boosting weak signals and relatively suppressing strong signals, thereby making the speech structure more balanced and the formants smoother. In the high-frequency subband, to avoid amplifying noise spikes, an extended dynamic range or a more conservative mapping can be used to significantly weaken weak noise and prevent further amplification of strong noise. By setting different dynamic range adjustment methods for different frequency bands, a power-law nonlinear spectral subtraction function that varies with frequency is constructed.

[0060] The fourth option is logarithmic or nonlinear adjustment in the perceptual domain. In this type of approach, instead of directly designing the function on the linear amplitude spectrum, the spectral subtraction result is first mapped to a domain more consistent with auditory characteristics. For example, the amplitude is logarithmically compressed or mapped to the energy of a Mel frequency sub-band. Then, nonlinear adjustments are implemented in this domain. For instance, in the low-frequency Mel sub-band, the boost to low-energy components can be designed to be relatively larger to enhance the low-frequency speech details that the human ear is sensitive to; in the high-frequency Mel sub-band, the boost to all components is designed to be relatively moderate, only moderately enhancing components that significantly exceed the noise floor, while maintaining or slightly attenuating components close to the noise level. Finally, the adjusted perceptual domain characteristics are mapped back to the actual spectral amplitude. By setting nonlinear gain curves for different sub-bands in the perceptual domain, it is possible to more closely match the subjective sensitivity of the human ear to different frequencies.

[0061] The fifth option is an adaptive nonlinear function incorporating posterior signal-to-noise ratio (SNR). In this design, the shape of the nonlinear spectral subtraction function depends not only on the frequency subband but also on the current SNR level of that subband. For example, in a low-frequency subband, when the posterior SNR is high, the function can be moderately boosted in the low-amplitude region to restore as much speech detail as possible. Once the posterior SNR of that subband decreases and the noise proportion increases, the boosting amplitude in the low-amplitude region is automatically tightened, or even slightly suppressed, to prevent noise amplification. In high-frequency subbands, when a low SNR is detected, the nonlinear function tends to be more suppressive overall, i.e., weak components are not boosted, and strong components are slightly compressed. When the SNR increases slightly, the suppression of mid-to-high amplitude components is moderately relaxed. In this way, the nonlinear spectral subtraction function deforms simultaneously on both the frequency and SNR axes, truly achieving frequency-dependent and state-dependent adaptive nonlinear adjustment.

[0062] Specifically, the previous stage completed preliminary spectral subtraction based on the spectral subtraction intensity parameter, obtaining linear spectral subtraction results within each frequency sub-band. Building upon this, the spectral values ​​within each frequency sub-band are iterated, with different adjustment strategies applied to frequencies with lower and higher amplitudes. For spectral components located near speech harmonics but whose amplitudes are significantly reduced after spectral subtraction, especially in low-frequency sub-bands, their amplitudes can be moderately increased to restore them to a level closer to real speech. For example, a formant frequency that should be clearly present, but whose amplitude after spectral subtraction is only slightly above the noise floor, can be given a slightly higher boost through a nonlinear function. For spectral peaks with amplitudes exceeding a set threshold or the average of adjacent peaks, especially in high-frequency sub-bands, if analysis suggests they may contain significant noise components, a conservative suppression approach is adopted. This involves slightly reducing the gain for these high amplitudes to avoid exceeding surrounding frequencies and thus preventing the formation of harsh noise spikes. In this way, the nonlinear spectral subtraction function is equivalent to further reshaping the spectrum by segmenting by amplitude and differentiating by frequency bands on the basis of linear spectral subtraction.

[0063] In the noise spectrum estimation process, the concepts of statistical extrema of noise spectrum in adjacent frames and peak values ​​of non-stationary noise changes are introduced primarily to address the problem of rapid changes in environmental noise over time and the potential for distortion in single-frame estimation. Specifically, for each frequency point or sub-band, not only is the noise estimate of the current frame recorded, but also the historical maximum, minimum, or more complex extrema characteristics of that frequency point are statistically analyzed within a sliding time window. By comparing the current noise estimate with these historical extrema, peak values ​​of non-stationary noise changes can be identified, such as sudden car horns, keyboard typing, and other short-duration, strong noises. When the system detects an abnormal surge in a frequency point in recent frames, exceeding the stationary noise level over a past period, this change can be considered a peak value of noise change, which can then be filtered or weakened when updating the noise model. For example, if the noise level of a certain high-frequency subband has remained at a moderate level for the past few dozen frames, but suddenly doubles in the current frame, the system will not immediately include this sudden increase in the noise spectrum estimation. Instead, it will only allow it to slowly approach the new level, or only use it partially, thereby avoiding the noise spectrum estimation being dragged up by short-term impulse signals.

[0064] Based on the analysis of statistical extrema and peak values, the system introduces a smoothing process into the noise spectrum estimation, resulting in a smoothed noise spectrum. A smoothed noise spectrum can be understood as imposing a gradual constraint on the noise estimate in the time dimension, preventing it from changing too rapidly between adjacent frames. In implementation, the system weights and fuses the noise estimate of the current frame with estimates from several previous frames. For example, the current estimate may only have a portion of the weight, while historical estimates have the remaining weight. Simultaneously, the weight of peak values ​​identified as changing values ​​is appropriately reduced, and historical values ​​may even be temporarily used to replace the current value at certain frequency points. The result is a noise spectrum that exhibits a continuous, smooth, and slowly changing trend over time, more realistically reflecting the long-term statistical characteristics of background environmental noise, while avoiding being dominated by transient impulse noise. The final smoothed noise spectrum serves as a reference background image in subsequent fusion processes, providing a stable noise baseline.

[0065] After completing the nonlinear adjustment of the speech spectrum and obtaining the smoothed noise spectrum, the system needs to fuse these two pieces of information to obtain the sub-band enhancement spectrum within each frequency sub-band. Conceptually, this step is not a simple superposition, but rather a dynamic allocation of weights between highlighting speech and suppressing noise based on the speech confidence and noise stability of each sub-band and each frequency point. For example, for frequency points in the low-frequency sub-band that have already undergone nonlinear enhancement and whose amplitude is significantly higher than the smoothed noise spectrum, the speech spectrum can be given greater weight, allowing it to dominate the fusion result; while for frequency points in the high-frequency sub-band that are close to the noise baseline and whose noise changes relatively smoothly, the smoothed noise spectrum can be given higher suppression weights, making the final enhancement spectrum closer to the form after the noise background is reduced. In some frequency bands close to the speech edge, the system can also fine-tune the fusion result based on the sub-band posterior signal-to-noise ratio, so that those ambiguous regions between speech and noise are smoothly transitioned without presenting obvious hard cuts. The resulting subband enhanced spectrum inherits the structural information of the nonlinearly adjusted speech spectrum and makes full use of the stable background constraints provided by the smoothed noise spectrum, thus forming a more natural, continuous, and robust spectral representation on different subbands for further refinement and reconstruction by subsequent deep models.

[0066] Step S103: Construct a metricGAN model including a generator and a discriminator.

[0067] Step S104: Input the initial enhanced speech features into the generator to obtain the second enhanced speech features, and input the second enhanced speech features into the discriminator. The discriminator performs a function fitting of the speech quality evaluation index on the second enhanced speech features to obtain the corresponding speech quality score.

[0068] In this embodiment, the metricGAN model can be understood as a generative adversarial network structure specifically designed for speech quality evaluation scenarios. Unlike traditional generative adversarial networks that primarily allow the discriminator to distinguish between real and fake samples, the core idea of ​​metricGAN is to use the discriminator to approximate the mapping relationship of an objective speech quality evaluation index, such as commonly used subjective or objective evaluation index scores, so that the scalar value output by the discriminator is as close as possible to the real speech quality score. Thus, the generator's optimization is directly guided by the goal of speech quality during training. In other words, metricGAN not only focuses on whether the enhanced speech is statistically close to the real speech, but also emphasizes the improvement of the perceived auditory quality index by the enhancement result.

[0069] Structurally, the metricGAN model mainly consists of a generator and a discriminator, which jointly model the speech in an adversarial manner. The generator takes noisy speech features or initial enhanced speech features as input and outputs second enhanced speech features, which are speech representations further refined based on traditional spectral subtraction or subband enhancement. The discriminator receives this second enhanced speech feature, or simultaneously receives the second enhanced speech feature and a reference clean speech feature, mapping them to a scalar speech quality score. This score is designed to approximate the true score of a certain objective quality indicator, thereby achieving an approximate fit to the functional relationship of the speech quality evaluation index. In this way, during the training phase, the metricGAN model constrains the generator to generate feature distributions that more closely resemble clean speech, while also directly driving the generator to optimize towards improving the speech quality index through the quality score output by the discriminator.

[0070] The generator can be built based on various deep neural network architectures. For example, a time-frequency mask estimation network based on a convolutional neural network or a convolutional encoder-decoder structure can be used to estimate an enhancement mask for the input time-frequency spectrum, and then multiply it with the original spectrum to obtain the enhancement result. Alternatively, a bidirectional recurrent neural network or a gated recurrent unit structure can be used to model the contextual dependencies in the time series, suitable for handling noise patterns and speech dynamics in long-term contexts. A time-series modeling framework based on attention mechanisms or transformer structures can also be used to perform global correlation modeling of multi-frame time-frequency features. In the example scenario, the generator can receive the initial enhanced speech features obtained through sub-band adaptive spectral subtraction and residual optimization. Based on this, it can further refine and adjust the energy distribution and time-frequency texture within each time frame and frequency sub-band, such as repairing details lost due to over-suppression, smoothing spectral discontinuities, and optimizing harmonic structures, to output a second enhanced speech feature that better meets the target speech quality requirements.

[0071] The discriminator is implemented as a regressive deep neural network. Its input can be a feature pair of enhanced speech and reference clean speech, or it can perform quality estimation based solely on enhanced speech features under no-reference conditions. In the reference case, the discriminator often employs a dual-path structure, encoding and extracting high-level representations for both enhanced and clean speech features separately. These representations are then compared and fused in the high-level space, for example, by concatenating, subtracting, or interactively modeling the two feature paths. Several fully connected or convolutional layers are then used to map this similarity and difference into a continuous quality score. In the no-reference case, the discriminator needs to learn speech quality cues from the single-path enhanced speech features, such as distortion level, residual noise level, spectral smoothness, and energy distribution rationality, and integrate these cues into a score. Regardless of the structure, the discriminator essentially approximates the functional mapping relationship from speech features to quality scores through learning from a large number of labeled samples, making its output an approximation of the speech quality evaluation metric.

[0072] In building a metricGAN, the first step is to determine the type of speech quality evaluation metric to be approximated. These metrics can be predicted values ​​of subjective listening scores, such as the average opinion score obtained from human listening tests, or existing objective speech quality evaluation metrics, such as evaluation measures specifically measuring speech clarity, intelligibility, or naturalness. Once the target metric is determined, a corresponding training dataset needs to be constructed. Each sample in the dataset contains the original clean speech, noisy speech with added noise, and a version of speech enhanced using traditional or existing methods, with a corresponding quality score calculated or labeled for each version. Subsequently, these samples are used to supervise the training of the discriminator, enabling it to output an estimate as close as possible to the true score when inputting enhanced speech features. Based on this, the generator and discriminator are arranged in an adversarial structure. On the one hand, the discriminator is expected to accurately approximate the metric score; on the other hand, the generator is expected to continuously deceive the discriminator, gradually improving the score of the enhanced speech, thereby optimizing the generator in terms of the target quality metric.

[0073] Regarding the enhanced speech features themselves, their form is typically a time-frequency domain representation suitable for deep model processing. In the embodiments of this application, the front end has obtained an initial enhanced speech spectrum through steps such as short-time transformation, subband division, spectral subtraction, and residual suppression. This spectrum contains two-dimensional structural information in both time and frequency dimensions. Enhanced speech features can be transformed in various ways based on this, such as directly using the amplitude spectrum or power spectrum as input, or performing logarithmic compression on the spectrum to enhance the learnability of the dynamic range. Furthermore, auditory perception-based features can be extracted, such as performing Mel filtering on the spectrum to obtain a Mel spectrum, or constructing a subband energy sequence based on frequency band energy. In some implementations, time-domain features such as phase-related information, short-time energy, zero-crossing rate, and pitch period can be jointly encoded with spectral features to provide the model with richer cues for distinguishing speech from noise. Regardless of the specific form, enhanced speech features should be able to reflect the time-frequency structure, energy distribution, formant positions, and residual noise morphology of the speech well, providing a sufficient information foundation for the generator and discriminator.

[0074] In the process of fitting a function to speech quality evaluation metrics based on these enhanced speech features, the discriminator essentially acts as a learnable metric calculator. During the training phase, for each training sample, the system feeds the enhanced speech features into the discriminator to obtain a predicted score, while simultaneously obtaining a target value from pre-calculated or labeled real quality scores. The training process continuously adjusts the parameters of the discriminator's internal network, gradually reducing the difference between the predicted and target scores, thus progressively approximating the calculated results of the real metrics. As training progresses, the discriminator learns to comprehensively judge speech quality from multiple aspects, including the distribution of residual noise in different frequency bands, the continuity of the speech signal in the time and frequency directions, spectral smoothness, and distortion patterns, thereby achieving function fitting for complex quality metrics. During the adversarial training phase, the generator's output is also scored by the same discriminator. By adjusting its enhancement strategy, the generator continuously improves the score given by the discriminator, essentially using this data-driven quality evaluation function in reverse, no longer relying solely on a manually set simple cost function, but directly optimizing towards the final speech quality goal. In this way, metricGAN tightly couples speech enhancement with speech quality assessment, ensuring that the generative network revolves around speech quality metrics throughout the structural design and parameter training process, thereby obtaining enhancement results that are more in line with the subjective perception of the human ear.

[0075] Step S105: Based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning. The joint training includes at least quality constraints on the second enhanced speech features at different time scales or spectral scales, and optimizing and updating the model parameters of the metricGAN model.

[0076] In this embodiment, the speech quality score refers to a scalar evaluation value used to quantify the subjective or objective quality of speech. During the training phase, the discriminator outputs a continuous quality score based on the input second enhanced speech feature (and optional reference clean speech features). This score can be designed to approximate traditional objective speech quality indicators, such as objective predicted values ​​related to subjective listening scores, or indicators related to speech clarity, intelligibility, and naturalness. In other words, the speech quality score can be seen as the output of a comprehensive scoring function for speech quality learned internally by the discriminator. Through learning from a large number of labeled samples, it gradually acquires a discriminative ability that is highly correlated with real quality indicators. During adversarial training, this score serves as a supervisory signal for the discriminator's training, making it more accurately approximate the real quality indicators. It also serves as an optimization target for the generator's training, causing the generator to tend to output enhanced speech features that achieve higher scores.

[0077] As an optional embodiment, in step S105, based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning, including: using Wasserstein distance as the discriminator loss function, combining gradient penalty to constrain the gradient norm of the discriminator, and using Wasserstein distance and speech perception features to jointly construct the generator loss function, thus establishing a multi-objective loss function; constructing multi-scale representations of the second enhanced speech features at different time resolutions or different frequency resolutions; calculating speech quality scores at different time scales or spectral scales based on the multi-scale representations; and jointly optimizing the model parameters of the generator and discriminator using the multi-objective loss function according to the speech quality scores at different time scales or spectral scales, so as to maximize the consistency of the speech quality scores of the second enhanced speech features output by the optimized generator at different time scales or spectral scales.

[0078] In joint training, this embodiment employs adversarial learning to alternately optimize the generator and discriminator. The discriminator uses Wasserstein distance as a loss function to measure the difference between the distribution of real, high-quality speech and the distribution of generated, enhanced speech. Compared to traditional binary classification methods that only distinguish between real and generated speech, Wasserstein distance constrains the discriminator from the perspective of distance between distributions, ensuring that its output not only differentiates between real and generated samples but also reflects the magnitude of the difference. To ensure the discriminator meets the smoothness requirement of this distance metric, this embodiment introduces gradient penalty to constrain the gradient norm of the discriminator in the sample space, preventing gradient explosion or severe oscillations. The generator's loss function consists of two parts: one part comes from the adversarial term of the Wasserstein distance, encouraging the generator to narrow the gap between the enhanced and target speech distributions; the other part comes from speech perception features, taking into account the differences between the enhanced speech and the reference clean speech in the perceptual domain (e.g., Mel spectrum, speech masking features, auditory filtering features, etc.), allowing the generator to improve its adversarial score while also closely resembling the speech structure that the human ear is sensitive to. Together, they constitute a multi-objective loss, which takes into account both adversarial distribution matching and perceptual consistency when optimizing the generator.

[0079] Regarding multi-scale representations at different time or spectral scales, this embodiment emphasizes quality constraints on the second enhanced speech features from multiple observation scales. The time scale can be understood as the observation dimension of speech at different temporal resolutions. For example, using shorter frame lengths and smaller time steps yields a high-temporal-resolution time-frequency map to capture details such as instantaneous transitions and plosives. Simultaneously, using longer frame lengths or downsampling / pooling in the time dimension yields a low-temporal-resolution representation to focus on sentence rhythm, energy envelope, and noise stability over a longer time range. The spectral scale refers to representing the frequency axis with different levels of precision. For example, using finer frequency divisions across the entire frequency band to characterize fine-grained harmonic structures, while constructing coarse-grained frequency representations based on sub-band aggregation, Mel filtering, or multi-band filtering to evaluate overall frequency band energy distribution and perceptual balance. Through these operations, a set of multi-scale feature representations can be constructed for the same enhanced speech segment, including features emphasizing local details, features emphasizing global smoothness, features highlighting low-frequency speech components, and features focusing on high-frequency noise residue.

[0080] After obtaining the multi-scale representation, the corresponding speech quality score is calculated at each time scale or spectral scale. Specifically, a multi-branch or multi-head output structure can be designed in the discriminator, or after sharing the feature extraction layer, features at different scales can be mapped to their respective quality scores through sub-networks. For example, for high temporal resolution representation, the discriminator can evaluate the instantaneous clarity of the speech, whether plosives are distorted, and whether transitions are natural. For low temporal resolution representation, it can evaluate the energy balance of the entire speech, the stability of background noise over time, and the overall naturalness. For fine-grained spectral representation, the discriminator can focus on the formant shape and the continuity of harmonic structure. For coarse-grained frequency band representation, it evaluates the energy ratio of low frequencies to high frequencies, the degree of inter-band balance, and whether residual noise is concentrated in certain frequency bands. Each scale corresponds to one or more quality scores, and these scores are combined to form the multi-scale quality evaluation result.

[0081] Quality constraints on the second-enhanced speech features at different time or spectral scales require that the quality scores of the second-enhanced speech features output by the generator be sufficiently high at each scale, and that they exhibit good consistency across scales. Quality constraints have two levels: an absolute level, aiming for scores at each scale to be as close as possible to the reference scores of high-quality speech, reflecting improvements in both detail and overall quality; and a relative level, aiming for no significant contradictions between scores at different scales. For example, there shouldn't be situations where local quality appears good but the overall listening experience is poor, or where low-frequency quality is high but high-frequency noise is heavy. In multi-objective loss, constraining the score differences across different scales prevents the generator from exploiting weaknesses at a single scale, instead ensuring a relatively balanced and coordinated quality improvement across all scales. This results in a more natural and stable subjective perception of the speech enhancement effect.

[0082] At the model parameter level, the metricGAN model consists of a generator and a discriminator, each with a large number of learnable parameters, such as convolutional kernel weights, recurrent unit weights, normalization layer parameters, fully connected layer weights, and bias terms. These parameters collectively determine how the generator maps input features to second-enhanced speech features, and how the discriminator outputs a quality score based on these features. During joint training, parameter optimization employs a typical adversarial training paradigm. In each training epoch, the generator is first fixed, and the discriminator parameters are updated to enable the discriminator to more accurately distinguish the quality differences between real high-quality speech features and generated enhanced speech features across multiple scales, while ensuring that its output score is as consistent as possible with the expected speech quality metric. Then, the discriminator is fixed, and the generator parameters are updated so that the second-enhanced speech features output by the generator receive a higher multi-scale quality score in the discriminator's eyes. During the optimization process, multi-objective loss is used to integrate the Wasserstein distance term, multi-scale quality constraint term, and speech perception feature term. These objectives are considered simultaneously during each parameter update, so that the generator and discriminator continuously iterate around the core requirement of improving speech quality scores and maintaining consistency across different time and spectral scales throughout the training process.

[0083] Through the aforementioned joint training mechanism, the metricGAN model in this embodiment no longer merely pursues noise reduction or artifact elimination at a single scale, but rather constrains the quality performance of the second enhanced speech features holistically across multiple temporal and spectral scales. The discriminator, through learning multi-scale quality scores, becomes a speech quality evaluator that considers both local and overall performance. The generator, under such evaluation feedback, continuously adjusts its parameters, gradually learning to simultaneously improve speech clarity, naturalness, and stability at different scales, thereby achieving comprehensive optimization of the overall auditory experience of the entire speech segment.

[0084] As an optional embodiment, step S106, reconstructing the enhanced speech signal based on the generated third enhanced speech features, includes: performing an inverse short-time Fourier transform on the third enhanced speech features to obtain a time-domain signal, and reconstructing the time-domain signal into a continuous speech waveform using an overlap-add algorithm; extracting cochlear perception-sensitive phase features from historical speech waveforms using a sensitive phase encoder; performing nonlinear adjustment on the phase spectrum of the reconstructed continuous speech waveform using a multi-head self-attention phase compensation network, combined with the extracted phase features, to reduce phase distortion in the phase spectrum; performing time-domain smoothing on the reconstructed continuous speech waveform using time-domain mask constraints; and outputting the phase-optimized and time-domain smoothed continuous speech waveform to the user terminal as the enhanced speech signal.

[0085] In this optional embodiment, step S106 focuses on converting the third enhanced speech features into a final playable high-quality enhanced speech signal. The entire process, from frequency domain reconstruction to phase optimization and then to time domain smoothing, involves the collaborative design of signal processing and deep network modeling. First, the third enhanced speech features are usually derived from the frequency domain representation output by the front-end metricGAN generator, which can be the enhanced amplitude spectrum, real and imaginary part spectrum, or the spectrum after subband rearrangement. Following parameter configurations that correspond exactly to the front-end analysis, these frequency domain features are restored to a complex spectrum form, i.e., each time frame and each frequency point is equipped with corresponding amplitude information and initial phase information. The initial phase can either directly use the phase of the original noisy speech or use a phase approximation obtained from the front-end network estimation. Then, an inverse short-time Fourier transform is performed on each frame of complex spectrum, using a window function, frame length, and frame shift that match the analysis stage to map the frequency domain signal back to a frame-level time domain waveform. The resulting frame-level waveforms have certain overlapping regions in time. The system employs an overlap-addition algorithm, shifting and superimposing each frame according to its starting sample position on the time axis. During superposition, a synthesis window function weights the overlapping regions to ensure that the reconstructed continuous waveform maintains smooth amplitude and avoids obvious splicing artifacts in the transition region. Through this process, the third enhanced speech feature is converted into a continuous, playable, but not fully corrected phase and temporal details enhanced speech waveform, providing input for subsequent neural network-level phase compensation and temporal smoothing.

[0086] After obtaining the initially reconstructed continuous speech waveform, to better correct the phase distortion problem that traditional frequency domain enhancement methods struggle to handle, this embodiment constructs a sensitive phase encoder to extract phase features closely related to cochlear perception from historical speech waveforms. The sensitive phase encoder can be considered a feature extraction network specifically designed for phase information. Its input includes historical speech waveforms from the previous few frames or a time window, and optionally incorporates derived features such as short-time phase spectrum or group delay. Structurally, this encoder can employ a multi-layer one-dimensional convolutional neural network or a time-frequency two-dimensional convolutional network. The kernel length and stride are set according to the pitch period and formant characteristics of the speech, enabling it to capture phase-time patterns at different time scales and perceive the coupling relationship between phase and amplitude in different frequency bands. To simulate the cochlea's non-uniform sensitivity to phase, the encoder can employ a cochlear filter-like sub-band structure in frequency band division. The input waveform is decomposed through a set of parallel bandpass convolution kernels, and phase-related instantaneous features, such as period alignment and cross-period phase shift, are extracted from each sub-band. Finally, these features are aggregated along the channel dimension by a feature fusion module to form a compact yet information-rich phase-aware feature representation. These features are implicitly correlated with specific speaker characteristics, pronunciation patterns, and environmental reverberation features within the model, providing auditory-meaning prior information for the subsequent phase compensation network.

[0087] Next, to perform high-precision nonlinear adjustment of the phase spectrum in the reconstructed speech, this embodiment designs a phase compensation network based on a multi-head self-attention mechanism. The network's input consists of the phase spectrum or complex spectrum of the current continuous speech waveform obtained from the short-time Fourier transform, and the phase-aware features output by the sensitive phase encoder. These two are combined through channel splicing, additive fusion, or gated injection to form a joint feature containing both the current phase state and historical phase priors. At the network level, the phase compensation network can adopt a structure similar to a transformer encoder, consisting of multiple stacked multi-head self-attention sublayers and feedforward sublayers. Each layer first establishes long-range dependencies for each time-frequency unit in the time and frequency dimensions through a self-attention mechanism. The multi-head structure allows different attention heads to focus on different frequency bands, different time spans, and different phase patterns. For example, some attention heads specifically model the phase continuity in the low-frequency region, while others focus on capturing phase fluctuations in high-frequency details. The feedforward network following the attention sublayer consists of several fully connected or convolutional layers, used to perform nonlinear transformations on the features after attention convergence, outputting a phase compensation amount or optimized phase value for each time-frequency point. The network as a whole maintains training stability and gradient flow through layer normalization and residual connections. During the training phase, the goal of the phase compensation network is to minimize the difference in phase structure between the compensated spectrum and the reference clean speech spectrum, and through subjective quality-oriented perceptual loss constraints, enable the network to learn to maintain the naturalness and spatiality of the speech while reducing phase distortion. The parameters required for training include the number of attention heads, attention dimension, number of layers, and convolutional kernel size. These parameters can be configured appropriately based on specific computational resources and target application scenarios (such as call enhancement, recording restoration, etc.).

[0088] After phase optimization, the frequency domain representation of the continuous speech waveform has been corrected in the phase dimension. However, there may still be a small amount of subtle temporal irregularities introduced by phase adjustment and early spectral subtraction, such as short-term spikes, slight ringing, or local waveform coarsening. To further improve temporal smoothness and auditory continuity, this embodiment uses a mask constraint method for smoothing in the temporal domain and introduces a Gaussian weighted self-attention mechanism similar to that in the TGSA structure, combining deep temporal modeling with local Gaussian weights. Specifically, the system constructs a temporal smoothing network. This network takes the phase-optimized continuous speech waveform as input and can first extract local temporal features, such as short-term energy changes, zero-crossing rate, and local envelope shape, through several one-dimensional convolutional layers. Then, based on these features, a Gaussian weighted self-attention module is introduced, which assigns higher weights to neighboring time points and gradually decreasing weights to distant time points in the self-attention calculation. This Gaussian weighting acts on the attention score or attention weight, enabling the network to perceive longer temporal contexts during sequence modeling while prioritizing local continuity in the smoothing operation. The temporal features output by the Gaussian-weighted self-attention module are then converted into a temporal mask sequence through a set of feedforward layers or convolutional residual blocks. The mask values ​​are constrained within an appropriate range, such as fluctuating slightly around a certain value, to perform dot-multiplication or additive fine-tuning on the original waveform. During training, the temporal mask network learns to suppress local anomalous fluctuations and unnatural ringing while maintaining the clarity of the main speech content through joint constraints from loss terms such as the difference from the reference clean speech waveform, short-term smoothness metrics, and spectral smoothness metrics. The introduction of Gaussian-weighted self-attention makes the model more adaptable to speech with different speeds and rhythms, especially suitable for continuous speech enhancement in call scenarios and noisy environments.

[0089] In summary, the input-output relationships between the modules in this embodiment have a clear intrinsic connection: the third enhanced speech feature, as the frequency domain enhancement result, is input to the inverse short-time Fourier transform and overlap-add module, and the output is a continuous time-domain waveform; this waveform is input to the sensitive phase encoder as a historical speech sequence in the time dimension to generate phase-aware features; the phase-aware features and the phase spectrum derived based on the same waveform are input to the multi-head self-attention phase compensation network, and the output is the phase-optimized spectrum or the corresponding time-domain waveform; these optimized time-domain data are further input to the time-domain smoothing network with Gaussian weighted self-attention to generate a time-domain mask and fine-tune the waveform, and the final output continuous speech waveform is the enhanced speech signal for the user end. In terms of the training process, the sub-networks can be trained jointly end-to-end or pre-trained in stages. For example, the sensitive phase encoder and phase compensation network can be trained first on the frequency domain enhancement and phase compensation datasets, and then the time-domain smoothing network can be trained on a large-scale corpus of clean speech with reference, and finally fine-tuned under a unified multi-task loss. Therefore, this embodiment closely integrates traditional frequency domain speech enhancement with deep phase modeling and temporal attention smoothing, so that the generated enhanced speech signal exhibits high quality and high robustness in both objective indicators and subjective listening experience.

[0090] As an optional embodiment, in step S106, after reconstructing the enhanced speech signal based on the generated third enhanced speech features, the enhanced speech signal can be input into a pre-trained speech quality assessment model; if the reconstructed speech quality score is lower than the set reconstruction quality threshold, the fine-tuning iteration of the joint training process of the metricGAN model is triggered.

[0091] In this optional embodiment, after completing frequency domain reconstruction, phase compensation, and temporal smoothing based on the third enhanced speech features to obtain the final enhanced speech signal, the system further introduces a pre-trained speech quality assessment model to perform quality loop detection on the enhanced speech, thereby achieving adaptive fine-tuning and updating of the metricGAN model. Specifically, the pre-trained speech quality assessment model can be constructed using a deep network structure highly correlated with objective speech quality evaluation standards such as PESQ or POLQA. This model typically includes three layers: a feature extraction front-end, a temporal modeling body, and a quality scoring regression head. The feature extraction front-end is responsible for converting the input temporal enhanced speech signal into a time-frequency or perceptual feature representation suitable for network processing. For example, it performs short-time Fourier transform or Mel-filter analysis on the waveform to obtain the energy distribution of speech in time and frequency, and then further extracts local time-frequency patterns through several convolutional layers. The temporal modeling body can employ a bidirectional long short-term memory network, a gated recurrent unit network, or a self-attention-based sequence encoder structure to integrate dynamic information of speech over a longer time span, capturing the changing trend of speech quality across the entire sentence and phenomena such as noise residue, distortion, and discontinuity. The quality score regression head consists of one or more fully connected layers or one-dimensional convolutional layers. It compresses the high-dimensional features output by the temporal modeling subject into a single scalar or a few score values ​​representing different dimensions of speech quality. Through pre-training on a large-scale speech quality annotation dataset (such as standard corpus with PESQ or POLQA scores), the score output by the model is highly aligned with the real target index in terms of numerical value.

[0092] During system operation, after the metricGAN model enhances the noisy speech and outputs the reconstructed enhanced speech signal, the enhanced speech signal, along with optional clean reference speech (available in certain controlled scenarios), is input into the pre-trained speech quality assessment model. The model provides one or more quality scores based on the time-frequency structure, residual noise morphology, and distortion characteristics of the input signal. The system pre-sets a reconstruction quality threshold based on the target application scenario. This threshold can be calibrated according to voice communication standards, terminal device requirements, or user subjective experience requirements; for example, for a call scenario, it can be set to a score level corresponding to a specific MOS score. When the score output by the quality assessment model is higher than or equal to the reconstruction quality threshold, it indicates that the current metricGAN's enhancement effect in this noisy environment and speaking conditions meets expectations, and there is no need to immediately adjust the model parameters; the enhanced speech signal can be directly output to the user terminal for playback or subsequent recognition tasks. When the score is lower than the set threshold, the system considers that the current model has certain quality deficiencies in this scenario, and the metricGAN model needs to be specifically optimized through online or quasi-online fine-tuning mechanisms.

[0093] After triggering a fine-tuning iteration, the system organizes the various data involved in the current enhancement process into a training sample chain. This chain includes the original noisy speech signal, the corresponding third-level enhanced speech features, the reconstructed enhanced speech waveform, optional clean reference speech, and the quality score and preset quality threshold or target score provided by the pre-trained quality assessment model. This input-output correspondence establishes an explicit link between scene conditions, model output, and quality feedback. During fine-tuning, the generator and discriminator parameters of metricGAN do not need to be completely retrained. Instead, some parameters are selectively made available for updating, such as the mid-to-high-level convolutional or attention modules responsible for time-frequency detail reconstruction in the generator, and the quality regression layers near the output in the discriminator. Simultaneously, the lower-level feature extraction layers and the overall structure remain stable to avoid catastrophic forgetting under limited sample conditions. Regarding training loss, in addition to the Wasserstein distance and perceptual feature losses, a new loss related to quality score bias is added. For example, the gap between the score output by the quality assessment model and the target quality threshold or ideal score is mapped as an additional optimization objective. This ensures that the generator update direction is not only driven by adversarial and spectral perception constraints, but also explicitly adjusted towards improving the score of the pre-trained quality assessment model. To integrate with specific domains or scenarios, the system can classify these low-quality samples according to noise type, scene label, and terminal device type. During fine-tuning, scene labels can be introduced as additional input, or a conditional generative adversarial structure can be adopted to enable metricGAN to learn different enhancement strategies under different scene conditions, achieving adaptive enhancement for specific scenarios.

[0094] In fine-tuning parameter settings, to ensure system stability and prevent overfitting, this embodiment typically uses a smaller learning rate than in the offline training phase and limits the number of iterations or update batches triggered for each fine-tuning. For example, after detecting that the quality scores of several voice samples are continuously below a threshold, a short-cycle fine-tuning training is performed on this batch of samples, and then the normal inference mode is restored. During training, the score changes of the fine-tuned generated voice samples under the pre-trained quality assessment model can be continuously monitored. The new parameters are only retained when the scores significantly improve without causing a performance degradation on other scenario test sets; otherwise, the parameter configuration can be reverted to the previous stable version. Through this closed-loop mechanism of quality assessment-driven and fine-tuning training feedback, this embodiment achieves dynamic adaptation between the metricGAN model and real-world application scenarios. The pre-trained quality assessment model is responsible for objectively quantifying the enhancement results at the perceptual quality level, while the metricGAN model continuously adjusts the internal connection relationship and weight distribution of the generator and discriminator under the guidance of this quantification feedback, so that it can maintain output performance close to the target quality threshold in different scenarios such as call noise reduction, in-vehicle voice enhancement, and conference sound pickup, thereby improving the actual user experience and environmental robustness of the voice enhancement system.

[0095] In this embodiment, by using a baseline facial model for identity matching and voice-driven facial motion state sequence modeling, cross-modal reconstruction from voice to facial image is achieved in diverse application scenarios and large-scale user environments. This improves the stability of identity features in cross-modal modeling, as well as the synchronization rate between facial image and voice changes, and enhances the image quality output by cross-modal modeling.

[0096] Having introduced the methods of exemplary embodiments of this application, the following describes a speech enhancement system based on the metricGAN model, according to an exemplary embodiment of this application. (Reference) Figure 2 The system includes the following modules: a preprocessing module, used to acquire the noisy speech signal to be processed and preprocess the noisy speech signal to obtain time-frequency features characterizing the speech energy distribution; an enhancement module, used to perform adaptive spectral domain noise suppression processing based on the time-frequency features to obtain initial enhanced speech features for subsequent modeling, wherein the adaptive spectral domain noise suppression processing is used to introduce noise suppression priors and reduce background noise interference, the adaptive spectral domain noise suppression processing includes noise estimation and spectral subtraction operations on different frequency sub-bands respectively, and adaptively adjusting the noise suppression intensity according to the signal-to-noise ratio of each frequency sub-band; a construction module, used to construct a metricGAN model including a generator and a discriminator; and a training module. The system is configured to input the initial enhanced speech features into a generator to obtain second enhanced speech features, and then input the second enhanced speech features into a discriminator. The discriminator fits a speech quality evaluation index function to the second enhanced speech features to obtain a corresponding speech quality score. Based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning. The joint training includes at least quality constraints on the second enhanced speech features at different time scales or spectral scales, and optimizing and updating the model parameters of the metricGAN model. The reconstruction module is used to process the noisy speech signal using the trained generator and reconstruct the enhanced speech signal based on the generated third enhanced speech features. The above system can implement the steps described in the above method implementation, and the specific implementation methods of each step will not be repeated here.

[0097] After introducing the methods and systems of exemplary embodiments of this application, a terminal device according to an exemplary embodiment of this application will be described next. This terminal device can implement the steps described in the above method embodiments, and the specific implementation of each step will not be repeated here. It should be noted that the above embodiments are only specific embodiments of this application, used to illustrate the technical solutions of this application, and not to limit it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in this application, or make equivalent substitutions for some of the technical features. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A speech enhancement method based on the metricGAN model, characterized in that, The method includes: Acquire the noisy speech signal to be processed, and preprocess the noisy speech signal to obtain time-frequency features characterizing the speech energy distribution; Adaptive spectral domain noise suppression processing is performed based on the time-frequency features to obtain initial enhanced speech features for subsequent modeling. The adaptive spectral domain noise suppression processing is used to introduce noise suppression priors and reduce background noise interference. The adaptive spectral domain noise suppression processing includes noise estimation and spectral subtraction operations for different frequency sub-bands, and adaptive adjustment of noise suppression intensity according to the signal-to-noise ratio of each frequency sub-band. Construct a metricGAN model including a generator and a discriminator; input the initial enhanced speech features into the generator to obtain the second enhanced speech features, and input the second enhanced speech features into the discriminator. The discriminator fits the second enhanced speech features with a function of speech quality evaluation index to obtain the corresponding speech quality score. Based on the speech quality score, the generator and discriminator are jointly trained using adversarial learning. The joint training includes at least quality constraints on the second enhanced speech features at different time scales or spectral scales, and optimizing and updating the model parameters of the metricGAN model. The trained generator is used to process the noisy speech signal, and the enhanced speech signal is reconstructed based on the generated third enhanced speech features.

2. The speech enhancement method based on the metricGAN model according to claim 1, characterized in that, The preprocessing of the noisy speech signal to obtain time-frequency features characterizing the speech energy distribution includes: The noisy speech signal is subjected to frame segmentation and windowing processing; Perform a short-time Fourier transform on the windowed speech frame to obtain the time spectrum of the speech signal; By combining auditory perception sensitivity, frequency-weighted processing is performed on the time spectrum to obtain the time-frequency characteristics, wherein the gain coefficients of the low-frequency and high-frequency ranges are adaptively adjusted through training data or psychoacoustic models.

3. The speech enhancement method based on the metricGAN model according to claim 1, characterized in that, The step of performing noise estimation and spectral subtraction on different frequency sub-bands, and adaptively adjusting the noise suppression intensity based on the signal-to-noise ratio of each frequency sub-band, includes: According to a preset frequency sensing scale, the time spectrum is divided into multiple continuous and non-overlapping frequency sub-bands; Within each frequency sub-band, based on statistical information from non-speech intervals or adjacent multiple frames, independent noise estimation is performed on the background noise spectrum corresponding to each frequency sub-band to obtain the speech spectrum and noise spectrum within each frequency sub-band. Based on the speech spectrum and noise spectrum within each frequency sub-band, the posterior signal-to-noise ratio (SNR) of each frequency sub-band is calculated; the posterior SNR is used to characterize the instantaneous ratio of speech signal to noise signal within each frequency sub-band. Based on the posterior signal-to-noise ratio, the spectral subtraction intensity parameter corresponding to each frequency sub-band is adaptively determined; wherein, the spectral subtraction intensity parameter changes with the posterior signal-to-noise ratio in a nonlinear mapping manner; Using the spectral subtraction intensity parameter, spectral subtraction is performed on the speech spectrum in each frequency sub-band, and the results of the spectral subtraction in each frequency sub-band are fused into the initial enhanced speech feature.

4. The speech enhancement method based on the metricGAN model according to claim 3, characterized in that, After performing spectral subtraction on the speech spectrum within each frequency subband using the spectral subtraction intensity parameter, the method further includes: Set a lower limit threshold for the spectral spectrum of the subtracted speech spectrum; When the amplitude of the speech spectrum after spectral subtraction is lower than the lower limit threshold, the lower limit threshold or the minimum amplitude value of the corresponding frequency sub-band in the adjacent frame is used as the replacement.

5. The speech enhancement method based on the metricGAN model according to claim 3, characterized in that, After performing spectral subtraction on the speech spectrum within each frequency subband using the spectral subtraction intensity parameter, the method further includes: The frequency-dependent nonlinear spectral subtraction function is used to adaptively adjust the spectral subtraction result to obtain the nonlinearly adjusted speech spectrum; wherein, the adaptive adjustment includes: enhancing the low amplitude value in the spectral subtraction result, or conservatively suppressing the high amplitude value in the spectral subtraction result; In the noise spectrum estimation process, the statistical extreme values ​​of the noise spectrum of adjacent frames are combined to obtain the changing peak values ​​of non-stationary noise, and the estimated value of the noise spectrum is filtered based on the changing peak values ​​to obtain a smooth noise spectrum. The nonlinearly adjusted speech spectrum is fused with the smoothed noise spectrum to obtain the sub-band enhanced spectrum within each frequency sub-band.

6. The speech enhancement method based on the metricGAN model according to claim 1, characterized in that, The step of jointly training the generator and discriminator based on the speech quality score using adversarial learning includes: Wasserstein distance is used as the discriminator loss function, combined with gradient penalty to constrain the gradient norm of the discriminator, and Wasserstein distance and speech perception features are used together to form the generator loss function, thus establishing a multi-objective loss function; Multi-scale representations of the second enhanced speech features are constructed at different time resolutions or different frequency resolutions; Based on the multi-scale representation, speech quality scores are calculated at different time scales or spectral scales respectively; Based on the speech quality scores at different time scales or spectral scales, the model parameters of the generator and discriminator are jointly optimized using a multi-objective loss function to maximize the consistency of the speech quality scores of the second enhanced speech feature output by the optimized generator at different time scales or spectral scales.

7. The speech enhancement method based on the metricGAN model according to claim 1, characterized in that, The process of reconstructing the enhanced speech signal based on the generated third enhanced speech features includes: The third enhanced speech feature is subjected to inverse short-time Fourier transform to obtain a time-domain signal, and the time-domain signal is reconstructed into a continuous speech waveform through an overlap-add algorithm; A sensitive phase encoder is used to extract cochlear-sensory sensitive phase features from historical speech waveforms; By using a multi-head self-attention phase compensation network and combining the extracted phase features, the phase spectrum in the reconstructed continuous speech waveform is nonlinearly adjusted. Temporal smoothing is performed on the reconstructed continuous speech waveform using temporal mask constraints. The phase-optimized and time-domain-smoothed continuous speech waveform is output to the user terminal as the enhanced speech signal.

8. The speech enhancement method based on the metricGAN model according to claim 1, characterized in that, After reconstructing the enhanced speech signal based on the generated third enhanced speech features, the process further includes: The enhanced speech signal is input into a pre-trained speech quality assessment model; If the reconstructed speech quality score is lower than the set reconstruction quality threshold, fine-tuning iterations of the joint training process of the metricGAN model are triggered.

9. A speech enhancement system based on the metricGAN model, characterized in that, The system includes the following modules: The preprocessing module is used to acquire the noisy speech signal to be processed and to preprocess the noisy speech signal to obtain time-frequency features characterizing the speech energy distribution. An enhancement module is used to perform adaptive spectral domain noise suppression processing based on the time-frequency features to obtain initial enhanced speech features for subsequent modeling. The adaptive spectral domain noise suppression processing is used to introduce noise suppression priors and reduce background noise interference. The adaptive spectral domain noise suppression processing includes noise estimation and spectral subtraction operations on different frequency sub-bands respectively, and adaptively adjusting the noise suppression intensity according to the signal-to-noise ratio of each frequency sub-band. Build modules are used to construct metricGAN models that include generators and discriminators; The training module is used to input the initial enhanced speech features into the generator to obtain the second enhanced speech features, and input the second enhanced speech features into the discriminator. The discriminator fits the second enhanced speech features with a function of speech quality evaluation index to obtain the corresponding speech quality score. Based on the speech quality score, the generator and discriminator are jointly trained using an adversarial learning approach. The joint training includes at least quality constraints on the second enhanced speech features at different time scales or spectral scales, and optimizing and updating the model parameters of the metricGAN model. The reconstruction module is used to process the noisy speech signal using the trained generator and reconstruct the enhanced speech signal based on the generated third enhanced speech features.

10. An electronic device, characterized in that, The electronic device is used to implement the speech enhancement method based on the metricGAN model as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Voice processing method and electronic equipment

    CN121985068A