An adaptive speech noise reduction method and system in a multi-noise scene

By combining wavelet analysis and nested network modulation techniques, the problem of speech denoising under multiple noise levels and scenarios is solved, improving speech clarity and naturalness, adapting to different noise scenarios and maintaining stable performance.

CN121565189BActive Publication Date: 2026-04-21HUNAN XIAOYU ZHIHE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN XIAOYU ZHIHE TECHNOLOGY CO LTD
Filing Date
2026-01-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing speech noise reduction technologies struggle to adapt in noisy and multi-scene environments, resulting in decreased speech clarity and naturalness, and they also struggle to maintain stable performance in devices with limited computing power.

Method used

Wavelet analysis is used to perform cross-scale structural decomposition and directional texture decomposition, constructing a multi-level scale feature set. Structural details are jointly modulated by the encoder and decoder ends of a nested network. Combined with multi-target loss constraints and hierarchical supervision, end-to-end parameter learning and optimization are achieved.

Benefits of technology

In complex noisy environments, robust extraction of speech backbone information and detailed textures is achieved, enabling adaptive adjustment of feature fusion and improving the dynamic enhancement and high-consistency reconstruction of speech in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565189B_ABST
    Figure CN121565189B_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive speech denoising method and system for multi-noise scenarios, belonging to the field of speech denoising technology. The method includes: S1, preprocessing input speech data, operational control training support data, and wavelet analysis configuration data; S2, performing cross-scale structural decomposition and directional texture decomposition analysis, and applying prior quality constraints; S3, constructing a structural morphology consistent with the internal feature maps of the deep network; S4, executing structural modulation and detail coupling mechanisms; S5, constructing a loss constraint system and a hierarchical supervision structure; and S6, performing coefficient adjustment, strategy switching, and network path control. This invention solves the problems of existing denoising methods in noisy environments and conference calls, which are prone to generalization failure under non-stationary noise and environmental / equipment changes, insufficient consideration of global energy and high-frequency details, and lack of adaptive control due to time delay and computational constraints, resulting in unstable achievement of clarity and audibility standards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech denoising technology, specifically to an adaptive speech denoising method and system for multi-noise scenarios. Background Technology

[0002] Currently, voice processing technology is widely used in communications, smart devices, in-vehicle systems, wearable terminals, remote conferencing, and voice interaction. The demand for voice signal acquisition, transmission, and processing in various environments is increasing daily. As device usage scenarios expand, the types of noise encountered in voice acquisition are becoming increasingly diverse, including traffic noise, conversations, mechanical equipment noise, keyboard typing, indoor air conditioning noise, wind noise, and reverberation. These background noises often have characteristics such as suddenness, wide frequency distribution, and rapid change, resulting in a complex time-frequency structure for the voice signal. Furthermore, as smart terminals and mobile devices continue to evolve towards lightweight, low-power, and high real-time performance, their voice processing links are typically constrained by edge computing power, storage space, and response speed. This necessitates maintaining voice quality while ensuring computational efficiency. Against the backdrop of multiple devices, multiple scenarios, and multiple noise types coexisting, the industry has gradually formed a development trend centered on high-quality voice output, real-time stable operation, and cross-scenario adaptability, promoting the evolution of voice noise reduction technology towards greater refinement, intelligence, and environmental understanding.

[0003] For example, invention patent CN109378013B discloses a speech denoising method and its training process. The method includes: S1, preprocessing the noisy speech signal by windowing, converting to the frequency domain, using a frequency domain algorithm for denoising, and reconstructing it into a time domain signal; S2, using speech endpoint detection (VAD) technology to detect endpoints in the preprocessed speech signal, determining effective speech segments based on signal characteristics, and pruning the entire speech; S3, converting the pruned speech signal to a predetermined format and slicing it into fixed-length speech segments as input to a deep denoising model; S4, outputting a clean speech signal through a deep neural network model. Simultaneously, the training process includes the acquisition, format conversion, and segment generation of noisy and clean speech samples, which are then used as input and output pairs for neural network training to obtain the deep denoising model. This invention ensures effective denoising processing of speech through multi-step preprocessing and deep model inference, guaranteeing clear output of the speech signal.

[0004] Traditional methods often rely on fixed frequency domain processing procedures or static deep neural network structures, making it difficult to provide timely and fine-grained responses to complex environments with varying noise levels and scenarios. In the context of sudden noise, high-frequency interference, and structural noise, the detailed features and key articulation structures of speech are easily over-suppressed, leading to a decrease in speech clarity and naturalness. Furthermore, existing deep noise reduction models typically lack dynamic perception and adjustment mechanisms for real-time operation, making it difficult to maintain stable performance in computationally limited and latency-sensitive devices. In addition, the relatively fixed preprocessing methods, feature slicing methods, and model input formats lack adaptability when facing multiple scene transitions, user differences, and changes in noise patterns, making it difficult to balance noise reduction effectiveness, structural fidelity, and real-time requirements. These shortcomings make it difficult for traditional speech noise reduction technologies to meet the requirements of higher quality and greater stability in complex, dynamic, and ever-changing application scenarios.

[0005] Therefore, in order to address the above problems, there is an urgent need for an adaptive speech denoising method and system for multi-noise scenarios. Summary of the Invention

[0006] Technical problems to be solved

[0007] To address the shortcomings of existing technologies, this invention provides an adaptive speech denoising method and system for multi-noise scenarios. It solves the problems that existing denoising methods are prone to generalization failure under non-stationary noise and environmental equipment changes in scenarios such as conference calls and noisy environments. They also fail to adequately consider global energy and high-frequency details and lack adaptive control due to time delay and computing power constraints, resulting in difficulty in consistently achieving the required clarity and audibility.

[0008] Technical solution

[0009] To achieve the above objectives, the present invention provides the following technical solution: an adaptive speech denoising method for multi-noise scenarios, comprising: S1, acquiring input speech data, obtaining operational control training support data and wavelet analysis configuration data; preprocessing the input speech data, operational control training support data, and wavelet analysis configuration data; S2, performing cross-scale structural decomposition and directional texture decomposition analysis on the input speech data based on the wavelet analysis mechanism, generating a multi-level scale feature set, and performing prior quality constraints in the feature injection stage; S3, uniformly rearranging and aligning the multi-level scale features in terms of spatial size, batch structure, channel dimension, and numerical domain, possessing the characteristics of features within the deep network. Figure 1The network adopts a fusionable structural form; S4, performs coding layer feature fusion analysis with joint modulation of structural details at the encoder and decoder ends of the nested network, and executes the structural modulation and detail coupling mechanism according to the coding layer feature fusion analysis results to construct a cross-level joint feature enhancement representation; S5, based on the input speech data and the operation control training support data, constructs a multi-objective loss constraint system and a hierarchical supervision structure to perform end-to-end parameter learning and optimization on the speech reconstruction network; S6, performs quality resource sensitive operation status evaluation on the input speech data, and performs coefficient adjustment, strategy switching and network path dynamic control according to the quality resource sensitive operation status evaluation results.

[0010] The second aspect of this invention provides an adaptive speech denoising system for multi-noise scenarios, including a data acquisition and preprocessing module for acquiring input speech data, obtaining operational control training support data and wavelet analysis configuration data; preprocessing the input speech data, operational control training support data, and wavelet analysis configuration data; a wavelet resolution feature decomposition module for performing cross-scale structural decomposition and directional texture decomposition analysis on the input speech data based on wavelet analysis mechanism, generating a multi-level scale feature set, and performing prior quality constraints in the feature injection stage; and a feature rearrangement and alignment module for uniformly rearranging and aligning the multi-level scale features in terms of spatial size, batch structure, channel dimension, and numerical domain, possessing the characteristics of features within the deep network. Figure 1 The system comprises a unified fusion structure; a neural network and downsampling layer feature injection module, used to perform co-modulation of structural details at the encoder and decoder ends of the nested network, and to execute structural modulation and detail coupling mechanisms based on the results of the co-modulation analysis to construct cross-layer joint feature enhancement representations; a training and hierarchical supervised optimization module, used to construct a multi-objective loss constraint system and hierarchical supervision structure based on input speech data and operational control training support data, and to perform end-to-end parameter learning and optimization on the speech reconstruction network; and an online evaluation and operational control module, used to perform quality resource-sensitive operational status evaluation on the input speech data, and to perform coefficient adjustment, strategy switching and dynamic network path control based on the evaluation results.

[0011] Beneficial effects

[0012] The present invention has the following beneficial effects:

[0013] (1) This invention, by constructing a multi-level, multi-scale speech structure perception framework, realizes the joint expression and dynamic enhancement from global structure to local details. It can robustly extract speech backbone information and detail texture in complex noise environments, thereby achieving the structure fidelity effect of speech in various environments and effectively solving the problem that speech details are easily masked by noise in the prior art.

[0014] (2) This invention utilizes multi-source time-frequency feature rearrangement, channel alignment and selective injection mechanism to enable the system to adaptively adjust feature fusion intensity in different noise scenarios, autonomously balance speech clarity and noise suppression, and thus achieve dynamic speech enhancement effect in multiple scenarios, effectively solving the problem of difficulty in handling different noise types in the prior art.

[0015] (3) The present invention is based on multi-scale speech structure features and dynamic feedback mechanism, which enables the system to select noise reduction path and feature fusion strategy according to different usage scenarios and device load, thereby realizing scenario-based and personalized intelligent speech enhancement effect, effectively solving the problems of fixed processing method and slow response in the prior art.

[0016] (4) The present invention adopts a multi-level supervised training and multi-objective joint optimization strategy to achieve a comprehensive balance between noise reduction, structure restoration and detail preservation, thereby improving the naturalness and intelligibility of speech and achieving a high consistency and high fidelity speech reconstruction effect, effectively solving the problems of harsh and distorted speech after noise reduction in the prior art.

[0017] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0018] Figure 1 This is a flowchart of an adaptive speech denoising method for multi-noise scenarios according to the present invention;

[0019] Figure 2 This is a structural diagram of an adaptive speech denoising system in a multi-noise scenario according to the present invention;

[0020] Figure 3 A schematic diagram of the local network structure for wavelet feature processing injection in this invention;

[0021] Figure 4 This is a comparison diagram of the feature modulation and enhancement effects of the present invention under different acoustic scenarios;

[0022] Figure 5 This is a flowchart of the voice processing monitoring and rollback decision-making process of the present invention;

[0023] Figure 6 This is a schematic diagram of the multi-level nested U-Net network structure of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figures 1-6 This invention provides a technical solution: an adaptive speech denoising method for multi-noise scenarios, comprising the following steps: S1, collecting input speech data, obtaining operation control training support data and wavelet analysis configuration data; preprocessing the input speech data, operation control training support data, and wavelet analysis configuration data; S2, performing cross-scale structural decomposition and directional texture decomposition analysis on the input speech data based on the wavelet analysis mechanism, generating a multi-level scale feature set, and performing prior quality constraints in the feature injection stage; S3, uniformly rearranging and aligning the multi-level scale features in terms of spatial size, batch structure, channel dimension, and numerical domain, possessing the characteristics of the deep network's internal features. Figure 1 The network adopts a fusionable structural form; S4, performs coding layer feature fusion analysis with joint modulation of structural details at the encoder and decoder ends of the nested network, and executes the structural modulation and detail coupling mechanism according to the coding layer feature fusion analysis results to construct a cross-level joint feature enhancement representation; S5, based on the input speech data and the operation control training support data, constructs a multi-objective loss constraint system and a hierarchical supervision structure to perform end-to-end parameter learning and optimization on the speech reconstruction network; S6, performs quality resource sensitive operation status evaluation on the input speech data, and performs coefficient adjustment, strategy switching and network path dynamic control according to the quality resource sensitive operation status evaluation results.

[0026] Specifically, the preprocessing process for input speech data, operation control training support data, and wavelet analysis configuration data is as follows: Input speech data is collected, which includes discrete speech signal sequences, speech sampling frequencies, and noisy speech samples, representing the basic information of the original speech in the time domain and sampling domain, and providing the original input basis for subsequent feature extraction.

[0027] Acquire operational control training support data, including: short-time Fourier transform window length, real-time GPU memory usage, reference GPU memory usage, target clean speech, actual GPU memory usage, floating-point operation volume, computing unit utilization rate, speech frame entry processing time, and output frame generation processing time. This data is used to characterize resource constraints, computational consumption, and structural configuration during the training and inference processes, enabling a unified optimization strategy under different computing power conditions. Training parameters are used for feature construction during the model learning phase, while real-time GPU memory usage, computing unit utilization rate, speech frame entry processing time, and output frame generation processing time are inference-phase operational monitoring metrics. These two metrics are used independently in the data domain to avoid confusion between training information and operational status.

[0028] Obtain wavelet analysis configuration data, which includes: time sampling index, wavelet base type, and maximum number of wavelet decomposition levels. This data is used to determine the time localization strategy and decomposition depth for multi-scale structural analysis, ensuring scale consistency between wavelet domain structural features and time-frequency domain features.

[0029] This paper addresses convolutional bounds at both ends of a discrete speech signal sequence using a symmetric extension method, maintaining energy continuity between smooth and abrupt regions and preventing energy attenuation or distortion of the convolution kernel during boundary calculations, thus improving the stability of the reconstructed original signal. Pre-emphasis processing and a spectrum tilt compensation algorithm enhance the high-frequency energy of the discrete speech signal sequence, improving the discriminability of speech boundaries, voiceless consonants, and detailed texture components. Framing and windowing algorithms structurally segment the input speech data according to its time-varying characteristics, reducing analysis distortion caused by cross-frame leakage. Finally, short-time Fourier transform and Mel filter bank transform convert the discrete speech signal sequence into a two-dimensional spectrogram, extracting time-frequency energy suitable for modeling. Distribution characteristics; the amplitude scale and distribution differences are normalized by the mean-variance standardization algorithm; the scale and unit are aligned and unified for the input speech data and the operation control training support data by the standardization and normalization algorithms to avoid gradient imbalance caused by inconsistent feature dimensions and improve the fusion quality of multi-source input data in the network. Among them, the frame length of the framing and windowing operations is 20-32ms, the frame shift is 8-12ms, and the window type is Hamming window; the number of FFT points of the short time Fourier transform is configured in the range of 256-1024 points; the number of bands of the Mel filter bank is between 32 and 64 bands according to the computing power and target accuracy, which can be adaptively adjusted according to the noise scene, model size and real-time requirements.

[0030] In this implementation scheme, the input speech data, operation control training support data, and wavelet analysis configuration data achieve a unified structural expression and dimensional consistency in the time, frequency, and scale domains. The boundary continuity of the original speech signal is maintained, the distinguishability of high-frequency energy components is enhanced, the time-frequency structure of time-varying speech segments is accurately segmented, the time-frequency energy distribution of the two-dimensional spectrogram is effectively extracted, the amplitude and scale differences of multi-source data are eliminated, the data domain boundaries of training parameters and inference monitoring quantities are clearly defined, and multidimensional features enter the network under a unified statistical distribution. This provides a stable, standardized, and fusionable data foundation for subsequent multi-scale wavelet decomposition, feature mapping, structure injection, and end-to-end modeling. This enables the entire speech reconstruction network to have consistent data input quality, controllable feature scale, and reliable model convergence conditions under complex noise environments and multi-computing power conditions.

[0031] Specifically, the process of performing cross-scale structural decomposition and directional texture decomposition analysis to generate multi-level scale feature sets is as follows: This is used to perform multi-scale feature decomposition on the input discrete speech signal, supplementing the shortcomings of deep networks in cross-scale representation through multi-resolution analysis, forming a parallel feature source that simultaneously possesses the overall structural skeleton and local detailed features; the input discrete speech signal sequence is sequentially processed through wavelet low-pass analysis filtering sequences and wavelet high-pass analysis filtering sequences to separate low-frequency smoothing components from high-frequency detailed components, forming the first wavelet decomposition; using the time sampling index as a sliding reference, the input discrete speech signal sequence is compared with the wavelet low-pass analysis filtering sequence and the wavelet high-pass analysis filtering sequence respectively. The two convolutional outputs are summed to obtain the complete low-frequency and high-frequency convolutional outputs. Then, both convolutional results are simultaneously downsampled by a factor of two, retaining the output values ​​at even-numbered sampling positions to obtain the first-level smoothing coefficient sequence and the first-level detail coefficient sequence, respectively. The first-level smoothing coefficient mainly characterizes the overall energy distribution, formant structure, and stable features of the vowel region in speech. The first-level detail coefficient mainly describes the characteristics of voiceless consonants, plosives, transient changes in speech, and local rapid changes in high-frequency noise. Since wavelet decomposition uses a factor of two downsampling, only even-numbered sampling points in the filtered output are retained. The lengths of the first-level smoothing coefficient and the first-level detail coefficient are approximately half the length of the input speech signal sequence. The specific calculation formulas are as follows: ;

[0032] ;

[0033] Wavelet decomposition is performed on the first-level smoothing coefficients to form a recursive multi-level processing mechanism. The smoothing coefficients generated in the previous layer are used as inputs for the next layer, and low-pass analysis, high-pass analysis, and downsampling are continuously performed. By convolving and summing with the wavelet low-pass analysis filter sequence and performing double downsampling, a coarser-scale low-frequency information is extracted to form the k-th level smoothing coefficient. This is then convolved and summed with the wavelet high-pass analysis filter sequence and performed double downsampling to extract high-frequency variation features at the corresponding scale, forming the k-th level detail coefficient. This process continues recursively until the maximum decomposition level is reached. As the decomposition level deepens, the smoothing coefficients gradually present a coarser-scale global structural information, while the detail coefficients retain high-frequency texture, speech contour transitions, local abrupt changes, and noise interference features at multiple scales, achieving a comprehensive expression from macro to micro and from stationary to non-stationary. The specific calculation formula is as follows:

[0034] ;

[0035] ;

[0036] In the formula, Let represent the input discrete speech signal, and represent the sampled noisy speech time-domain sequence, i.e., the... One sampling point; This represents the time sampling index, the domain of the discrete signal, and the position of the signal on the discrete time axis. This represents a wavelet low-pass analysis filter sequence, from which low-frequency trends and smoothing components are extracted to generate smoothing coefficients; The wavelet high-pass analysis filter sequence is used to extract high-frequency details and transient changes to generate detail coefficients. The wavelet basis two-scale recursive equation determined by the wavelet basis type determines the wavelet low-pass analysis filter sequence and the wavelet high-pass analysis filter sequence. The low-pass analysis filter sequence is obtained by discretization through the scaling function. The high-pass analysis filter sequence is derived from the low-pass analysis filter sequence through orthogonal mirror transformation. Indicates the index of the convolution summation. This represents the summation index and the convolution sliding window variable. This represents the first-level smoothing coefficient, reflecting the main energy, formants, and vowel smoothing structure of the speech, and is output by wavelet decomposition. Represents first-level detail coefficients; characterizes voiceless consonants, transients, and noise components in speech, output by wavelet decomposition; Indicates the first The smoothing coefficients are obtained from multi-level recursive wavelet decomposition and represent a coarser scale after further downsampling based on the previous low-frequency layer. Indicates the first Level detail coefficients are obtained through multi-level recursive wavelet decomposition, capturing the first level. High-frequency texture and structural abrupt change features This represents the input of the smoothing coefficients from the previous level, obtained from the wavelet decomposition recursive chain, the th... The low-frequency output of the first layer is used as the input of the next layer. This indicates the maximum number of wavelet decomposition layers, used to control the multi-scale information granularity and frequency resolution.

[0037] The input 2D spectrogram is processed in the first wavelet decomposition stage, yielding four types of subband coefficients: coarse-scale smooth subband, horizontal detail subband, vertical detail subband, and diagonal detail subband, each corresponding to half the size of the input spectrogram. As the decomposition level increases, the spatial size of the k-th level subband decreases sequentially to half, a quarter, and an eighth of the original spectrogram, and so on. After the k-th convolution and pooling operation, the 2D spectrogram forms multi-channel time-frequency features, representing the k-th layer feature map. The coarse-scale smooth subband handles the global structural representation, while the directional detail subband handles local high-frequency details and speech patterns. For the task of characterizing speech texture changes and noise response, wavelet subbands of corresponding one-dimensional discrete sequence dimensions and two-dimensional spectrogram dimensions are selected for fusion based on the dimension of the injected nodes and task requirements. One-dimensional wavelet decomposition focuses on characterizing the temporal scale variation trend of the speech signal, while two-dimensional wavelet decomposition is used to characterize the directional texture features of the spectrogram. Both serve as optional sources of multi-scale structure. During the feature injection stage, wavelet feature branches matching the feature map dimension of the corresponding network layer are selected: two-dimensional feature maps are preferentially matched with two-dimensional wavelet subbands, while one-dimensional feature sequences are used for running quality prediction, gating adjustment, or additional temporal structure enhancement. By adhering to the principle of dimensional consistency, the two branches can be used collaboratively without directly merging wavelet features of different dimensions.

[0038] In this implementation scheme, the cross-scale structural decomposition and directional texture decomposition analysis in this step constructs a multi-level scale feature set covering both temporal and time-frequency scales. This enables the input speech to obtain a systematic description in the dimensions of coarse-scale structural skeleton, formant morphology, energy trend, speech contour transition, transient changes, directional texture, and high-frequency noise response. It allows the smoothing coefficient and detail coefficient to form a scale expression from global to local and from stationary to non-stationary in a multi-level recursive process. It enables the coarse-scale smoothing sub-band and directional detail sub-band to establish a directional expression of global structure and local texture in two-dimensional space. It enables one-dimensional wavelet decomposition and two-dimensional wavelet decomposition to undertake the feature modeling of temporal structure and time-frequency texture, respectively. It enables the multi-scale sub-bands to be seamlessly injected into the deep network feature stream after matching the feature map dimension of the network layer. This gives the network a more complete structural modeling capability, detail preservation capability, noise identification capability, and cross-scene robustness under multi-scale, multi-directional, and multi-frequency information coverage.

[0039] Specifically, the prior quality constraint process in the feature injection stage is as follows: After multi-level wavelet decomposition, the speech signal-to-noise ratio (SNR) is obtained by calculating the ratio of speech energy to noise energy for the discrete speech signal sequence using a silence segment statistical algorithm; the amplitude spectrum is calculated by performing a short-time Fourier transform on the discrete speech signal sequence, and the ratio of the geometric mean to the arithmetic mean of the spectrum is obtained to obtain the speech spectral flatness; the SNR and speech spectral flatness of each frame are evaluated for quality, and the SNR and SNR threshold, and the speech spectral flatness and speech spectral flatness threshold are compared in real time.

[0040] When the speech signal-to-noise ratio is less than the speech signal-to-noise ratio threshold or the speech spectrum flatness is greater than the speech spectrum flatness threshold, the noise-dominant feature is obvious. The corresponding decomposed detail texture subbands, including horizontal detail subbands, vertical detail subbands and diagonal detail subbands, are weighted candidate labels for subsequent feature fusion stage weight suppression, gating filtering and bandwidth compression processing, so as to avoid the noise being excessively amplified in the deep network and improve stability and noise reduction generalization ability.

[0041] When the speech signal-to-noise ratio (SNR) is greater than or equal to the SNR threshold and the speech spectrum flatness is less than or equal to the speech spectrum flatness threshold, no special processing is performed. The output is the coarse-scale smooth sub-band and directional detail texture sub-band for each decomposition level, including the time-frequency statistical feature map, energy distribution map, bandwidth ratio, and scale level index information of the k-th level coarse-scale smooth sub-band, k-th level horizontal detail sub-band, k-th level vertical detail sub-band, and k-th level diagonal detail sub-band. At the same time, the wavelet low-pass analysis filter sequence and wavelet high-pass analysis filter sequence used, wavelet base type, maximum decomposition level, and downsampling strategy are archived and created and archived in the wavelet feature decomposition database to provide a basis and auditability support for subsequent feature injection, model training traceability, experimental reproduction, and online operation monitoring.

[0042] In this implementation scheme, prior quality constraints enable the coarse-scale smooth subband and directional detail texture subband obtained from multi-level wavelet decomposition to have reliable noise sensitivity screening capabilities before entering the deep network. This allows the noise-dominant directional detail subband to receive clear weighted candidate labels before feature injection. It also enables the detail texture components in high-noise scenes to receive controllable weight suppression and gating filtering in the subsequent fusion process. Furthermore, it prevents the accumulation and amplification of invalid high-frequency noise in the network during the feature injection stage. This ensures that the multi-scale features received by the deep network maintain higher speech dominance and structural stability. It also ensures that the coarse-scale smooth subband and directional detail texture subband in low-noise scenes can be completely preserved and entered into the network. Finally, it ensures that speech structure information, energy distribution information, and scale hierarchy information are effectively transmitted in the reconstruction path supporting the backbone network. This makes the feature injection process auditable, traceable, and repeatable, and enables subsequent training and online operation to obtain more robust noise suppression capabilities, structure preservation capabilities, and cross-scene generalization capabilities.

[0043] Specifically, it possesses characteristics similar to those of deep networks. Figure 1 The specific process for achieving the desired fusionable structural form is as follows: Input the set of scale subband coefficients from the wavelet decomposition outputs of layers 1 to N, including coarse-scale smooth subbands, horizontal detail subbands, vertical detail subbands, and diagonal detail subbands. Before injecting the multi-scale wavelet decomposition features into the deep network, the scale subbands at different levels are rearranged and aligned to ensure consistency with the multi-level nested U-shaped neural network in terms of spatial structure, channel representation, data type, and numerical distribution. Using bilinear interpolation, nearest-neighbor interpolation, and average pooling methods, the coarse-scale smooth subbands and directional detail subbands are adjusted to align with the k-th layer features at the encoding end. Figure 1 To achieve pixel-level fusion, spatial size adjustment is performed on the coarse-scale smooth subband and the directional detail subband, ensuring that the spatial size of the scale subband matches the corresponding layer features of the target network's encoding end. Figure 1 To avoid calculation errors caused by inconsistent tensor dimensions when running in batch mode, the scale subbands are expanded in batch dimension under batch processing conditions so that their number matches the number of samples in the current input batch.

[0044] The wavelet subband has a single-channel structure, and the feature map at the encoding end contains multi-channel semantics. Through one-dimensional convolutional mapping, the wavelet features are mapped to the target number of channels, ensuring that its semantic capacity, feature contrast, and network learning ability are consistent. Channel mapping is performed on the scale subband, converting the single-channel subband into a multi-channel feature with the same number of channels as the corresponding layer feature map at the encoding end of the target network. When the scale subband is marked as a candidate for weight reduction, its channel weight and participation ratio are reduced in the channel mapping. During the channel mapping process, by setting linear and exponential decay weight functions, its channel weights are adaptively reduced in the range of 0.1 to 0.6 to limit the participation ratio and avoid feature bias during multi-channel fusion, while ensuring the stability and trainability of the overall feature representation. To further ensure training stability and operator compatibility, all multi-channel features are unified to the same floating-point data type. According to the network internal specifications, a numerical range calibration strategy consistent with Z-score and group normalization style is adopted to obtain multi-scale features, so that the wavelet domain features can be seamlessly integrated into the main network feature stream.

[0045] In the feature selection and fusion stage, not all subbands are injected indiscriminately. Adaptive selection is performed based on multi-scale speech expression patterns and task characteristics: in the deep encoding stage, coarse-scale smooth subbands are preferentially retained to express overall energy trends, prosodic structure, and speech skeleton; while in the shallow encoding and decoding stages, horizontal, vertical, and diagonal detail subbands are selectively introduced when gating conditions are met to recover transient features related to plosives, voiceless consonants, edge transitions, and speech intelligibility. Multi-scale features are input to the corresponding layers of the target network via channel-dimensional concatenation, providing the network with a fusion feature source that combines global structural information with high-frequency detail expression. Simultaneously, the alignment mapping operator configuration, the number of target channels, the feature selection strategy, and the threshold set are recorded to provide a basis for subsequent training analysis, deployment optimization, visualization checks, and performance tracking.

[0046] like Figure 3 The diagram shows a schematic of the local network structure for wavelet feature injection, including a first-level wavelet decomposition unit, a feature map resizing unit, convolution and pooling units, deconvolution units, and a bilinear interpolation fusion unit. First, the time-frequency features of the input speech are decomposed using a first-level wavelet to obtain multiple detail subbands and low-frequency subbands, which are shown in the diagram as the original subband features. These subband features are processed by the feature map resizing module to ensure their spatial resolution matches the feature map of the corresponding layer in the backbone network, thus meeting the size requirements for subsequent fusion. At the encoding end, a progressively downsampled feature extraction path is formed through consecutive 3×3 convolutions, ReLU activation, and 2×2 max pooling operations. After resizing, the wavelet subband features at the corresponding positions are introduced into the decoding path through a combination of 2×2 upsampling and 1×1 convolutions. The 1×1 convolution is used for channel mapping and feature compression, and bilinear interpolation is used to align the spatial scale differences between the backbone feature map and the injected wavelet features. The right branch in the figure illustrates the feature fusion process after first-level wavelet decomposition injection. Through upsampling, 1×1 convolution, and interpolation operations, the decoder features can simultaneously utilize convolutional features from the encoder and structural information from the wavelet domain. This local structure reflects the fusion method of the first-level wavelet subband and the U-Net backbone network at the corresponding scale, demonstrating the injection position of the wavelet decomposition results in the decoding path, the channel mapping method, and its collaborative relationship with the convolutional and upsampling layers, providing a clear structural illustration for the realization of multi-scale feature fusion.

[0047] In this implementation, this step ensures that the coarse-scale smoothing subband and the directional detail texture subband obtained from multi-level wavelet decomposition form a fusionable structure that is completely consistent with the feature maps inside the backbone network before being injected into the deep network. This ensures that the scale subband maintains strict alignment with the feature maps at the encoder and decoder ends in terms of spatial size, number of channels, batch dimension, data type, and numerical distribution. It also enables the single-channel wavelet subband to acquire expressive power consistent with multi-channel semantics through channel mapping, and allows the directional detail texture subband, marked as a weight reduction candidate, to achieve controllable linear and exponential decay during the channel mapping stage, thus enhancing the detail texture under high noise conditions. The automatic reduction of component participation in the fusion process enables the deep feature extraction stage to focus on utilizing the global energy skeleton provided by the coarse-scale smooth subband, and the shallow reconstruction stage to fully recover transient details and boundary structures when the gating conditions are met. This allows feature injection to perform stable pixel-level alignment and channel-level fusion on a scale-by-scale and level-by-level basis, and allows wavelet domain structures to be naturally incorporated into multi-level nested U-shaped networks. This significantly improves the network's ability to restore structure, preserve details, and express speech intelligibility, and makes the entire fusion process traceable in terms of structural configuration, inspectable injection nodes, and reproducible experimental paths.

[0048] Specifically, the process of implementing the coding layer feature fusion analysis for joint modulation of structural details is as follows: Based on a multi-level nested U-shaped network structure, multi-scale features are extracted from the convolutional layer at the coding end. After the k-th pooling operation, the original network feature map of the k-th layer is obtained. The wavelet scale sub-bands, after feature rearrangement and alignment, are injected into the k-th layer of the target network according to their corresponding levels to enhance the cross-scale speech feature representation capability. Based on the original network feature map of the k-th layer at the coding end, a feature modulation mapping function of the k-th layer is constructed to jointly model the structure and details of the features. The modulation features are constrained and normalized by a nonlinear gated mapping function. The obtained modulation weights are multiplied element-wise with the original network feature map to obtain the modulation features. The modulation features and the original network features are residually superimposed to generate the enhanced feature map of the k-th layer. The specific calculation formula is as follows:

[0049]

[0050] In the formula, This represents the enhanced feature map of the k-th layer, which will be used as the input for subsequent network encoding and decoding; This represents the original network feature map of the k-th layer at the encoding end, reflecting the semantic and structural information extracted by the network at this layer; This represents the feature modulation mapping function of the k-th layer, used for joint structural and detail modeling of the input feature map. The layer input feature map first extracts spatial structure response and neighborhood change information through local convolution operation, and then performs feature transformation in the channel dimension to characterize the response intensity distribution of different semantic channels. If necessary, nonlinear activation is combined to re-encode structural trends and detail changes, thereby generating a modulation feature map that contains both overall structural information and local detail differences, which is used for subsequent gating and enhancement calculations to extract modulation features that reflect overall structural trends and local detail changes. This represents a nonlinear gated mapping function that compresses and maps the output of the feature modulation mapping function to a finite interval between 0 and 1, thereby achieving adaptive adjustment of the response intensity of features at different positions and channels to suppress invalid or noisy features and highlight semantically valuable details. This represents element-wise multiplication, used to apply modulation weights to the original feature map to dynamically enhance or attenuate the feature amplitude. By superimposing the modulated features with the original features using residuals, an adaptive enhancement mechanism based on structural and detail information is introduced while maintaining the stability of the original feature representation. This enables the network to balance global structural consistency and local detail representation capabilities during multi-scale feature fusion.

[0051] In this embodiment, Table 1 presents a statistical table of feature modulation and enhancement effects under different acoustic scenarios. It quantifies the modulation response and enhancement results of coding layer features under various typical acoustic conditions, characterizing the impact of modulation mapping functions and nonlinear gating mechanisms on network feature enhancement behavior in different speech and noise environments. The table provides statistical results such as the original feature intensity, mean modulation mapping response, mean gating coefficient, feature enhancement contribution, and enhanced feature intensity, reflecting the overall trend and relative differences of feature enhancement mechanisms under different acoustic conditions. Scenario 1, "Quiet speech + moderate detail," had an original feature strength of 0.80, a mean modulation mapping response of 0.42, a mean gating coefficient of 0.60, and a corresponding feature enhancement contribution of 0.48. The enhanced feature strength reached 1.28, indicating that under conditions of low background interference and moderate detail, the modulation mechanism can stably improve feature representation capabilities. Scenario 2, "Normal speech + significant detail," also had an original feature strength of 0.80, but the mean modulation mapping response increased to 0.55, the mean gating coefficient was 0.63, and the feature enhancement contribution was 0.50. The enhanced feature strength reached 1.30, showing that under conditions of richer detail information in speech, structure-detail joint modulation can further amplify semantically valuable feature responses. Scenario 3, "Speech..." In scenario 4, "weak speech + light noise", the original feature strength remained at 0.80, but the mean modulation mapping response increased significantly to 0.95, the mean gating coefficient increased to 0.72, the feature enhancement contribution was 0.58, and the enhanced feature strength reached 1.38. This reflects that in scenarios with significant high-frequency noise, the modulation mapping function responds more strongly to feature changes, while the gating mechanism constrains the enhancement amplitude. Scenario 4, "weak speech + light noise", saw the original feature strength decrease to 0.40, the mean modulation mapping response was 0.30, the mean gating coefficient was 0.57, the feature enhancement contribution was 0.23, and the enhanced feature strength was 0.63. The overall enhancement amplitude was significantly lower than in other scenarios, demonstrating the suppression characteristics of the modulation mechanism on feature amplification under weak speech energy conditions.

[0052] Table 1. Statistical table of feature modulation and enhancement effects in different acoustic scenarios.

[0053] Scene Number Scene Description Original feature strength Modulation-mapped response mean Mean of gating coefficient Feature enhancement contribution Enhanced feature strength Scene 1 Quiet voice + moderate detail 0.80 0.42 0.60 0.48 1.28 Scene 2 Normal speech + obvious details 0.80 0.55 0.63 0.50 1.30 Scene 3 Voice + strong high-frequency noise 0.80 0.95 0.72 0.58 1.38 Scene 4 Weak speech + light noise 0.40 0.30 0.57 0.23 0.63

[0054] like Figure 4The figure shows a comparison of feature modulation and enhancement effects under different acoustic scenarios. Table 1 shows significant differences in the original feature intensity, gating coefficient, and enhanced feature intensity across different acoustic scenarios. Scenario 3, "speech + strong high-frequency noise," has the highest feature enhancement contribution, with an enhanced feature intensity of 1.38. This indicates that under complex acoustic conditions with significant high-frequency noise, the modulation mapping and gating mechanism respond most fully to feature enhancement. Scenario 1, "quiet speech + moderate detail," and Scenario 2, "normal speech + obvious detail," have enhanced feature intensities of 1.28 and 1.30, respectively, remaining at a relatively high and stable level. This suggests that feature modulation can continuously improve feature representation under normal speech conditions. Scenario 4, "weak speech + light noise," has an original feature intensity of only 0.40 and an enhanced feature intensity of 0.63, with a relatively low overall enhancement magnitude. This reflects that the gating mechanism has a suppressive effect on feature amplification under weak speech energy scenarios. Overall, the figure intuitively reflects the synergistic changes of feature modulation mapping, gating coefficients, and enhanced feature intensity under different acoustic scenarios, and can be used to illustrate the adaptive adjustment effect of the feature modulation and gating enhancement mechanism in multi-noise environments.

[0055] This implementation scheme achieves collaborative modeling and adaptive enhancement of structural and detail information during the multi-level nested network encoding stage. By introducing rearranged and aligned wavelet-scale subbands into the corresponding layers at the encoding end, the network gains additional cross-scale structural priors on top of the original convolutional feature extraction. Based on the original feature representation of the current layer at the encoding end, a joint feature modulation mechanism of structure and detail is constructed to uniformly characterize the overall energy distribution, spatial structure trend, and local detail changes in the features. A nonlinear gating method is used to impose bounded constraints on the modulation results, ensuring the stability and controllability of the enhancement process. Furthermore, the modulated features are residually superimposed with the original features, enabling the network to effectively highlight effective feature components related to speech structure and clarity without disrupting the continuity of the original semantic expression, while suppressing invalid responses and noise amplification. In this way, the encoding layer can simultaneously maintain the consistency of the overall speech structure and the ability to express local details at different scales, providing a more stable, clear, and discriminative feature foundation for feature reconstruction in the subsequent decoding stage, thereby improving the robustness of the speech enhancement process in complex noise environments and the overall reconstruction quality.

[0056] Specifically, based on the feature fusion analysis results of the coding layer, the structural modulation and detail coupling mechanism is executed to construct a cross-layer joint feature enhancement representation. The specific process is as follows: In the network structure, the deep coding stage uses the k-th level coarse-scale smooth subband and its convolutional response as the main external scale supplementary features, focusing on robust modeling of the global speech energy framework, structural morphology, and prosodic direction. The shallow coding and decoding stages are based on the detail energy map composed of the k-th level horizontal detail subband norm, vertical detail subband norm, and diagonal detail subband norm, combined with the ratio regularization modulation mechanism, to perform directional enhancement of local transition boundaries, plosive details, unvoiced consonant textures, and transient structures, achieving bottom-up fine-grained reconstruction. In the decoding stage, under the action of the cross-layer skip connection mechanism, the k-th level enhanced feature map after wavelet scale injection and the corresponding shallow features are jointly input into the decoding module, forming a multi-scale recovery path that gradually feeds back from coarse-scale structure to directional detail texture.

[0057] like Figure 5 The diagram shows the speech processing monitoring and rollback decision-making process. Computational power consumption is calculated from the actual memory usage, floating-point operations, and computing unit utilization during the network forward inference process; real-time runtime latency is measured from the end-to-end processing time between the input speech frame and the generation of the corresponding output frame; and a latency and computational power constraint control mechanism is implemented by comparing real-time runtime latency with real-time latency thresholds, and comparing computational power consumption with computational power consumption thresholds in real time.

[0058] When the end-to-end real-time runtime latency exceeds the real-time latency threshold, a fallback strategy is implemented, which includes high-frequency subband channel pruning, subband bandwidth compression, and reducing the number of wavelet decomposition layers. This reduces the computational load while ensuring basic speech intelligibility and main structure reconstruction. When the computational power consumption exceeds the computational power consumption threshold, the maximum number of wavelet decomposition layers is reduced, the number of deep injections is decreased, and a lightweight decoding head structure is enabled. This works in conjunction with the multi-scale injection strategy to meet the real-time performance and computational power constraints of the terminal device.

[0059] When the end-to-end real-time runtime latency is less than or equal to the real-time latency threshold and the computing power consumption is less than or equal to the computing power consumption threshold, no special processing is performed; the output includes the training discrete speech signal sequence, network feature map, hierarchical sub-band mapping table, configuration of the number of channels in each layer, bandwidth allocation record, and rollback strategy execution log.

[0060] like Figure 6The diagram shows a multi-level nested U-Net network structure, which consists of three parts: an encoder, a decoder, and cross-layer skip connections. The overall structure employs a symmetrical network framework with top-down and bottom-up progressive downsampling. Each layer in the encoder consists of two convolution operations and one pooling operation. The convolution operation uses a 3×3 convolution kernel and a linear rectified activation function, while the pooling operation uses 2×2 max pooling to reduce the dimensionality of the feature map. Each layer in the decoder uses a 2×2 upsampling convolution to enlarge the feature map, and then splices and fuses it with the feature map of the corresponding layer in the encoder after cropping and alignment. Further feature extraction is then performed using a 3×3 convolution operation. Injection nodes for first-level, second-level, and third-level wavelet decomposition are introduced between each downsampling layer. The wavelet sub-band features of the corresponding resolution are mapped to a spatial scale consistent with the network feature map size and then jointly input with the feature map of that layer to achieve the fusion of structural information and detailed texture information at different scales at the corresponding depth. The spatial dimensions of wavelet decomposition injection points and network downsampling layers are marked with annotations in the figure, acting on spatial resolutions of approximately 284², 140², and 68², respectively, thus enhancing and supplementing coarse-scale structure and fine-scale details at different stages of the encoding process. The figure uses different colors and line types to distinguish the functions of convolution, upsampling, pooling, and 1×1 convolution operations; dashed lines represent cross-layer jump connections, enabling the decoder to combine features from the same layer of the encoder to achieve progressive recovery and detail reconstruction of multi-scale information. The annotations of the feature map dimensions at each scale visually reflect the changes in spatial resolution caused by each convolution and sampling operation.

[0061] This implementation scheme enables the deep coding stage to obtain a stable global energy framework and main morphological structure, the shallow coding and decoding stages to obtain directional high-frequency detail enhancement capabilities, the multi-scale structure to form a coarse-to-fine hierarchical expression in the downsampling path, the directional detail texture to complete step-by-step backfeeding in the upsampling path, the cross-layer jump connections to maintain spatial consistency when fusing structural and detail features, the enhancement features to maintain a stable distribution and controllable amplitude at each scale level, the overall speech contour, prosodic trend, and formant morphology to be stably constructed, the plosives, voiceless consonants, transition boundaries, and transient textures to be locally enhanced, and the detail energy-dominated and noise-dominated regions to be separated. Adaptive separation is achieved, enabling the coupling of structural domain modulation and detail domain to form a constrained enhanced representation. This allows the multi-scale recovery path to establish a balance between global consistency and local clarity, enables controllable recovery of high-frequency details during the end-to-end upsampling stage, maintains scale traceability in feature fusion at different resolutions, ensures stable convergence of the network in noisy environments, weak speech scenarios, and complex interference scenarios, and allows the multi-level nested structure to exhibit continuous detail compensation capabilities during the decoding stage. This ensures robust speech reconstruction performance under real-time and computational constraints, and the final output feature representation simultaneously possesses high-fidelity speech structure, clear boundary details, and good noise suppression.

[0062] Specifically, the process of constructing a multi-objective loss constraint system and a hierarchical supervision structure for end-to-end parameter learning and optimization of the speech reconstruction network is as follows: Noisy speech samples and corresponding clean target speech samples form paired training data. Training discrete speech signal sequences are used as model training input. A speech reconstruction model is constructed through end-to-end supervised learning. Training samples cover telephone calls, conference conversations, traffic noise, crowd noise, mechanical operation sounds, keyboard typing sounds, reverberation scenarios, and multi-source composite interference sound fields. This is used to establish the mapping relationship between noisy speech and target clean speech, ensuring the generalization ability of the speech reconstruction model in time-varying, strong interference, and non-stationary sound fields. The training objective is constrained by a weighted vector composed of speech signal-to-noise ratio reconstruction loss weights, logarithmic amplitude spectrum reconstruction loss weights, time-frequency structure perception loss weights, and high-frequency band edge protection loss weights, forming a multi-objective joint optimization framework to balance noise suppression capability, speech structure fidelity, and detail preservation capability. The training loss function consists of the following parts, used to guide backpropagation optimization during the end-to-end supervised learning process: speech signal-to-noise ratio... Consistency reconstruction loss is used to constrain the consistency of the output speech of the speech reconstruction model with the reference speech in the temporal energy structure; logarithmic amplitude spectrum reconstruction loss is used to improve the realism and naturalness of the spectrogram amplitude; hierarchical short-time Fourier transform structure perception loss improves the speech reconstruction model's ability to express transient structure, harmonic details, and acoustic texture through different window lengths and frequency resolutions; high-frequency band edge protection loss is used to suppress excessive smoothing and attenuation of high-frequency structures and maintain the clarity of voiceless consonants, plosives, and boundary information. In this embodiment, the speech quality metrics involved in the multi-objective loss constraint system adopt an objective calculation index system to ensure the quantifiability of training objectives and the consistency of evaluation standards. Speech perception quality can be measured by perception evaluation index, speech intelligibility is characterized by short-time objective intelligibility index, speech structure consistency is measured by signal-to-noise ratio consistency index, and speech naturalness is measured by subjective mean opinion score prediction estimation model. The indexes can be calculated uniformly in conjunction with the validation set to monitor the training effect, dynamic trend, and possible structural degradation risk of the speech reconstruction model. The various loss weights in the multi-objective loss constraint system are obtained from the analysis of the training objectives and noise scene requirements within a certain range. Among them, the loss weight for speech signal-to-noise ratio consistency reconstruction can be 0.2 to 0.5, the loss weight for logarithmic amplitude spectrum reconstruction is 0.1 to 0.4, the loss weight for hierarchical short-time Fourier transform structure perception is 0.2 to 0.5, and the loss weight for high-frequency band edge protection is 0.05 to 0.3. The weights are adjusted according to the plosive sounds, friction sounds, and fast transition sound fields with prominent high-frequency energy to balance noise suppression capability, structure fidelity, and high-frequency detail preservation capability.

[0063] The supervision method employs a hierarchical time-frequency supervision mechanism, setting supervision heads at different depths in the encoding, decoding, and final output layers to constrain the consistency of cross-layer feature representations, forming a progressive convergence and structured reconstruction process. During training, data enhancement strategies such as noise enhancement, signal-to-noise ratio perturbation, reverberation synthesis, echo superposition, and background replacement are implemented to maintain stable generalization performance under time-varying, multi-source, and non-stationary noise conditions. The learning rate scheduling adopts a preheating incremental cosine annealing joint strategy, combined with exponential moving average, gradient pruning, and weight stabilization control, to accelerate training convergence and suppress gradient oscillations. Training quality evaluation focuses on the comprehensive evaluation index of speech quality on the validation set, including speech intelligibility, speech clarity, signal-to-noise ratio improvement, and dynamic perception score, to monitor the performance trend, training stability, and overfitting risk of the speech reconstruction model. When the quality of the validation set continuously falls below the quality backoff threshold, an early stop or loss weight freeze strategy is executed. In scenarios with strong real-time constraints or high band edge energy, the high-frequency band edge protection loss weights are increased to stabilize the expressive power of high-frequency features. Finally, the best checkpoint file, exponential moving average model version, loss weight configuration set, and data augmentation configuration are output after training. The evaluation results of speech perception quality, speech intelligibility, structural consistency, and speech naturalness are output. A model training database is created and archived in the model training database.

[0064] In this implementation scheme, the speech reconstruction network is trained end-to-end to form a quantifiable speech quality target system, enabling simultaneous constraints on speech perception quality, speech intelligibility, speech structure consistency, and speech naturalness. This allows the network to establish a stable convergence path for mapping noisy speech to clean target speech, enhances generalization ability in multi-source noise scenarios, and achieves a balance between different loss terms in speech structure, amplitude spectrum details, time-frequency texture, and high-frequency band edge information. This results in synergistic improvements in noise suppression, structure preservation, detail enhancement, and high-frequency protection. Furthermore, the multi-scale supervisor provides cross-layer consistency constraints at different depths, enabling... The feature representation maintains stable gradient propagation during stepwise convergence, enabling the training process to achieve strong robustness under various enhancement strategies. It keeps the reconstruction accuracy stable under dynamic noise, reverberation environment, and complex interference conditions. The learning rate scheduling, gradient pruning, and weight balancing mechanisms ensure the controllability and convergence speed of the training process. The quality backtracking detection and early stopping strategies prevent training degradation and overfitting. It achieves an overall improvement in speech clarity, intelligibility, energy structure, detail texture, and naturalness of sound. All parameters, configurations, and versions of the training process are traceable, and a stable structural representation foundation and detail recovery capability are obtained in the inference stage.

[0065] Specifically, the process of performing a quality resource-sensitive operational status assessment on the input speech data, and then adjusting coefficients, switching strategies, and dynamically controlling network paths based on the assessment results, is as follows: Speech perception quality scoring, speech intelligibility estimation, structural consistency measurement, and speech naturalness assessment are performed on the real-time denoised speech. After standardization, these are fused according to quality weights to obtain a comprehensive speech quality evaluation index, which characterizes the current denoising effect. Simultaneously, end-to-end real-time indicators and computing power consumption are collected to reflect energy consumption, fluency, and equipment capacity.

[0066] Inputting a comprehensive speech quality evaluation index, a runtime quality prediction model is created. This model uses the time mean of the comprehensive speech quality evaluation index as a base term, and incorporates a resource constraint suppression term consisting of end-to-end real-time runtime latency and computing power consumption into the denominator. The high-frequency detail complexity metric and sensitivity adjustment coefficient of the k-th layer are used to correct the predicted runtime quality. Specifically, the time mean of the comprehensive speech quality evaluation index is used as the basic quantity reflecting the stability of the current speech processing effect, which is used to characterize the consistency of speech denoising and the level of perceived quality within a certain time window. At the same time, a resource constraint mechanism related to the real-time running state is introduced into the runtime quality prediction model. By jointly modeling the latency pressure and the computing power consumption pressure, a suppression term for the speech quality prediction result is constructed to reflect the constraint influence of the real-time running environment on the model output quality. A detail complexity metric related to the intensity of high-frequency features of the k-th layer is introduced to characterize the activity level of the current network in the high-frequency detail enhancement stage, and the prediction quality is corrected by energy-sensitive adjustment. Thus, a runtime quality prediction value that comprehensively considers speech quality, real-time constraints, computing power consumption, and high-frequency detail complexity is output. The calculation formula is as follows:

[0067] in, This represents the predicted operational quality value corresponding to the k-th layer, which reflects the overall quality performance and stability of the current speech enhancement system in the actual operating environment, and serves as an important basis for subsequent gating strategies, feature injection ratio adjustments, and operational mode switching. This represents the time mean of the comprehensive speech quality evaluation index, which is used to characterize the overall perception quality level and noise reduction consistency of the current speech over a short time scale, and serves as a basic item for predicting operational quality. This represents the end-to-end real-time runtime latency, which reflects the degree to which the current system meets real-time requirements. The greater the latency, the stronger the suppression effect on the predicted runtime quality. This represents the computing power consumption indicator, which reflects the computing resource usage during the current inference process. The computing power consumption is obtained from factors such as comprehensive computing load, video memory utilization, and computing unit activity. This represents the latency-sensitive adjustment coefficient, used to control the weight of the impact of real-time runtime latency on the quality prediction results. It is obtained by normalizing the volatility of real-time runtime latency and has a value range between 0.05 and 0.3. This represents the computing power sensitivity adjustment coefficient, used to control the suppression strength of computing power consumption on the prediction results of running quality. It is calculated by calculating the relative change between real-time video memory usage and reference video memory usage to obtain the video memory usage growth rate, and by calculating the relative growth ratio between current video memory usage and reference video memory usage to obtain the computing power utilization rate. The computing power sensitivity adjustment coefficient is obtained by applying the resource pressure mapping function to the video memory usage growth rate and the computing power utilization rate. The value range is between 0.05 and 0.3. This represents the high-frequency detail complexity metric of the k-th layer, used to reflect the feature activity level of the network in the high-frequency detail enhancement stage at the current scale. It is calculated from the statistical intensity and energy distribution of the high-frequency feature responses. This represents the high-frequency complexity sensitive adjustment coefficient, used to suppress the risk of noise amplification caused by excessively strong high-frequency details in the prediction results of operational quality. It is determined by the first... The amplitude dispersion, energy concentration, and statistical entropy of the high-frequency characteristic response within the sliding window are obtained through nonlinear mapping and normalization, with values ​​ranging from 0.1 to 0.6. By running the quality prediction model, while maintaining speech quality evaluation as the main reference, real-time constraints, computing resource constraints, and high-frequency detail complexity constraints are introduced to enable the running quality prediction results to more realistically reflect the comprehensive operating status of the system in the actual working environment.

[0068] Real-time comparison of predicted runtime quality (MQ) values ​​and MQ prediction thresholds: When the MQ prediction value is greater than the MQ prediction threshold, high-frequency detail feature injection is maintained to improve speech clarity and detail; when the MQ prediction value is less than or equal to the MQ prediction threshold, high-frequency subband gating is triggered to reduce the high-frequency subband fusion bandwidth and switch to a lightweight network to ensure speech integrity and real-time performance. During policy switching, the MQ evaluation value must continuously meet the conditions within a stable window length to suppress frequent state switching caused by short-term fluctuations. The stable window length is determined through statistical analysis of the short-term fluctuations, cross-frame variation characteristics, and noise disturbance duration of the MQ evaluation value, ensuring that the policy triggering conditions remain stable within a continuous window. After policy triggering, a cooling-off period is entered. The duration of the cooling-off period is obtained through analysis of the recovery time after policy switching, the trend change rate of the MQ evaluation value, and the stability of high-frequency subband injection. Parameter adjustments are frozen during the cooling-off period to avoid policy oscillations.

[0069] This implementation scheme achieves simultaneous quantitative evaluation of the quality performance and operational resource status of real-time noise-reduced speech, forming a comprehensive evaluation index based primarily on the time average of speech quality, and influenced by end-to-end real-time runtime latency, computing power consumption, and other factors. The operational quality prediction value, which is jointly constrained by the high-frequency detail complexity metric of the layer, enables the system to stably output operational status criteria that can be used for decision-making under different noise environments and different equipment loads. Based on this, dynamic control strategies such as high-frequency detail feature injection preservation or high-frequency subband gating and lightweight network switching are triggered. At the same time, the frequent switching caused by short-term fluctuations is suppressed by the stable window length and cooling period mechanism, thereby improving the overall operational stability, real-time adaptability and controllability of the voice enhancement system.

[0070] like Figure 2 As shown, the second aspect of this invention provides an adaptive speech denoising system for multi-noise scenarios, including an acquisition and preprocessing module for acquiring input speech data, obtaining operational control training support data and wavelet analysis configuration data; preprocessing the input speech data, operational control training support data, and wavelet analysis configuration data; a wavelet resolution feature decomposition module for performing cross-scale structural decomposition and directional texture decomposition analysis on the input speech data based on the wavelet analysis mechanism, generating a multi-level scale feature set, and performing prior quality constraints in the feature injection stage; and a feature rearrangement and alignment module for uniformly rearranging and aligning the multi-level scale features in terms of spatial size, batch structure, channel dimension, and numerical domain, possessing the characteristics of features within the deep network. Figure 1 The system comprises a unified fusion structure; a neural network and downsampling layer feature injection module, used to perform co-modulation of structural details at the encoder and decoder ends of the nested network, and to execute structural modulation and detail coupling mechanisms based on the results of the co-modulation analysis to construct cross-layer joint feature enhancement representations; a training and hierarchical supervised optimization module, used to construct a multi-objective loss constraint system and hierarchical supervision structure based on input speech data and operational control training support data, and to perform end-to-end parameter learning and optimization on the speech reconstruction network; and an online evaluation and operational control module, used to perform quality resource-sensitive operational status evaluation on the input speech data, and to perform coefficient adjustment, strategy switching and dynamic network path control based on the evaluation results.

[0071] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0072] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. An adaptive speech denoising method for multi-noise scenarios, characterized in that, include: S1, collect input voice data, obtain operation control training support data and wavelet analysis configuration data; Preprocess the input speech data, operation control training support data, and wavelet analysis configuration data; Acquire operational control training support data, which includes: short-time Fourier transform window length, real-time video memory usage, reference video memory usage, target clean speech, actual video memory usage, floating-point operation volume, computing unit utilization, speech frame entry processing time, and output frame generation processing time. S2, based on wavelet analysis mechanism, performs cross-scale structural decomposition and directional texture decomposition analysis on input speech data, generates multi-level scale feature sets, and performs prior quality constraints in the feature injection stage. The specific process of the prior quality constraint in the feature injection stage is as follows: The speech signal-to-noise ratio (SNR) is obtained by calculating the ratio of speech energy to noise energy using a silence segment statistical algorithm on a discrete speech signal sequence. The speech spectral flatness is obtained by calculating the amplitude spectrum using a short-time Fourier transform and then determining the ratio of its geometric mean to its arithmetic mean. Quality assessments are performed on the SNR and spectral flatness of each frame, with real-time comparisons between the SNR and its threshold, and between the spectral flatness and its threshold. When the speech signal-to-noise ratio is less than the speech signal-to-noise ratio threshold or the speech spectrum flatness is greater than the speech spectrum flatness threshold, the corresponding decomposed detail texture subbands will be assigned as candidates for weight reduction. When the speech signal-to-noise ratio is greater than or equal to the speech signal-to-noise ratio threshold and the speech spectrum flatness is less than or equal to the speech spectrum flatness threshold, output the time-frequency statistical feature map, energy distribution map, bandwidth ratio and scale level index information of the coarse smooth sub-band and directional detail texture sub-band of the decomposition level; create a wavelet feature decomposition database, and archive the wavelet low-pass analysis filter sequence and wavelet high-pass analysis filter sequence used, wavelet base type, maximum decomposition level and downsampling strategy to the wavelet feature decomposition database; S3 performs unified rearrangement and alignment of multi-level scale features in terms of spatial size, batch structure, channel dimension and numerical domain, and has a fusionable structural form consistent with the feature maps inside the deep network. S4. Implement coding layer feature fusion analysis with joint modulation of structural details at the coding and decoding ends of the nested network. Based on the coding layer feature fusion analysis results, execute the structural modulation and detail coupling mechanism to construct a cross-level joint feature enhancement representation. S5, based on input speech data and operation control training support data, constructs a multi-objective loss constraint system and hierarchical supervision structure to perform end-to-end parameter learning and optimization on the speech reconstruction network; S6 performs quality and resource-sensitive operational status assessment on the input voice data, and performs coefficient adjustment, strategy switching and dynamic network path control based on the quality and resource-sensitive operational status assessment results. The specific process of performing quality resource-sensitive operational status assessment on the input voice data, and executing coefficient adjustment, strategy switching, and dynamic network path control based on the assessment results, is as follows: The high-frequency detail complexity measure of the k-th layer is calculated from the statistical intensity and energy distribution of the high-frequency feature response. A comprehensive speech quality evaluation index is obtained by fusing speech perception quality scoring, speech intelligibility estimation, structural consistency measurement, and speech naturalness assessment on real-time denoised speech. This comprehensive speech quality evaluation index is then input to create a runtime quality prediction model. By using the time mean of the comprehensive speech quality evaluation index as a base term, a resource constraint suppression term consisting of end-to-end real-time runtime latency and computing power consumption is introduced into the denominator. The high-frequency detail complexity metric and sensitivity adjustment coefficient of the layer are corrected, and the predicted value of the running quality is output. Real-time comparison of predicted operation quality values ​​and predicted operation quality thresholds: When the predicted operation quality value is greater than the predicted operation quality threshold, high-frequency detailed feature injection is maintained; When the predicted operational quality value is less than or equal to the threshold, a high-frequency subband gating operation is triggered to reduce the high-frequency subband fusion bandwidth and switch to a lightweight network. After the policy is triggered, a cooling-off period begins. The duration of the cooling-off period is determined based on the recovery time after the policy switch, the rate of change of the operational quality evaluation value trend, and the stability analysis of high-frequency subband injection. Parameter adjustments are frozen during the cooling-off period.

2. The adaptive speech denoising method in a multi-noise scenario according to claim 1, characterized in that: The specific process for preprocessing the input speech data, operation control training support data, and wavelet analysis configuration data is as follows: Collect input speech data, which includes: discrete speech signal sequence, speech sampling frequency and noisy speech sample; Obtain wavelet analysis configuration data, which includes: time sampling index, wavelet basis type, and maximum number of wavelet decomposition levels; The discrete speech signal sequence is processed by symmetric extension to handle convolutional overflows at both ends, maintaining the energy continuity between smooth and abrupt regions. The input speech data is structurally segmented according to the time-varying characteristics of speech through framing and windowing algorithms. The discrete speech signal sequence is converted into a two-dimensional spectrogram through short-time Fourier transform and Mel filter bank transform. The amplitude of the input speech data and the training support data for operation control are scale-aligned and units are unified through standardization and normalization algorithms.

3. The adaptive speech denoising method in a multi-noise scenario according to claim 1, characterized in that: The specific process of performing cross-scale structural decomposition and directional texture decomposition analysis to generate multi-level scale feature sets is as follows: Obtain the time sampling index and the maximum number of wavelet decomposition levels; generate wavelet low-pass and high-pass analysis filter sequences based on the two-scale recursive equations of the wavelet basis determined by the wavelet basis type. The low-pass analysis filter sequence is obtained by discretization using a scaling function, and the high-pass analysis filter sequence is derived from the low-pass analysis filter sequence through orthogonal mirror transformation; convolve the input discrete speech signal sequence with the wavelet low-pass and high-pass analysis filter sequences respectively, and perform double downsampling on the convolution output to obtain the first-level smoothing coefficient sequence and the first-level detail coefficient sequence; The wavelet decomposition operation is continuously performed on the first-level smoothing coefficients. The smoothing coefficients of the previous level are used as the input of the next level. Convolution and summation with the wavelet low-pass analysis filter sequence and the wavelet high-pass analysis filter sequence are performed respectively, and then downsampled by a factor of 2 to obtain the k-th level smoothing coefficients and the k-th level detail coefficients. This process is repeated until the maximum number of decomposition levels is reached. The input is a two-dimensional spectrogram. In the first wavelet decomposition stage, coarse smooth subband, horizontal detail subband, vertical detail subband and diagonal detail subband are obtained. After the two-dimensional spectrogram undergoes the k-th convolution and pooling operation, it forms multi-channel time-frequency features, which represent the k-th layer feature map. According to the dimension of the injected node and the task requirements, the wavelet subbands of the corresponding one-dimensional discrete sequence dimension and two-dimensional spectrogram dimension are selected to participate in the fusion.

4. The adaptive speech denoising method in a multi-noise scenario according to claim 1, characterized in that: The specific process of having a fusionable structural form consistent with the internal feature maps of the deep network is as follows: Input the scale subband coefficient set output from wavelet decomposition at levels 1 to N, perform rearrangement and alignment on scale subbands at different levels, and use interpolation and pooling to adjust the coarse smooth subband and directional detail subband to the same spatial size as the k-th level feature map at the encoder; under batch conditions, perform batch dimensionality expansion on the scale subband and perform channel mapping to convert it from a single channel to a multi-channel feature with the same number of channels as the feature map at the encoder; When a scale subband is marked as a candidate for weight reduction, its channel weight is reduced during the channel mapping process; After mapping is completed, the multi-scale features are concatenated according to the channel dimension and input into the corresponding layer of the target network.

5. The adaptive speech denoising method in a multi-noise scenario according to claim 1, characterized in that: The specific process of implementing the coding layer feature fusion analysis for joint modulation of structural details is as follows: In a multi-level nested U-shaped network structure, the original network feature map of the kth layer is obtained by the convolutional layer at the encoding end after the kth pooling, and the aligned scale subbands are injected into the corresponding layers according to the level. Based on the original network feature map of the k-th layer at the encoding end, a feature modulation mapping function of the k-th layer is constructed to jointly model the structure and details of the features. The modulation features are constrained and normalized by a nonlinear gated mapping function. The obtained modulation weights are multiplied element-wise with the original network feature map to obtain the modulation features. The modulation features are then residually superimposed with the original network features to generate the enhanced feature map of the k-th layer.

6. The adaptive speech denoising method in a multi-noise scenario according to claim 1, characterized in that: The specific process of constructing a cross-layer joint feature enhancement representation by executing a structure modulation and detail coupling mechanism based on the feature fusion analysis results of the coding layer is as follows: The deep coding stage uses coarse-scale smooth subbands and their convolutional responses as the main external scale supplementary features; the shallow coding and decoding stages use the detail energy map composed of the horizontal detail subband norm, vertical detail subband norm and diagonal detail subband norm to complete directional detail enhancement through ratio regularization; the decoding stage uses cross-layer skip connections to input the enhanced feature map after wavelet scale injection and the corresponding shallow features into the decoding module. The computing power consumption is obtained by statistically analyzing the actual video memory usage, floating-point operation volume, and computing unit utilization rate during the network forward inference process. The end-to-end processing time from the input of a voice frame to the generation of an output frame yields the real-time runtime latency; real-time comparisons are made between the real-time runtime latency and the real-time latency threshold, and between computing power consumption and the computing power consumption threshold. When the real-time runtime latency exceeds the real-time latency threshold, a fallback strategy is implemented, including high-frequency subband channel pruning, subband bandwidth compression, and reducing the number of wavelet decomposition layers. When the computing power consumption exceeds the computing power consumption threshold, the maximum number of wavelet decomposition layers is reduced, the number of deep injections is reduced, and the quantization decoding head is reduced. When the real-time runtime latency is less than or equal to the real-time latency threshold and the computing power consumption is less than or equal to the computing power consumption threshold, the following are output: training discrete speech signal sequence, network feature map, hierarchical sub-band mapping table, channel number configuration, bandwidth allocation record and rollback strategy execution log.

7. The adaptive speech denoising method in a multi-noise scenario according to claim 1, characterized in that: The specific process of constructing a multi-objective loss constraint system and a hierarchical supervision structure to perform end-to-end parameter learning and optimization on the speech reconstruction network is as follows: Noisy speech samples and corresponding clean target speech samples are used to form paired training data. The training discrete speech signal sequence is used as the model training input. The speech reconstruction model is constructed through end-to-end supervised learning. The training loss function is used to guide backpropagation optimization. The supervision method adopts a hierarchical time-frequency supervision mechanism. During the training process, data augmentation strategies such as noise enhancement, signal-to-noise ratio perturbation, reverberation synthesis, echo superposition, and background replacement are implemented. The learning rate scheduling adopts a preheating incremental cosine annealing mechanism, combined with exponential moving average, gradient pruning, and stability control. Output the best checkpoint file, exponential moving average model version, loss weight configuration and data augmentation configuration, and generate evaluation results for speech perception quality, speech intelligibility, structural consistency and speech naturalness, and create and archive them to the model training database.

8. An adaptive speech denoising system for multi-noise scenarios, employing the adaptive speech denoising method for multi-noise scenarios as described in any one of claims 1-7, characterized in that, include: The acquisition and preprocessing module is used to acquire input speech data, obtain operation control training support data and wavelet analysis configuration data; Preprocess the input speech data, operation control training support data, and wavelet analysis configuration data; The wavelet resolution feature decomposition module is used to perform cross-scale structural decomposition and directional texture decomposition analysis on input speech data based on the wavelet analysis mechanism, generate multi-level scale feature sets, and perform prior quality constraints in the feature injection stage. The feature rearrangement and alignment module is used to uniformly rearrange and align multi-level scale features in terms of spatial size, batch structure, channel dimension and numerical domain, and has a fusionable structural form consistent with the feature maps inside the deep network. The neural network and downsampling layer feature injection module is used to perform coding layer feature fusion analysis by implementing joint modulation of structural details at the encoding and decoding ends of the nested network. Based on the coding layer feature fusion analysis results, it executes the structural modulation and detail coupling mechanism to construct a cross-level joint feature enhancement representation. The training and hierarchical supervision optimization module is used to construct a multi-objective loss constraint system and hierarchical supervision structure based on input speech data and operation control training support data, and to perform end-to-end parameter learning and optimization on the speech reconstruction network. The online evaluation and operation control module is used to evaluate the quality and resource-sensitive operation status of the input voice data, and to perform coefficient adjustment, strategy switching and network path dynamic control based on the evaluation results.

Citation Information

Patent Citations

  • A speech noise reduction method

    CN109378013B

  • Cross-lingual ai voice cloning method, system, and storage medium thereof

    CN120766655B

  • Keyword Detection In The Presence Of Media Output

    US20200090647A1