Voice noise reduction method based on reasoning optimization and Bluetooth earphone
By employing a speech denoising method based on inference-optimized U-shaped networks and causal temporal weighted deep filtering, the problem of joint suppression of non-steady-state noise and reverberation in Bluetooth headsets and conference terminals is solved, achieving low-latency, high-quality speech enhancement effects, suitable for scenarios such as online meetings and Bluetooth calls.
Patent Information
- Application Number
- CN202511622355.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-11-07
AI Technical Summary
Existing technologies have shortcomings in areas such as the coexistence of non-steady-state noise and reverberation, strict latency and power consumption limitations, complex domain amplitude and phase consistency constraints, and coordination with downstream speech understanding and synthesis links, making it difficult to achieve stable, low-latency, and highly intelligible speech enhancement on devices such as Bluetooth headsets and conference terminals.
A speech denoising method based on inference optimization is adopted. By using short-time Fourier transform, U-shaped network and causal time-weighted deep filtering, combined with minimum mean square error post-processing, it achieves denoising and dereverberation with consistent amplitude and phase. It only relies on current and historical speech information, reduces computation and storage overhead, and adapts to sudden non-steady-state noise and reverberation.
It achieves joint suppression of non-steady-state noise and reverberation under strict time delay constraints, outputting stable, continuous, low-latency high-quality voice, suitable for scenarios such as online meetings and Bluetooth calls, reducing power consumption and improving voice clarity and intelligibility.
Smart Images

Figure CN121506162A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent speech processing technology, and in particular to an edge-side single-channel speech noise reduction and de-reverberation method based on inference optimization, which can be applied to devices such as Bluetooth headsets, mobile phones and conference terminals to achieve low-latency speech enhancement in offline simultaneous interpretation and voice interaction scenarios. Background Technology
[0002] With the improvement of mobile computing power and terminal-side chips, voice enhancement is gradually migrating from the cloud to local devices such as Bluetooth headsets, smartphones, and conference terminals to meet the requirements of privacy protection and low latency. In real-world applications, sudden and rapidly changing non-steady-state noise and room reverberation often coexist. Meanwhile, battery-powered devices are limited in terms of computing power and power consumption, making it difficult for traditional processing chains to achieve a balance between audible quality, robustness, and end-to-end latency. Therefore, the industry needs to collaboratively optimize model structure, inference computation, and engineering deployment. For scenarios such as Bluetooth calls, offline translation, and simultaneous interpretation in live conferences, a stable integrated closed loop needs to be formed with local speech recognition, machine translation, and speech synthesis modules.
[0003] Existing traditional methods mainly include spectral subtraction, Wiener filtering, and gain estimation targeting minimum mean square error. These methods are usually based on the assumption of "nearly stationary noise" and are effective for relatively stable interference (such as white noise and some traffic noise). However, under non-steady-state noise conditions such as wind noise when Bluetooth headsets are close to the pickup port, keyboard typing and footsteps in meeting rooms, and mutual interference from multiple people talking, they are prone to problems such as "musical noise," speech distortion, and sharp sound. In reverberant environments, these methods have limited ability to distinguish between early and late reflection energy, making it difficult to significantly improve clarity and intelligibility. Furthermore, they rely on manually setting noise estimation and smoothing parameters, resulting in insufficient adaptability across devices and scenarios, and significant performance fluctuations.
[0004] Currently, deep learning-based speech enhancement has achieved better results in complex scenarios. A common approach is to use a symmetrical encoding, enhancement, and decoding structure, combining convolutional and recurrent units, to perform mask estimation or direct end-to-end reconstruction of noisy speech to suppress wind noise, venue reverberation, and human voice interference. However, many solutions, in pursuit of performance metrics, rely on bidirectional time modeling or require the use of future frame information. The inference phase must cache multiple input frames, making it difficult to meet the stringent latency requirements of Bluetooth calls and conference sound reinforcement. Some models only optimize the amplitude spectrum while neglecting phase consistency, easily resulting in a "hollow" or "metallic" sound under strong reverberation or low signal-to-noise ratio conditions. Furthermore, these networks have large parameter scales and computational demands, high memory consumption, and significant accuracy degradation after quantization and compression, leading to high power consumption and heat generation, which is detrimental to long-term stable operation on battery-powered devices.
[0005] In terms of engineering and training strategies, existing methods often fail to adequately cover noise types, signal-to-noise ratio distributions, and reverberation times. They also suffer from unreasonable low-frequency weighting, affecting the richness and prosodic fidelity of speech. Furthermore, they lack a coordinated design that matches causal constraints, relying solely on the current frame and its preceding frames for both temporal weighted fusion and statistical post-processing. This makes it difficult to achieve a robust balance between residual noise and speech distortion after enhancement. In Bluetooth headset local offline translation and simultaneous interpretation links for live conferences, front-end enhancement is often designed separately from subsequent speech recognition, machine translation, and speech synthesis. There is a lack of unified planning for latency jitter, caching strategies, and resource contention, resulting in an unstable end-to-end experience.
[0006] Chinese patent CN112767959B discloses a method that first performs a short-time Fourier transform on the speech to obtain a time-frequency representation, then uses an encoding and decoding network for enhancement, and introduces information from the previous frame when decoding the current frame to reduce latency. This approach utilizes only historical information as a causal path. While this approach helps reduce buffering, it still has the following shortcomings: First, it does not explicitly constrain the amplitude and phase consistency in the complex domain, making it prone to auditory distortion under strong reverberation or low signal-to-noise ratio conditions. Second, it does not provide a timing mechanism that performs non-negative, normalized weighted fusion of adjacent times and frequencies based solely on the current frame and its preceding frames, making it difficult to stably suppress sudden and alternating non-steady-state noise. Third, it does not coordinate with statistical post-processing aimed at minimizing mean square error, and it does not systematically explain key aspects such as prior and posterior signal-to-noise ratios and frequency-based gain correction. In terms of applications, it does not provide lightweight inference details for low-power and low-computing-power environments for Bluetooth headsets, does not explain the latency and resource coordination between offline translation links and local speech recognition, machine translation, and speech synthesis, and does not provide specific adaptation strategies for complex interferences such as overlapping speakers, keyboard typing, and wind noise in live meetings. As a result, it is difficult to simultaneously achieve enhanced quality, real-time performance, and energy consumption constraints in the above scenarios.
[0007] Chinese patent CN109949821B employs a symmetrical encoder-decoder structure with cross-layer connections for dereverberation in far-field scenarios to improve clarity. While this solution focuses on reverberation suppression, it suffers from several issues: First, it primarily relies on amplitude domain features, failing to explicitly limit the amplitude-phase consistency in the complex domain, leading to subjective defects such as "hollowness" and "metallic sound" in the recovered signal. Second, it lacks strict causal modeling requirements and a neighboring-time / neighboring-frequency weighted fusion mechanism that "only depends on the current frame and its preceding frames," resulting in insufficient adaptability to sudden noise and speaker rotation. Third, it lacks a collaborative framework with statistical post-processing, failing to provide supporting designs for key steps such as gain estimation and noise tracking. At the application level, it does not propose lightweighting and quantification measures to address the computing power, storage, and power consumption constraints of Bluetooth headsets and conference terminals. It also fails to explain the closed-loop coordination and latency control with local speech recognition, machine translation, and speech synthesis, and does not provide a systematic solution for handling latency jitter, resource contention, and output stability in simultaneous interpretation at live conferences. Therefore, in scenarios that are sensitive to latency and have limited resources, such as Bluetooth calls, offline translation, and on-site meetings, the above solutions are difficult to meet the comprehensive requirements of stability, low latency, and high intelligibility.
[0008] Therefore, existing technologies still have shortcomings in terms of the coexistence of non-steady-state noise and reverberation, strict time delay and power consumption limitations, complex domain amplitude and phase consistency constraints, and coordination with downstream speech understanding and synthesis links. Comprehensive improvements are needed to achieve model lightweighting, enhanced quality, and system-level stability. Summary of the Invention
[0009] To address the aforementioned technical problems, this invention provides a speech denoising method based on inference optimization. This method is capable of handling complex acoustic scenarios such as conferences, phone calls, and simultaneous interpretation, including sudden non-steady-state noises like clapping, keyboard clicks, and footsteps, wind noise highly overlapping with low-frequency speech, and reverberation tails caused by both early and late reflections. Under strict time-delay constraints, it relies solely on current and historical speech information to achieve amplitude and phase consistent denoising and de-reverberation, avoiding subjective distortions such as "musical noise" and "metallic sounds." Furthermore, it achieves continuous, stable, and robust output across various scenarios while maintaining controllable computational and storage overhead.
[0010] A speech denoising method based on inference optimization includes the following steps:
[0011] Step S1: Perform a short-time Fourier transform on the noisy speech to obtain time-frequency features;
[0012] Step S2: Input the features into a U-shaped network that includes encoding, enhancement, and decoding and has skip connections to obtain an initial mask;
[0013] Step S3: Based on causal time-weighted depth filtering that only utilizes the time-frequency information of the current frame and its preceding frames, optimize the initial mask to obtain an optimized mask; use minimum mean square error class post-processing to estimate the prior and posterior signal-to-noise ratios based on the power spectra of noisy speech and estimated clean speech, calculate the gain to correct the optimized mask, and obtain the final mask;
[0014] Step S4: Weight the noisy spectrum with the final mask, and output the denoised speech through inverse short-time Fourier transform. Multiply the corrected mask with the noisy spectrum to obtain the denoised spectrum, and output the denoised speech through inverse short-time Fourier transform.
[0015] Furthermore, the data processing in step S1 includes: acquiring room impulse response, clean speech, and noise data; adjusting the reverberation duration of the room impulse response to maintain attenuation characteristics; convolving the clean speech with the adjusted room impulse response to obtain reverberant speech; and, based on preset noise selection and intensity configuration, enhancing and normalizing the amplitude consistency of the reverberant speech and noise in the spectral domain, and mixing them according to a preset signal-to-noise ratio strategy to obtain training input, with the clean speech serving as the corresponding training target.
[0016] The reverberation time of the room's impulse response is adjusted so that the target reverberation time T60 is between 0.10 seconds and 0.60 seconds, while maintaining its original decay pattern;
[0017] Noise is selected with a probability of p < 0.7. One or more noise sources are selected, and the signal-to-noise ratio of the training samples is set in the range of -10 to +20 dB.
[0018] Enhancement processing and amplitude consistency normalization are performed on the reverberant speech and / or the noise in the spectral domain, and the level of at least one signal is normalized to a predetermined range; the processed signals are mixed according to the signal-to-noise ratio strategy to form a training input.
[0019] Furthermore, the neural network construction in step S2 satisfies the following limitations: the encoding part downsamples and extracts the input spectral features; the high-dimensional features are calculated separately by the grouped recurrent neural network of the enhancement part and then normalized and fused; the decoding part upsamples and reconstructs the enhanced features; a skip connection is set between encoding and decoding, and the skip connection performs feature transformation in the channel dimension and passes it by feature addition; wherein, the encoding part and the decoding part are respectively composed of multiple levels of convolutional layers and deconvolutional layers, except for the first level of convolutional layer of the encoding part and the last level of deconvolutional layer of the decoding part, the rest... The size of the convolutional kernel and deconvolutional kernel at each level is 3×1; the size of the convolutional kernel of the first level convolutional layer of the encoding part and the last level deconvolutional layer of the decoding part is 5×2; the enhancement part adopts a grouped recurrent neural network, which divides the input features into two groups in the channel dimension and configures corresponding recurrent units for independent calculation. The outputs of each group are normalized by an instantaneous normalization layer and then fused; a skip connection is set between the encoding part and the decoding part. The skip connection performs channel matching through 1×1 convolution and fuses the matched encoding side features and decoding side features by feature addition.
[0020] Furthermore, the decoding part uses sub-pixel convolution for upsampling, rearranging the frequency dimension step by step at integer multiples to restore the resolution until it matches the frequency resolution before encoding; the features of the encoding part are first matched by 1×1 convolution; the matched features are then fed into a gated convolution branch with a sigmoid activation function to generate channel-wise weights; these weights are used to weight the encoded features and then fused with the corresponding decoded features element-wise; the output of the enhancement part is normalized for each time frame and each frequency point in the instantaneous normalization layer before fusion; the two sets of recurrent units in the enhancement part receive only the temporal features of the current frame and its preceding frames in a causal manner for calculation, and the two sets of parameters are not shared; the output of the U-shaped network is a complex-valued mask, which weights the real and imaginary parts of the input spectrum, and the amplitude of the mask is limited to the range of 0 to 1.
[0021] Furthermore, step S3 includes: performing a short-time Fourier transform on the training input to obtain spectral features and inputting them into the U-shaped network to obtain a mask; using a joint loss function composed of a time-domain signal-to-noise ratio term, a speech component term, and a noise component term, setting weighting coefficients for different frequency sub-bands to distinguish processing priorities, and performing backpropagation based on the joint loss to update network parameters; performing a short-time Fourier transform on the training input to obtain complex spectral features and inputting them into the U-shaped network to obtain a mask;
[0022] The joint loss consists of one or more of the following three terms: a time-domain signal-to-noise ratio (SNR) term, used to improve the SNR of the estimated speech relative to the target clean speech; a speech component term, used to constrain the estimated speech to approximate the target clean speech in amplitude and / or phase; and a noise component term, used to penalize residual noise energy in the estimated speech. The frequency range is divided into multiple sub-bands, and weight coefficients are assigned to each sub-band. These weights are preset or learnable and can be adaptively adjusted according to training epochs, sample SNR, or noise steady-state characteristics, ensuring that the weight of low-frequency sub-bands is not lower than that of mid-to-high-frequency sub-bands and / or assigning higher weights to sub-bands with prominent non-steady-state noise. Backpropagation is performed using a gradient descent-type optimization method based on the joint loss to update the network parameters.
[0023] Further, step S4 includes: performing a short-time Fourier transform on the noisy speech to obtain complex spectral features, and inputting them into the U-shaped network to obtain an initial mask; employing a time-weighted deep filtering strategy, using only the time-frequency information of the current frame and its preceding frames at this frequency point and its adjacent frequency points, fusing them according to non-negative and normalized weighting coefficients to obtain an optimized mask, without using any subsequent frame information; using post-processing based on minimum mean square error to correct the optimized mask: calculating the posterior signal-to-noise ratio γ based on the power spectra of the noisy speech and the estimated clean speech, estimating the prior signal-to-noise ratio ξ in a decision-oriented manner, and obtaining the gain according to a predetermined frequency point gain formula to correct the mask; multiplying the corrected mask by the noisy spectrum to obtain the denoised spectrum, and outputting the time-domain denoised speech through an inverse short-time Fourier transform.
[0024] Furthermore, the number of back-view frames L in the temporal weighted depth filtering is a positive integer; the neighborhood width K of adjacent frequency points is a positive integer; the weighting coefficients are normalized frequency-by-frequency point and can be generated by gated convolution or attention mapping, and depend only on the information of the current frame and its preceding frames; the frequency point gain formula for the minimum mean square error post-processing is any one of Wiener gain, logarithmic amplitude minimum mean square error gain, or spectral amplitude minimum mean square error gain; the prior signal-to-noise ratio ξ is obtained by exponentially smoothing the estimation result of the previous time step with the current posterior signal-to-noise ratio using a decision-oriented method, with the smoothing factor valued between 0 and 1; the noise power spectrum is estimated using any one of online noise tracking or speech presence probability methods.
[0025] A Bluetooth headset, employing the method described above, includes: a pickup unit, an audio processing unit, a memory, a Bluetooth communication unit, and a speaker unit; the audio processing unit is configured to perform the following inference noise reduction process on the pickup signal: performing a short-time Fourier transform on the noisy speech to obtain spectral features; calling a neural network for encoding, enhancement, and decoding to output an initial mask; generating an optimized mask based on temporal weighted depth filtering using only the current frame and its preceding frame information; calculating the prior / posterior signal-to-noise ratio based on minimum mean square error post-processing and correcting the mask according to gain; multiplying the corrected mask by the noisy spectrum and performing an inverse short-time Fourier transform to obtain the denoised speech, which is then output to the speaker unit and / or transmitted via the Bluetooth communication unit;
[0026] Its features include: performing inference noise reduction on the picked-up signal to obtain denoised speech; performing local speech recognition on the denoised speech to generate source language text in the absence of a network connection; using a locally stored neural machine translation model to perform offline translation of the source language text to obtain target language text; using a locally stored speech synthesis model to offline synthesize the target language text into target language speech and outputting it to the speaker unit; presenting the source language text and target language text on the display unit and / or sending them through the communication unit; wherein the speech recognition, translation, and speech synthesis models are all stored in the memory and inferred locally on the processor.
[0027] Furthermore, the neural machine translation model automatically determines the translation direction and generates target language text in real time. The machine translation model is subjected to low-bit integer quantization, symmetrical scaling factor is used, and linear dequantization is performed during the inference stage to reduce storage bandwidth and computing power consumption.
[0028] The quantized model is converted into a terminal-specific inference format, which includes at least binary weights, a unified sub-word vocabulary, and model configuration.
[0029] The model is run on a general-purpose processor on a local inference engine, employing a small beamwidth search combined with a repetition suppression strategy, and allocating threads according to a predetermined proportion of available processing resources to achieve a balance between real-time performance and energy consumption.
[0030] Configure early stop criteria and set single-step generation waiting delay threshold during the decoding stage to ensure that the end-to-end delay of simultaneous interpretation is controllable.
[0031] The input speech is incrementally decoded using streaming coding and attention caching mechanisms, with fixed or adaptive windows, and look-ahead and stabilization windows configured to reduce short-term writeback and output flicker.
[0032] When the observed end-to-end delay exceeds the threshold, decoding degradation is performed, including beamwidth reduction or greedy decoding; when the direction confidence is below the threshold, manual confirmation or default routing is triggered to ensure output controllability.
[0033] Memory and power management strategies are set up to reduce peak resource consumption through weighted segmented loading, cache reuse, and low-power scheduling, and unknown sub-word backoff rules are used to handle symbols outside the word list to maintain decoding continuity.
[0034] Furthermore, the neural machine translation model employs a shared sub-word segmenter jointly trained for at least two languages to construct a unified vocabulary, and uses this unified vocabulary to uniformly encode and decode the source and target languages; wherein:
[0035] The shared word segmenter is based on a statistical word segmentation algorithm, which automatically learns the set of word units and their corresponding mappings from mixed corpora containing multilingual text;
[0036] The unified vocabulary includes functional tags and sets vocabulary size and character coverage parameters to simultaneously cover language elements from different writing systems and reduce out-of-vocabulary symbols;
[0037] The addition of new languages is achieved by updating the mixed corpus and jointly or sequentially retraining the word segmenter and translation model; during the expansion process, the existing sub-word indexes are frozen and retained and the new sub-words are placed in the expansion area to maintain backward compatibility with the existing model;
[0038] During the inference phase, unregistered symbols are processed using rules that fall back to finer-grained units to ensure the continuity and stability of cross-language decoding.
[0039] The speech denoising method and system based on inference optimization provided by this invention have the following advantages: under strict time delay and power consumption constraints, it only relies on the current frame and its preceding frames to complete causal processing, thereby achieving joint suppression of non-steady-state noise and room reverberation, and outputting stable, continuous, low-latency high-quality speech. It is especially suitable for scenarios that are sensitive to time delay and power consumption, such as online meetings, on-site meetings and Bluetooth calls.
[0040] This invention employs a time-weighted deep filtering mechanism to apply non-negative and frequency-point-by-frequency-point normalized weighted fusion to the time-frequency information of the current frequency and adjacent frequencies without using any subsequent frames. This significantly reduces waiting and buffering overhead, effectively suppressing sudden and rapidly changing interference such as clapping, keyboard sounds, footsteps, paper turning, and wind noise, ensuring a consistent listening experience. Simultaneously, amplitude and phase-consistent mask estimation is performed in the complex domain. The network directly outputs a complex-valued mask, weighted by the real and imaginary parts of the spectrum, and imposes interval constraints on the mask amplitude. Based on this, a post-processing link targeting minimum mean square error is introduced. Combining prior and posterior signal-to-noise ratios, gain is calculated frequency-point-by-frequency, and the network results are further corrected, forming a collaborative closed loop of "network estimation - statistical correction." This effectively reduces subjective artifacts such as "musical noise," "metallic sounds," and "hollowness," significantly improving clarity, intelligibility, and naturalness, especially under conditions of strong reverberation and low signal-to-noise ratio.
[0041] In terms of structural design, a "U-shaped" network is adopted, which includes encoding, enhancement, and decoding with skip connections: the main layer uses small-sized convolutional kernels to reduce computational load, while the first and last layers moderately enlarge the convolutional kernels to ensure representational power. The enhancement part uses grouped recurrent units combined with instantaneous normalization, the decoding end uses subpixel rearrangement for upsampling, and the cross-layer skip connections fuse features by adding them together after one-to-one channel matching. This combination significantly reduces multiply-accumulate operations and parameter size while ensuring reconstruction quality, reducing memory usage and power consumption, and making it more stable for long-term operation. It also allows for rapid parameter tuning by reviewing parameters such as the number of frames, neighborhood width, and weight generation method, facilitating smooth migration from small meeting rooms to large lecture halls.
[0042] In terms of training and data construction, comprehensive coverage of the real acoustic distribution of conferences and calls is achieved: the duration of room impulse response is adjusted to ensure reverberation time falls within a reasonable range. One or more noise sources are selected according to a set probability, covering signal-to-noise ratios from low to high. Enhancement and amplitude consistency normalization are performed in the spectral domain; a joint loss composed of time-domain signal-to-noise ratio, speech components, and noise components is adopted, and differentiated weights are set for each frequency band, highlighting low-frequency and non-steady-state sub-bands. As a result, good robustness and cross-site consistency are maintained even under complex conditions such as high wind noise, weak direct sound, and multiple speakers taking turns.
[0043] In terms of application effects, in online meetings, causal and low-latency noise reduction and dereverberation are performed on the voice before it enters the platform for encoding, decoding, and transmission, which can significantly shorten end-to-end latency and reduce buffer jitter. When there are network fluctuations or limited bitrates, due to the higher proportion of direct sound, subtitle recognition is more stable, sentence segmentation is more accurate, word errors and repetitions are significantly reduced, automatic mute and endpoint detection are more reliable, and listening fatigue at a distance is reduced. In Bluetooth headset scenarios, lightweight inference and low-power design reduce heat generation and extend battery life, and alignment is stable under fixed frame length and buffering strategies, reducing interruptions and "squeezing" sensations. It has stronger suppression of near-field interference such as wind noise, clothing friction, and breathing near the pickup port, while maintaining the richness and naturalness of the voice. Front-end local processing also reduces dependence on the network and cloud, improves privacy and usability, and maintains clear and intelligible voice output even in weak or offline networks. In summary, this invention offers systematic advantages in areas such as low latency due to causality, consistency across multiple domains, temporal weighted deep filtering, collaborative statistical post-processing, lightweight structure and training coverage, and overall integration with the conference link, significantly improving the practical usability and deployment efficiency of speech enhancement. Attached Figure Description
[0044] Appendix Figure 1 This is a flowchart of the speech denoising training path based on inference optimization in this invention;
[0045] Appendix Figure 2 This is a flowchart of the speech noise reduction inference path based on inference optimization in this invention;
[0046] Appendix Figure 3 This is a flowchart illustrating the construction of training inputs and training targets in this invention;
[0047] Appendix Figure 4 This is a schematic diagram of the U-shaped network structure with encoding-enhancement-decoding and skip connections in this invention;
[0048] Appendix Figure 5 This is a schematic diagram of the skip connection fusion mechanism in this invention;
[0049] Appendix Figure 6 This is a schematic diagram of the causal depth filtering mask in this invention. Detailed Implementation
[0050] The technical solution of the present invention will now be clearly and completely described in conjunction with the accompanying drawings. In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0051] This embodiment provides a low-latency noise reduction and reverberation reduction method for real-time speech enhancement on the edge side. Its overall process consists of a training path and an inference path: as follows... Figure 1 As shown, during the training phase, the training data production module first generates paired samples, including reverberant speech after convolution of the room impulse response, multiple types of noise configured according to the signal-to-noise ratio range, and noisy speech formed by mixing them according to the strategy. Amplitude consistency normalization and spectral domain enhancement are then performed. The noisy speech undergoes short-time Fourier transform to obtain complex time-frequency features containing both real and imaginary parts, which are then fed into a U-shaped neural network containing encoding, enhancement, and decoding stages with multi-level skip connections. The encoding end uses small-size convolutions to extract multi-scale representations, the middle enhancement unit performs causal temporal modeling, and the decoding end recovers resolution step-by-step through upsampling. The skip connections first perform channel alignment using one-to-one defined channel-matching convolutions, then generate channel-by-channel weights using a gating method and fuse them element-wise with the decoding end features. The network directly outputs a complex-valued mask with limited amplitude. The network output and the target clean spectrum are then fed into the loss calculation module, which uses a joint loss consisting of a time-domain signal-to-noise ratio term, a speech component term, and a noise component term, combined with frequency band weights to highlight key frequency bands. The model undergoes backpropagation and parameter updates until convergence on the validation set.
[0052] Reasoning stage, such as Figure 2As shown, the noisy speech acquired in real time undergoes short-time Fourier transform to obtain complex time-frequency features, which are input into a trained neural network to obtain an initial mask. This initial mask is then fed into an inference optimization module: First, without using future information, temporally weighted deep filtering is performed based solely on the neighboring time and frequency information of the current frame and its preceding frames. Non-negative weights, normalized at each frequency point, are used to causally fuse the initial mask, resulting in an optimized mask. Next, minimum mean square error post-processing is implemented, calculating the gain at each frequency point based on the posterior and prior signal-to-noise ratios, and statistically correcting the optimized mask to obtain the final mask. The final mask is then multiplied point-by-point with the noisy spectrum to obtain the enhanced spectrum, which is then restored to a time-domain signal via inverse short-time Fourier transform to output the denoised speech. Through the aforementioned collaborative link of "network estimation—causal deep filtering—minimum mean square error correction," this invention effectively suppresses non-steady-state noise and reverberation while maximizing the preservation of speech components, ensuring low latency and end-device deployability.
[0053] Appendix Figure 3 The complete process from raw data preparation to the generation of training input and training target pairing is presented, including consecutive steps such as "raw data preparation → reverberation processing → speech and noise mixing → training data generation". The source and role of the three basic types of data, room impulse response, clean speech and noise, are clearly listed, and key parameters and specifications are defined.
[0054] At the starting point of the data production pipeline, three types of basic data are prepared: Room Impulse Response (RIR), clean speech, and noise. RIR is used to characterize the spatial reflection and attenuation characteristics; clean speech serves as the target reference in supervised learning; and noise is used to construct interference conditions of varying intensities and forms from low to high. To ensure consistency in subsequent processing, a file format with a uniform sampling rate and bit depth is preferred. If necessary, the sampling rate is resampled, and the amplitude is initially normalized to eliminate distribution shifts caused by device differences. This three-dimensional data base directly supports the claim's requirement to "acquire RIR, clean speech, and noise as the starting point for data processing."
[0055] Appendix Figure 3The chain of "RIR → T60 targeting → convolution operation → reverberant speech" was clearly defined. Specifically, the reverberation time of the RIR was shortened, and the target reverberation time was set to T60 = 0.15 seconds while maintaining the exponential decay pattern. To cover different spatial conditions, it can be adjusted within the range of 0.10 to 0.60 seconds. The purpose of this step is to construct a sample distribution with controllable reverberation intensity, so that the model can learn the influence of direct sound and early and late reflections during training, reducing overfitting to a single scene. The clean speech was then convolved with the target-processed RIR in the temporal domain to obtain the reverberant speech. The convolved speech contains information about the speech itself and is superimposed with room reflections and attenuation effects, which is the basis for subsequent mixing, enhancement, and loss calculation. When processing the convolution operation at the frame level, the frame length and stride can be consistent with the subsequent short-time Fourier transform to reduce the mismatch between time-frequency features in the training and inference paths. This step supports the claim's limitation regarding "convolving clean speech with processed RIR to obtain reverberant speech".
[0056] To enhance sample diversity and coverage, the noise branch employs a two-stage strategy: the first stage selects noise sources at the sample level with a set probability, allowing for the selection of one or more sources at a time; the second stage configures a target signal-to-noise ratio (SNR) for each sample at the signal level. A directly reproducible implementation example is a selection probability p < 0.5 and an SNR range of -5 to 15 dB; specifically, for stationary noise, the SNR is standardized across the entire range to -55 to -65 dB to avoid training instability caused by excessively high or low levels. To meet the broader coverage requirements of the claims, a configuration of p < 0.7 and -10 to 20 dB can also be used to encompass weaker and stronger interference scenarios. This design ensures both sample randomness and representativeness, while also allowing for adjustability under different task focuses through parameter control.
[0057] Before speech-noise mixing, spectral domain enhancement and identical scaling are performed on the reverberant speech and noise respectively. Then, the decibel level of at least one signal is uniformized across the entire range of -20 to -35, achieving amplitude consistency normalization. Spectral domain enhancement may include minor perturbations to the time scale, frequency masking, or random bandstops to simulate variations in speaker, device, and environment. Identical scaling and level uniformity ensure alignment of amplitude distribution across different data sources, which is beneficial for maintaining stable convergence of the joint loss across different batches.
[0058] After the above processing is completed, the reverberant speech and the processed noise are linearly superimposed according to the target SNR to obtain the training input. The training target paired with this input is the clean speech after reverberation time reduction processing. The input and target are organized into supervised learning sample pairs in a one-to-one correspondence to ensure that the network output can be directly compared with the target in the frequency domain or time domain. Backpropagation and parameter updates can be performed in the training path using a joint loss including time domain signal-to-noise ratio, speech component, and noise component. Without limiting the protection range, the example parameters in this embodiment are as follows: T60 = 0.15 seconds, p < 0.5, SNR = -5 to 15 dB, the stationary noise dB is uniformly reduced to -55 to -65 dB across the entire range, and at least one signal before mixing is uniformly reduced to -20 to -35 dB. To adapt to different devices and scenarios, the settings can be adjusted as needed within T60∈[0.10,0.60] seconds, p less than or equal to 0.7, and SNR∈[-10,20] dB, to achieve multi-domain training such as "weak reverberation, strong reverberation, extremely low signal-to-noise ratio, and high signal-to-noise ratio", thereby improving the robustness and transferability of the model during edge inference.
[0059] To ensure data quality and parameter implementation, on the one hand, metadata such as T60, p, SNR, and dBfs within a unified range are recorded when each batch of data is generated for offline statistics and monitoring of the training process. On the other hand, after mixing, the amplitude and loudness of the training input are verified a second time to ensure stable level distribution among different samples. In addition, minimum coverage ratios can be set for different noise categories (steady-state, non-steady-state, and impulsive) to avoid class imbalance in the training set and improve the model's ability to suppress non-steady-state interference.
[0060] Appendix Figure 3 The output "training input / target" goes directly into Figure 1 The training path shown, after short-time Fourier transform and feature extraction, is fed into a network containing encoding, enhancement, and decoding stages with skip connections; parameter learning is completed under the influence of joint loss and frequency band weights. After training, Figure 2 The inference path shown performs temporally weighted deep filtering based solely on the current frame and its preceding frames, and performs statistical correction using minimum mean square error post-processing. This robustly transfers the "speech-noise-reverberation" separation capability learned during training to low-latency edge scenarios. The system design of the data pipeline in terms of reverberation targeting, noise sampling and intensity configuration, spectral domain enhancement, and amplitude consistency normalization significantly reduces the distribution offset between training and inference, ensuring the model's performance on the ground.
[0061] The speech denoising model in this embodiment adopts a U-shaped topology of encoding-enhancement-decoding. (See attached...) Figure 4The input shown is a complex time-frequency feature obtained through a short-time Fourier transform, which first enters the "encoding end". From bottom to top, the encoding end consists of Convolution 1 (kernel: 5×2, downsampling), Convolution 2 (kernel: 3×1, downsampling), and Convolution 3 (kernel: 3×1, downsampling), progressively compressing and extracting multi-scale representations along the frequency dimension. The dashed box in the middle represents the enhancement part – the recurrent unit, which uses LSTM / GRU (causal processing) to perform temporal modeling of the high-dimensional features. The "decoding end" consists of Deconvolution 1 (kernel: 3×1, upsampling) and Deconvolution 2 (kernel: 5×2, upsampling), progressively restoring the resolution to the same level as the input along the frequency dimension. The cross-layer arrows pointing from the encoding end to the decoding end represent multiple pairs of skip connections: features of the same scale on the encoding side are bypassed and sent to the decoding side for fusion with the main decoding features, thereby compensating for details and stabilizing training. The network finally outputs a complex-valued mask at the decoding end, which weights the real and imaginary parts of the input complex spectrum (the mask amplitude is limited to [0,1] to ensure stability), completes frequency domain enhancement, and provides input for the subsequent inverse short-time Fourier transform output of time-domain denoised speech.
[0062] Each pair of cross-layer connections is shown in the appendix. Figure 5 As shown, features from the encoding layer (feature dimension: C×H×W) first undergo channel matching / compression in a 1×1 convolution (continuous positive definite) to match the corresponding input channels of the decoding layer. A gated convolution (1×1 convolution + Sigmoid activation) is then connected in parallel on the matched branch to generate channel-wise weights α, where α∈[0,1]. Subsequently, the features enter the top-to-bottom fusion sequence shown in the diagram: first, the "channel-weighted" module uses weights to scale the encoding-side features channel-wise; then, at the "feature addition (⊕)" point, they are fused element-wise with the decoding backbone features of the same scale to obtain the "decoding layer input (dimensionality: C×H×W)". After convolution / upsampling in this layer, the "fused features (dimensionality: C×H×W)" are output. Compared with traditional channel concatenation, the additional... Figure 5 The path shown, “1×1 channel matching → gated weights → channel-weighted → element-wise addition”, significantly reduces the channel stacking and deconvolution computation at the decoding end. At the same time, it adaptively controls the cross-layer information injection intensity through α, ensuring stable fusion, lightweight operation, and strict alignment with the backbone decoding layer.
[0063] The recurrent unit (causal processing) performs temporal enhancement on the top-level encoded features, providing a context-consistent high-level representation for the decoder. The decoder then applies the following... Figure 4 The resolution is restored step-by-step through two-stage upsampling (deconvolution 1: 3×1; deconvolution 2: 5×2). The above decoding layers are horizontally related to the attached layers. Figure 5 The jump connections correspond one-to-one: after the same-scale encoded features are generated by 1×1 channel matching and gated convolution to generate weights α, they are injected into the decoding backbone at the point where the channel weights are added to the features, so that the transition of "decoding layer input → fused features" is consistent with the annotation in the figure.
[0064] To achieve a balance between speech fidelity and noise suppression in the outputs of the backbone and skip connections, this embodiment introduces a joint loss consisting of time-domain and frequency-domain terms during the training phase, and assigns differentiated weights to different frequency bands. The time-domain signal obtained by multiplying the complex-valued mask of the network output with the noisy spectrum and performing an inverse transform is used as... Clean voice with the target Together, they are used for time-domain loss; simultaneously, speech and noise components are modeled separately in the frequency domain. Details are as follows:
[0065] (1) Time-domain negative signal-to-noise ratio loss
[0066]
[0067] (2) Loss of speech components
[0068]
[0069] in Spectral features of clean speech; For predictive features.
[0070] Equation (3) Noise component loss:
[0071]
[0072] in The spectral characteristics of noise; This is a noise prediction feature.
[0073] Equation (4) Total loss function and weighting strategy:
[0074]
[0075] in Time-domain negative signal-to-noise ratio loss, These are weighting coefficients. Different values are set according to frequency bands to enhance low-frequency noise suppression while maintaining mid-to-high frequency speech details. Different values are set for different frequency bands. This primarily enhances low-frequency noise suppression, improving noise reduction for low-frequency noise such as wind noise. The frequency band weights are shown in the table below:
[0076] The output complex mask is obtained by multiplying with the noisy spectrum and performing an inverse transform. Substitute into equation (1) to calculate the time domain term; simultaneously recover the frequency domain estimate from the network output ( )and( ) respectively with the target ( )and( Substitute these values into equations (2) and (3) to calculate the frequency domain terms. When calculating equations (2) and (3), divide the spectrum into the four sub-bands mentioned above, calculate the MSE in each sub-band, and then multiply it by the corresponding... (voice) or (Noise) is aggregated and finally substituted into equation (4) along with equation (1) to obtain the total loss, thereby achieving enhanced suppression of low-frequency wind noise, etc. In this way, the constraint of "low frequency is heavier" is explicitly applied to the output after gating and fusion, while reducing speech distortion.
[0077] After optimizing the output mask using "causal time-weighted deep filtering", a minimum mean square error post-processing (deepmmse) is introduced to statistically correct the mask. The calculation order is as follows:
[0078] B1. Input quantities (from the preceding network and STFT): The real and imaginary parts of the spectrum of the noisy speech are respectively... The real and imaginary parts of the prediction mask of the model (after causal deep filtering) are respectively
[0079] B2. Complex-valued mask estimation of clean speech spectrum:
[0080]
[0081]
[0082] B3. Power Spectrum and Signal-to-Noise Ratio Calculation:
[0083] Noisy power spectrum:
[0084] Estimating the clean power spectrum:
[0085] Estimating the noise power spectrum:
[0086] Prior signal-to-noise ratio:
[0087] Posterior signal-to-noise ratio:
[0088] B4. Frequency gain and mask optimization:
[0089] Gain calculation:
[0090]
[0091] Optimized mask:
[0092]
[0093]
[0094] B5. Output Formation: The optimized mask is multiplied point-by-point with the noisy spectrum to obtain the final enhanced spectrum, which is then output as time-domain denoised speech via ISTFT. This statistical correction relies only on the current frame and its preceding information, forming a two-level inference optimization of "causality-statistics" together with causal deep filtering.
[0095] This invention weights the real and imaginary parts of complex time-frequency features in the frequency domain and enhances them using a complex-valued mask. The mask amplitude is limited to [0,1] to ensure stability and physical plausibility. To further improve the stability of the background after noise reduction and meet the low-latency constraints of real-time performance at the edge, this invention uses complex-valued masks in conjunction with deep filtering: the network first provides an initial complex-valued mask, then performs causal deep filtering to perform temporally weighted fusion of the mask during the inference phase, and finally uses minimum mean square error post-processing to perform statistical correction at frequency points. The entire mask optimization process only uses information from the current frame and its preceding frames, without using future frames, ensuring that real-time constraints are met.
[0096] To clearly illustrate the areas for improvement, refer to Figure 6 The improved depth filtering diagrams illustrate three scenarios:
[0097] Figure 6 (a) Traditional masking scheme
[0098] Input: A single frequency point (green square) in the time-frequency grid, corresponding to the energy at a certain time and frequency.
[0099] Processing: Perform direct "point-to-point" scaling on this frequency point without explicitly utilizing information from adjacent frames.
[0100] Output: Only the scaled energy at this frequency point is retained, which is insufficient to suppress non-steady-state noise and reverberation, and the background residue is prone to fluctuation over time.
[0101] Figure 6 (b) Complete depth filtering scheme
[0102] Input: The time and frequency regions that cover the current frame and several frames before and after it (the example is a small window composed of several frames before and after and several adjacent frequency points).
[0103] Processing: Assign weights to each frequency point within the region and perform weighted fusion to output a joint estimate of the current frequency point.
[0104] Output: Noise is suppressed by utilizing redundant information from preceding and following frames, but because subsequent frames are used, additional latency is introduced, making it unsuitable for real-time scenarios on the edge.
[0105] Figure 6(c) The improved depth filtering scheme of the present invention
[0106] Input: The time and frequency regions covering the current frame and several previous frames, excluding any future frames.
[0107] Processing: Weighted fusion is performed only in the "current + previous" direction to make the fusion process causal and avoid the uncertainty and delay caused by future frames not yet being available.
[0108] Output: Mask optimization results obtained solely from historical and current information, suitable for low-latency scenarios such as real-time calls and end-to-end streaming enhancement, which can significantly improve background stability and residual noise suppression without sacrificing real-time performance.
[0109] and Figure 2 As shown in the inference link alignment, the connection and constraints of the improved deep filtering of this invention in the system are as follows:
[0110] Placement: The improved deep filtering is placed after the initial complex value mask of the network output and before the minimum mean square error post-processing. The overall sequence is: "Network estimation → Causal deep filtering (improved scheme) → Minimum mean square error statistical correction → Inverse short time Fourier transform output".
[0111] Causality and normalization: The weights used for fusion come only from the current frame and the previous frame, and do not include any future frames; the fusion weights for each frequency point are non-negative and normalized at each frequency point to prevent abnormal energy amplification and ensure timing stability.
[0112] Neighborhood morphology and parameters: The time dimension uses a historical window with a length of "current + several previous frames"; the frequency dimension uses symmetrical adjacent frequency bands. The above window only covers the previous and current frames, effectively avoiding interference from future noise mutations on the current estimate, and the parameters are selected as small integers to meet the constraints of edge computing power and latency.
[0113] Complex value consistency: Causal fusion directly applies to the complex value mask (maintaining consistency in the weighting of the real and imaginary parts), and the fused mask is still limited to [0,1], which is naturally compatible with subsequent frequency point statistical gain calculations.
[0114] In conjunction with minimum mean square error post-processing: The mask obtained by improved depth filtering is first multiplied by the noisy complex spectrum to obtain the estimated spectrum, and then the power spectrum, prior and posterior signal-to-noise ratios, and frequency point gains are calculated, and the mask is statistically corrected accordingly. Causal fusion improves the timing consistency of the mask, making subsequent frequency point gain estimation smoother and residual noise suppression more sufficient.
[0115] A Bluetooth headset based on the method of this invention includes a microphone pickup and analog-to-digital converter module, a voice processing chip, a memory, a Bluetooth communication and audio output module, a display unit (or an equivalent prompt unit), and a power management module. The microphone, after front-end amplification and anti-aliasing, enters the analog-to-digital converter and is written into a small circular buffer, retaining only the current frame and its preceding frames to satisfy causal constraints. The voice processing chip performs short-time Fourier transform, network inference, causal deep filtering, minimum mean square error post-processing, and inverse transform in a fixed sequence. The resulting denoised voice is sent to the local digital-to-analog converter and power amplifier to drive the speaker unit, and also to the Bluetooth protocol stack for external transmission. The processing link does not use any subsequent frames, ensuring low latency and timing stability at the end side. The buffer and intermediate storage only cover the causal window and the necessary complex spectrum, mask, and weight tensor.
[0116] Network inference is implemented within the speech processing chip using an encoding-enhancement-decoding structure: The encoding end employs downsampling with convolutional kernel sizes of 5×2 (input layer only), 3×1, and 3×1 to extract multi-scale features; the enhancement unit uses a grouped recurrent structure, with each group independently performing causal computation followed by instantaneous normalization and fusion; the decoding end uses deconvolutions of 3×1 and 5×2 (last layer only) for progressive upsampling. Cross-layer skip connections first use 1×1 channel matching on the encoding side, followed by 1×1 gated convolution and sigmoid activation to generate channel-wise weights α (0-1). Subsequently, the encoded features are channel-weighted and element-wise added to the decoded features of the same scale. The network output complex mask weights the real and imaginary parts of the input complex spectrum separately, with amplitudes limited to 0-1. Causal depth filtering uses only the neighboring time and neighboring frequency information of the current frame and the previous frame to perform weighted fusion of the complex-valued mask. The fusion weight is non-negative and normalized to one at each frequency point. Then, the prior and posterior signal-to-noise ratios are calculated based on the noisy and estimated power spectra to obtain the frequency point gain. The mask is statistically corrected, and the corrected spectrum is inversely transformed to generate time-domain enhanced speech and output.
[0117] In the absence of a network connection, this device performs local speech recognition on the denoised speech to generate source language text. The recognized incremental text is fed into a locally stored neural machine translation model for offline translation, outputting target language text. Then, a local speech synthesis model offline synthesizes the target language text into target language speech and outputs it to the speaker unit. The source and target language texts can be displayed on the display unit and / or transmitted via the communication unit for prompting at the peer or local end. The aforementioned recognition, translation, and synthesis models are all stored in the device's rewritable memory and perform local inference on the processor of the speech processing chip. Their scheduling priority is lower than the main denoising link. A small circular buffer and hierarchical scheduling are used to ensure the real-time performance of the main link. When computing power or power consumption is limited, translation and synthesis can be downgraded or paused according to a strategy without changing the main link's order and causal constraints. Model parameters are obtained externally from the device according to the training path and then written into the device. Upgrades are performed wirelessly, replacing only the parameters and strategies without changing the fixed order of "short-time Fourier transform → network estimation → causal deep filtering → minimum mean square error post-processing → inverse transform" or the technical limitations of "using only the current frame and previous frames, non-negative weights and frequency-normalized weights, and complex value mask amplitude limited to 0 to 1". This allows for the realization of a local closed loop of noise reduction / reverberation reduction and offline recognition-translation-synthesis within the conventional hardware framework of a Bluetooth headset for calls.
[0118] While maintaining the main "noise reduction / reverberation reduction" link unchanged, the Bluetooth headset adds a human-computer interaction closed loop for local speech recognition, neural machine translation, speech synthesis, and display. Specifically, the enhanced speech obtained from the inverse transform is branched into a parallel streaming branch within the speech processing chip, input to the local recognition module at the same sampling rate, frame length, and step size as the main link; the recognition module outputs text segments and time stamps incrementally, which are then sent to the translation module for block-level translation, outputting incremental text in the target language. The synthesis module receives the incremental text and generates corresponding speech segments; the human-computer interaction module displays the recognized / translated text via the screen or indicator lights, or plays the synthesized speech locally as a prompt tone. This branch only consumes the enhanced audio and text, without reverse-engineering the processing order or causal constraints of the main link.
[0119] To ensure feasibility and real-time performance on the edge, a unified streaming buffer and scheduling mechanism is deployed within the chip. Enhanced speech input is fed into the recognition input buffer, which only covers the current frame and a small number of previous frames, without caching subsequent frames. Incremental text output from recognition enters the translation input buffer, and the translation output enters the synthesis input buffer. All three buffers employ a first-in-first-out (FIFO) approach with capacity limits and a backpressure strategy to prevent branch blockage from affecting the main link. Task scheduling uses a priority hierarchy: the main link has the highest priority, followed by recognition, then translation and synthesis. When frame-level latency or chip temperature rise approaches a threshold, branch degradation is triggered sequentially—synthesis switches from high-quality to lightweight, translation adopts a faster decoding strategy, and recognition reduces the frequency of auxiliary function calls, retaining only text recognition and display when necessary. All models and operating strategies are stored in rewritable storage, supporting over-the-air updates.
[0120] In terms of interface and presentation, enhanced speech is output to the recognition module with a fixed frame configuration. Recognition generates incremental text based on silence detection or fixed segment lengths while maintaining temporal continuity. Translation and recognition are aligned at incremental boundaries, updating smoothly with short text granularity to avoid display jitter. Synthesis outputs the target speech in segments, which can be played locally or overlaid with sidetone paths. All audio is synchronized within the same clock domain to avoid stuttering and crosstalk. Display and prompts have lower priority than call audio and will not interrupt uplink speech under any circumstances. Regarding privacy and robustness, recognition, translation, and synthesis all run locally offline and do not rely on external networks.
[0121] When Bluetooth headsets are used for translation, a "local translation branch" is set up after the enhanced main link on the headset. The enhanced speech is first converted into incremental text by the local recognition module, and then enters the neural machine translation module. This translation module automatically determines the direction between the source and target languages and outputs the target language text in real time. To adapt to the computing power of the device, the model adopts low-bit integer quantization and symmetric scaling coefficients, and performs linear dequantization during inference, significantly reducing storage bandwidth and computing power consumption. After quantization, the model is converted into a terminal-specific inference format, which includes binary weights, a unified sub-vocabulary, and model configuration, all of which are loaded by the local inference engine. Translation inference runs on a general-purpose processor, using a small beamwidth search combined with a repetition suppression strategy. Threads are allocated according to a predetermined proportion of available processing resources to ensure a balance between real-time performance and energy consumption. In the decoding stage, an early stopping criterion and a single-step generation waiting delay threshold are set to keep the end-to-end latency of simultaneous interpretation controllable.
[0122] Regarding multilingual support, this embodiment employs a shared sub-word segmenter jointly trained for at least two languages, constructing a unified vocabulary and performing unified encoding and decoding for both the source and target languages. The shared segmenter automatically learns the sub-word set and mapping relationships from the mixed corpus. The unified vocabulary includes necessary functional tags and sets vocabulary size and character coverage parameters to simultaneously cover different writing systems and reduce the proportion of out-of-vocabulary symbols. When a new language is added, expansion is achieved by updating the mixed corpus and jointly or sequentially retraining the segmenter and translation model. During expansion, existing sub-word indices are frozen, and newly added sub-words are placed in the expansion area, thus maintaining backward compatibility with existing models. If out-of-vocabulary symbols still appear during the inference phase, decoding continues at a finer granular level according to the fallback rules, ensuring continuous and stable cross-language output. The aforementioned local translation branch connects with the previous enhancement, recognition, synthesis, and display modules via standardized interfaces, meeting the requirements of real-time, low-power, and iteratively upgradeable implementation on the edge.
[0123] The present invention is applied in the following scenarios:
[0124] Scenario 1 (Vehicle-mounted, predominantly low-frequency noise): When a vehicle is in motion, the microphone picks up low-frequency noise such as engine noise and road noise. The device processes the noise in a fixed sequence: Short-Time Fourier Transform → Network Estimation → Causal Deep Filtering → Least Mean Square Error Post-processing → Inverse Transform. The complex-valued mask of the network output is limited to 0-1; then, causal fusion is performed using only the current frame and the previous frame, with weights at each frequency point being non-negative and normalized, focusing on low frequencies to suppress continuous low-frequency booming while minimizing the weakening of the speaker's low-frequency energy. The final statistical post-processing provides frequency point gains based on the strength of noise and speech, further refining the noise suppression; for instantaneous changes such as starting or acceleration, post-processing is only appropriately strengthened in the low-frequency range, maintaining overall latency and resulting in a smoother background.
[0125] Scenario 2 (Conference room, medium reverberation and multi-person conversation): Indoor reflections can cause reverberation trailing and ghosting, and multi-person conversations can easily interfere with each other. The device operates on the edge in the same sequence: first, the network performs initial enhancement (mask 0-1), then causal deep filtering uses weighted fusion of information from the current frame and previous frames within nearby time and frequency ranges, directly shortening the reverberation trail without using future frames. Subsequently, statistical post-processing further provides gain at each frequency point, suppressing residual room echoes and fluctuations in background voices. If seat movement or head turning causes abrupt reflection changes, cross-layer gating weights automatically reduce the injection intensity, preventing amplification of sudden reflections as useful details, thus ensuring that the call remains clear and natural.
[0126] Scenario 3 (Open Office Area, Keyboard Typing and Background Voice): This scenario presents common open office environments with keyboard typing (impulse noise) and continuous background voice. The invention processes speech continuously in a fixed order throughout the call, periodically checking processing time and buffer usage to ensure smooth operation. Causal deep filtering utilizes information from historical frames to help the current frame "see clearly" the speech, suppressing both impulse and steady-state background noise. Statistical post-processing further refines the signal at the frequency level, reducing the rise of background voice over extended periods. To balance power consumption, the device prioritizes strong post-processing in low and key mid-frequency ranges without altering the processing order or causal constraints, while operating at normal intensity in other frequency bands. It automatically reduces post-processing frequency during silent segments, returning to full processing upon resumption of speech. This ensures sustained speech intelligibility and background stability even in complex office environments.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech denoising method based on inference optimization, characterized in that, Includes the following steps: Step S1: Perform a short-time Fourier transform on the noisy speech to obtain time-frequency features; Step S2: Input the features into a U-shaped network that includes encoding, enhancement, and decoding and has skip connections to obtain an initial mask; Step S3: Optimize the initial mask based on causal time-weighted depth filtering that only utilizes the time-frequency information of the current frame and its preceding frames to obtain an optimized mask; The minimum mean square error post-processing is adopted. The prior and posterior signal-to-noise ratios are estimated based on the power spectra of the noisy speech and the estimated clean speech. The gain is calculated to correct the optimized mask and obtain the final mask. Step S4: Weight the noisy spectrum with the final mask and output the denoised speech through inverse short-time Fourier transform.
2. The method according to claim 1, characterized in that, The data processing in step S1 includes: acquiring room impulse response, clean speech, and noise data; adjusting the reverberation duration of the room impulse response to maintain attenuation characteristics; convolving the clean speech with the adjusted room impulse response to obtain reverberant speech; and, based on preset noise selection and intensity configuration, enhancing and normalizing the amplitude consistency of the reverberant speech and noise in the spectral domain, and mixing them according to a preset signal-to-noise ratio strategy to obtain training input, with the clean speech serving as the corresponding training target. The reverberation time of the room's impulse response is adjusted so that the target reverberation time T60 is between 0.10 seconds and 0.60 seconds, while maintaining its original decay pattern; Noise is selected with a probability of p < 0.
7. One or more noise sources are selected, and the signal-to-noise ratio of the training samples is set in the range of -10 to +20 dB. Enhancement processing and amplitude consistency normalization are performed on the reverberant speech and / or the noise in the spectral domain, and the level of at least one signal is normalized to a predetermined range; the processed signals are mixed according to the signal-to-noise ratio strategy to form a training input.
3. The method according to claim 1, characterized in that, The neural network construction in step S2 satisfies the following constraints: the encoding part downsamples and extracts the input spectral features; the high-dimensional features are calculated separately by the grouped recurrent neural network of the enhancement part and then normalized and fused; the decoding part upsamples and reconstructs the enhanced features; a skip connection is set between encoding and decoding, which performs feature transformation in the channel dimension and passes the features by feature addition; wherein, the encoding part and the decoding part are each composed of multiple levels of convolutional layers and deconvolutional layers, except for the first level of convolutional layer in the encoding part and the last level of deconvolutional layer in the decoding part, the remaining levels are all composed of multiple levels of convolutional layers and deconvolutional layers. The size of both the convolutional kernel and the deconvolutional kernel is 3×1; the size of the convolutional kernel of the first-level convolutional layer of the encoding part and the last-level deconvolutional layer of the decoding part is 5×2; the enhancement part adopts a grouped recurrent neural network, which divides the input features into two groups in the channel dimension and configures corresponding recurrent units for independent calculation. The outputs of each group are normalized by an instantaneous normalization layer and then fused; a skip connection is set between the encoding part and the decoding part. The skip connection performs channel matching through 1×1 convolution and fuses the matched encoding side features and decoding side features by feature addition.
4. The method according to claim 3, characterized in that, The decoding part uses sub-pixel convolution for upsampling, and rearranges the resolution level by level along the frequency dimension at integer multiples until it is consistent with the frequency resolution before encoding. The features of the encoding part are first matched by 1×1 convolution. The matched features are then fed into a gated convolution branch with a sigmoid activation function to generate channel-wise weights. The encoded features are then weighted using these weights and fused with the decoded features of the corresponding layer by element-wise addition. The output of the enhancement part is normalized for each time frame and each frequency point in the instantaneous normalization layer before being fused; the two sets of cyclic units of the enhancement part only receive the temporal features of the current frame and its preceding frame in a causal manner for calculation, and the two sets of parameters do not share each other; the output of the U-shaped network is a complex-valued mask, which weights the real part and the imaginary part of the input spectrum respectively, and the amplitude of the mask is limited to the range of 0 to 1.
5. The method according to claim 1, characterized in that, Step S3 includes: performing a short-time Fourier transform on the training input to obtain spectral features and inputting them into the U-shaped network to obtain a mask; using a joint loss function composed of a time-domain signal-to-noise ratio term, a speech component term, and a noise component term, setting weighting coefficients for different frequency sub-bands to distinguish processing priorities, and performing backpropagation based on the joint loss to update network parameters; performing a short-time Fourier transform on the training input to obtain complex spectral features and inputting them into the U-shaped network to obtain a mask; The joint loss consists of one or more of the following three terms: a time-domain signal-to-noise ratio (SNR) term, used to improve the SNR of the estimated speech relative to the target clean speech; a speech component term, used to constrain the estimated speech to approximate the target clean speech in amplitude and / or phase; and a noise component term, used to penalize residual noise energy in the estimated speech. The frequency range is divided into multiple sub-bands, and weight coefficients are assigned to each sub-band. These weights are preset or learnable and can be adaptively adjusted according to training rounds, sample SNR, or noise steady-state characteristics, ensuring that the weight of low-frequency sub-bands is not lower than that of mid-to-high-frequency sub-bands and / or assigning higher weights to sub-bands with prominent non-steady-state noise. Based on the joint loss, a gradient descent-type optimization method is used for backpropagation to update the network parameters.
6. The method according to claim 1, characterized in that, Step S4 includes: performing a short-time Fourier transform on the noisy speech to obtain complex spectral features, and inputting them into the U-shaped network to obtain an initial mask; employing a time-weighted deep filtering strategy, using only the time-frequency information of the current frame and its preceding frames at the current frequency and its adjacent frequency points, fusing them according to non-negative and normalized weighting coefficients to obtain an optimized mask, without using any subsequent frame information; using post-processing based on minimum mean square error to correct the optimized mask: calculating the posterior signal-to-noise ratio γ based on the power spectra of the noisy speech and the estimated clean speech, estimating the prior signal-to-noise ratio ξ in a decision-oriented manner, and obtaining the gain according to a predetermined frequency gain formula to correct the mask; multiplying the corrected mask by the noisy spectrum to obtain the denoised spectrum, and outputting the time-domain denoised speech through an inverse short-time Fourier transform.
7. The method according to claim 6, characterized in that: The number of back-view frames L in the temporal weighted depth filtering is a positive integer; the neighborhood width K of adjacent frequency points is a positive integer; the weighting coefficients are normalized frequency-by-frequency point and can be generated by gated convolution or attention mapping, and depend only on the information of the current frame and its preceding frames; the frequency point gain formula of the minimum mean square error post-processing is any one of Wiener gain, logarithmic amplitude minimum mean square error gain, or spectral amplitude minimum mean square error gain; the prior signal-to-noise ratio ξ is obtained by exponentially smoothing the estimation result of the previous time step with the current posterior signal-to-noise ratio using a decision-oriented method, with the smoothing factor valued between 0 and 1; the noise power spectrum is estimated using any one of online noise tracking or speech presence probability methods.
8. A Bluetooth headset, comprising the method described in any one of claims 1-7, including: The system includes a pickup unit, an audio processing unit, a memory, a Bluetooth communication unit, and a speaker unit. The audio processing unit is configured to perform the following inference noise reduction process on the pickup signal: performing a short-time Fourier transform on the noisy speech to obtain spectral features. The system calls the neural network for encoding, enhancement, and decoding to output an initial mask; generates an optimized mask based on temporal weighted depth filtering using only the information of the current frame and its preceding frames; calculates the prior / posterior signal-to-noise ratio based on minimum mean square error post-processing and corrects the mask according to the gain; multiplies the corrected mask by the noisy spectrum and performs an inverse short-time Fourier transform to obtain the denoised speech, which is then output to the speaker unit and / or transmitted via Bluetooth communication unit. Its features include: performing inference noise reduction on the picked-up signal to obtain denoised speech; performing local speech recognition on the denoised speech to generate source language text in the absence of a network connection; using a locally stored neural machine translation model to perform offline translation of the source language text to obtain target language text; using a locally stored speech synthesis model to offline synthesize the target language text into target language speech and outputting it to the speaker unit; presenting the source language text and target language text on the display unit and / or sending them through the communication unit; wherein the speech recognition, translation, and speech synthesis models are all stored in the memory and inferred locally on the processor.
9. A Bluetooth headset according to claim 8, characterized in that, The neural machine translation model automatically determines the translation direction and generates target language text in real time. The machine translation model is subjected to low-bit integer quantization, symmetrical scaling factor is used, and linear inverse quantization is performed during the inference stage to reduce storage bandwidth and computing power consumption. The quantized model is converted into a terminal-specific inference format, which includes at least binary weights, a unified sub-word vocabulary, and model configuration. The model is run on a general-purpose processor on a local inference engine, employing a small beamwidth search combined with a repetition suppression strategy, and allocating threads according to a predetermined proportion of available processing resources to achieve a balance between real-time performance and energy consumption. Configure early stop criteria and set single-step generation waiting delay threshold during the decoding stage to ensure that the end-to-end delay of simultaneous interpretation is controllable. The input speech is incrementally decoded using streaming coding and attention caching mechanisms, with fixed or adaptive windows, and look-ahead and stabilization windows configured to reduce short-term writeback and output flicker. Decoding degradation is performed when the observed end-to-end delay exceeds a threshold, including bundle width reduction or greedy decoding; When the direction confidence level is lower than the threshold, manual confirmation or default routing is triggered to ensure output controllability; Memory and power management strategies are set up to reduce peak resource consumption through weighted segmented loading, cache reuse, and low-power scheduling, and unknown sub-word backoff rules are used to handle symbols outside the word list to maintain decoding continuity.
10. A Bluetooth headset according to claim 9, characterized in that, The neural machine translation model employs a shared word segmenter jointly trained for at least two languages to construct a unified vocabulary, and uses this unified vocabulary to uniformly encode and decode the source and target languages; wherein: The shared word segmenter is based on a statistical word segmentation algorithm, which automatically learns the set of word units and their corresponding mappings from mixed corpora containing multilingual text; The unified vocabulary includes functional tags and sets vocabulary size and character coverage parameters to simultaneously cover language elements from different writing systems and reduce out-of-vocabulary symbols; The addition of new languages is achieved by updating the mixed corpus and jointly or sequentially retraining the word segmenter and translation model; during the expansion process, the existing sub-word indexes are frozen and retained and the new sub-words are placed in the expansion area to maintain backward compatibility with the existing model; During the inference phase, unregistered symbols are processed using rules that fall back to finer-grained units to ensure the continuity and stability of cross-language decoding.
Citation Information
Patent Citations
A method for far-field speech dereverberation using the U-NET structure of CNN
CN109949821B
Speech enhancement methods, devices, equipment and media
CN112767959B
Full-band speech enhancement method based on spectrum compression and self-attention neural network
CN115273885A
Multi-sound-source automatic gain control method based on sound field perception
CN116543784A
Beam forming method based on dynamic reconstruction noise covariance
CN118136039A