A speech intelligibility enhancement method and system based on sound field noise cognition

CN122598674BActive Publication Date: 2026-09-11GUANGZHOU LANDE ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611015127.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-11
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

[0003]针对现有技术存在的因无法区分混响与噪声而导致语音清晰度与自然度受损的问题,本申请通过一种基于声场噪声认知的语音清晰度增强方法及系统,实现混响与背景噪声的双通道解耦认知与选择性增强

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598674B_ABST
    Figure CN122598674B_ABST
Patent Text Reader

Abstract

This application relates to the fields of speech signal processing and intelligent acoustics, and provides a speech intelligibility enhancement method and system based on sound field noise cognition. The method includes: acquiring prior room acoustic information of the target sound field and the current input signal; estimating the reverberation energy spectrum and performing pre-whitening processing based on the prior room acoustic information through a first cognitive channel to obtain a pre-whitened signal; estimating background noise on the pre-whitened signal through a second cognitive channel to obtain a background noise power spectrum; generating simulated acoustic samples based on the prior room acoustic information and adaptively fine-tuning the probability estimation model to output the probability of clean speech presence and the probability of late reverberation dominance; generating a hybrid gain mask based on the probability of clean speech presence and the probability of late reverberation dominance, and combining it with the background noise power spectrum to enhance the current input signal to obtain an enhanced signal. This application achieves decoupled cognition and selective enhancement of reverberation and noise, improving speech intelligibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of speech signal processing and intelligent acoustics, specifically to a speech intelligibility enhancement method and system based on sound field noise cognition. Background Technology

[0002] In engineering practices such as building automation, intelligent installation, and conference control, microphones are often fixed to ceilings or walls due to structural limitations, resulting in low direct sound energy and a high proportion of room reverberation in the acquired signal. Existing speech enhancement methods typically treat reverberation and background noise as similar interference and suppress them uniformly. Without on-site professional acoustic calibration, algorithms are prone to misclassifying high-energy late-stage reverberation as diffuse field noise, triggering excessive déreverberation processing. While this reduces background noise, it causes the speech to sound dry, fragmented, or even broken, severely impairing speech clarity and naturalness. The root cause of this problem lies in the lack of an effective mechanism to distinguish between reverberation and noise. Reverberation is a deterministic acoustic process with exponential decay, while background noise is random interference unrelated to room acoustics. Their physical causes and statistical characteristics differ, and conflating them inevitably leads to inappropriate processing strategies. Therefore, there is an urgent need for a technical solution that can automatically distinguish between reverberation and noise and accordingly achieve refined speech clarity enhancement. Summary of the Invention

[0003] To address the problem that existing technologies suffer from impaired speech clarity and naturalness due to the inability to distinguish between reverberation and noise, this application proposes a speech clarity enhancement method and system based on sound field noise cognition, which achieves dual-channel decoupled cognition and selective enhancement of reverberation and background noise.

[0004] To achieve the above objectives, this application adopts the following technical solution: A speech intelligibility enhancement method based on sound field noise cognition includes: Acquire the room acoustic prior information and current input signal of the target sound field; Based on prior acoustic information of the room, the reverberation energy spectrum of the current input signal is estimated through the first cognitive channel, and the current input signal is pre-whitened based on the reverberation energy spectrum to obtain the pre-whitened signal; Background noise power spectrum is obtained by estimating the background noise of the pre-whitened signal through the second cognitive channel; Simulated acoustic samples are generated based on prior room acoustic information, and a pre-deployed probability estimation model is adaptively fine-tuned based on the simulated acoustic samples. The fine-tuned probability estimation model outputs the probability of clean speech and the probability of late reverberation dominance. A hybrid gain mask is generated based on the probability of the presence of clean speech and the probability of late reverberation dominance. The current input signal is then enhanced using the hybrid gain mask and the background noise power spectrum to obtain the enhanced signal.

[0005] The above scheme establishes a dual-channel decoupled cognitive model of reverberation and noise, uses prior information on room acoustics to accurately guide the separation of reverberation energy estimation and noise estimation, and combines a probabilistic model with on-site adaptive fine-tuning to generate soft decision gain. This solves the problem of speech distortion caused by confusion between the two types of interference in traditional methods while preserving beneficial early reflections.

[0006] As one implementation method, the pre-whitening process of the current input signal based on the reverberation energy spectrum includes: The transfer function is calculated based on the amplitude spectrum and reverberation energy spectrum of the current input signal, and a time-varying pre-whitening filter is constructed based on the transfer function; wherein, the transfer function is determined based on the difference between the amplitude spectrum and the reverberation energy spectrum, a preset safety factor, and a minimum gain lower limit; The current input signal is filtered by a time-varying pre-whitening filter to obtain a pre-whitening signal.

[0007] This implementation method, by introducing a safety factor and the calculation of the transfer function constrained by the minimum gain lower limit, retains a processing margin while deducting the reverberation energy that conforms to the exponential decay law, preventing signal distortion caused by excessive attenuation, and ensuring that the pre-whitened signal is statistically similar to a clean signal without reverberation plus background noise, thus laying the foundation for subsequent high-precision noise estimation.

[0008] As one implementation, estimating the reverberation energy spectrum of the current input signal through the first cognitive channel includes: Based on prior acoustic information of the room, the reverberation energy spectrum is calculated using a recursive formula. : ; in, For the current frame, For frequency points, This is the reverberation energy spectrum of the previous frame. This represents the frequency-dependent sound pressure attenuation coefficient from the room's acoustic prior information. The number of frame shift sampling points. The system sampling rate, For the current input signal at frequency point The amplitude spectrum at the location; The background noise estimation of the pre-whitened signal through the second cognitive channel includes: By using a minimum-controlled recursive averaging algorithm, the minimum value of each frequency point is smoothly searched on the spectrum of the pre-whitened signal and recursively smoothed to obtain the background noise power spectrum.

[0009] This implementation uses the physically measured sound pressure attenuation coefficient to accurately determine the smoothing factor in the recursive formula, so that the reverberation energy estimation strictly follows the physical law of exponential decay and satisfies all boundary conditions such as strong sound absorption, long reverberation and silent segment; combined with the pre-whitening process to remove the reverberation tail, the minimum value control recursive averaging algorithm can truly reflect the power level of irrelevant background noise and avoid reverberation contaminating the noise estimation.

[0010] As one implementation method, the adaptive fine-tuning of the pre-deployed probability estimation model based on simulated acoustic samples includes: Obtain clean speech samples and background noise samples; Based on the early reflection spectrum and late reverberation model in the prior information of room acoustics, the clean speech samples are convolved and superimposed with the reverberation tail, and combined with the background noise samples to generate simulated acoustic samples. Based on simulated acoustic samples, the weights of the probability estimation model are sampled and updated in multiple batches using a meta-learning algorithm to obtain the adapted model weights; the probability estimation model is a deep neural network model. The probability estimation model is fine-tuned by adapting the model weights to obtain the fine-tuned probability estimation model.

[0011] This implementation method utilizes prior acoustic information of the room to synthesize simulated samples on-site, generating a large amount of labeled training data without collecting real speech. Through meta-learning algorithms, the general model can quickly adapt to the specific acoustic environment of the target sound field within seconds, becoming an acoustic expert for that room, while avoiding the risk of privacy leakage.

[0012] As one implementation method, the step of outputting the probability of clean speech presence and the probability of late reverberation dominance through the fine-tuned probability estimation model includes: The current input signal is subjected to spectral and spatial feature extraction to obtain signal features; The signal features are input into the fine-tuned probability estimation model. The output layer of the fine-tuned probability estimation model has two parallel Sigmoid activation units, which output the probability of clean speech presence and the probability of late reverberation dominance within a preset value range, respectively. The probability of clean speech presence represents the probability of the presence of direct sound and early reflection components. The sum of the probability of pure speech presence, the probability of late reverberation dominance, and the probability of noise presence determined based on the background noise power spectrum satisfies the preset normalization constraint.

[0013] This implementation uses a soft decision mechanism with dual-dimensional probability output to provide a complete probabilistic description of the sound field components. The normalization constraint ensures that the sum of the probabilities of each component is constant, enabling the subsequent gain mask generation to selectively process based on the precise proportion of each component, rather than filtering them out indiscriminately.

[0014] As one implementation method, generating a hybrid gain mask based on the probability of pure speech presence and the probability of late reverberation dominance includes: Obtain the signal-to-mixing ratio estimate for the current time-frequency cell; The initial gain is obtained by exponentially weighting the probability of pure speech presence and the probability of late reverberation dominance. The exponent of the late reverberation dominance probability is adaptively adjusted based on the signal-to-mixing ratio estimate, and the initial gain is recalculated to obtain the adjusted gain. A hybrid gain mask is obtained by performing asymmetric recursive smoothing on the adjusted gain based on the time axis; in the execution of asymmetric recursive smoothing, the start-up time of gain increase is shorter than the release time of gain decrease.

[0015] This implementation adaptively adjusts the reverberation suppression intensity through the signal-to-mixing ratio to prevent reverberation from masking speech in low signal-to-mixing ratio sections; asymmetric time smoothing causes the gain to rise rapidly at the beginning of the speech to preserve key phonemes, and to fall slowly in the decay section to preserve a natural reverberation transition, ensuring a smooth and uninterrupted listening experience.

[0016] As one implementation, the enhancement processing of the current input signal by using a hybrid gain mask and background noise power spectrum includes: Obtain a language acoustic feature enhancement template that matches the current language; the language acoustic feature enhancement template includes preset weights for the frequency bands of key phonemes in the current language; The current input signal is enhanced using a hybrid gain mask to obtain the initial enhanced signal. The global masking threshold of the initial enhancement signal is calculated based on the background noise power spectrum and a psychoacoustic model. Calculate the amplitude difference between the global masking threshold and the initial enhanced signal, and calculate the sharpness compensation gain based on the speech acoustic feature enhancement template; The initial enhanced signal is compensated for intelligibility based on the intelligibility compensation gain, and then the amplitude is limited by a preset listening comfort threshold to obtain the enhanced signal.

[0017] This implementation elevates speech enhancement from the signal processing level to the auditory perception level, compensating only frequency bands that are masked by noise and contribute significantly to clarity. Combined with dynamic amplitude limiting of the auditory loudness model, it prevents harsh overshoot in consonant bursts, thereby achieving personalized perceptual enhancement specific to language.

[0018] As one implementation method, obtaining the room acoustic prior information of the target sound field includes: A nonlinear logarithmic sweep frequency signal containing a sinusoidal frequency modulation term is emitted into the target sound field, and multi-channel synchronous acquisition is performed through a sub-band time-division strategy to obtain the sweep frequency response signal; A reference signal containing an inverse model of harmonic distortion is constructed, and the swept frequency response signal is deconvolved based on the reference signal to obtain the room impulse response; Based on the atomic norm minimization method, the room impulse response is sparsely reconstructed in the continuous time delay domain, separating the early reflection spectrum and the late reverberation model, and labeled as the room acoustic prior information; the early reflection spectrum consists of multiple time delay-amplitude-direction reflection peaks, and the late reverberation model includes frequency-related reverberation time or sound pressure attenuation coefficient.

[0019] This implementation robustly acquires high-precision room impulse responses in complex engineering environments through specially designed nonlinear sweep frequency signals and harmonic distortion-tolerant deconvolution; the atomic norm minimization framework structurally decomposes the impulse response into early reflections that positively contribute to sharpness and a late reverberation model that characterizes the decay rate, providing accurate physical priors for subsequent dual-channel cognition.

[0020] As one implementation method, it also includes: The system continuously calculates multidimensional internal features in the background, including reverberation prediction bias, noise estimation abruptness rate, probabilistic chaos, and signal stationarity. Multidimensional internal features are input into a pre-deployed isolated forest anomaly detection model for anomaly scoring. When the anomaly score continues to be higher than a preset threshold, a hierarchical adaptive recalibration strategy is triggered. The hierarchical adaptive recalibration strategy includes: maintaining the current state for minor drifts; recalibrating and updating the parameters of the late reverberation model based on maximum likelihood estimation for significant steady-state changes; and re-triggering the full acquisition of room acoustic prior information and online adaptive fine-tuning of the probability estimation model after the idle condition is met for drastic non-steady-state changes.

[0021] This implementation method achieves self-supervised quality monitoring based on internal system state consistency checks, without relying on external labels or potentially invalid objective scores; the event-driven hierarchical response strategy adopts differentiated maintenance measures according to the degree of environmental change, avoiding frequent full recalibration from affecting normal use, and giving the system autonomous maintenance capabilities throughout its entire lifecycle.

[0022] In addition, this application also provides a speech intelligibility enhancement system based on sound field noise cognition, including a distributed speaker array, a microphone array, a processor, and a memory; The memory stores a computer program. When the processor executes the computer program, it coordinates the distributed speaker array and the microphone array to implement a speech intelligibility enhancement method based on sound field noise cognition as described above.

[0023] The system described above reuses existing distributed speaker and microphone array hardware in fixed installation scenarios, enabling automatic acquisition and closed-loop processing of room acoustic prior information without additional equipment. Combined with software algorithm packages executed by the processor, it achieves low-cost, privacy-friendly, and highly robust voice clarity enhancement deployment.

[0024] Beneficial effects: The technical solution provided in this application establishes a dual-channel decoupled cognitive model for reverberation and noise. It utilizes the sound pressure attenuation coefficient from prior room acoustic information to precisely drive recursive energy estimation. Furthermore, a pre-whitening mechanism is used to remove reverberation tails before background noise estimation, completely separating the two types of interference from a physical perspective. This avoids the speech dryness and fragmentation problems caused by obfuscation processing in traditional methods. Through on-site synthesized simulated samples and rapid fine-tuning via meta-learning, the probabilistic estimation model can adapt to specific sound fields without actual recordings. This achieves dual-dimensional probabilistic soft decision-making while protecting privacy, ensuring the selective preservation of beneficial early reflections. Combining a language-specific perception enhancement template with psychoacoustic masking threshold calculation, compensation is only applied to key phoneme frequency bands masked by noise. Dynamic constraints using an auditory loudness model elevate speech enhancement to the auditory perception level, ensuring a natural and comfortable listening experience. A self-supervised anomaly detection and hierarchical adaptive recalibration strategy based on multi-dimensional internal features enables the system to autonomously perceive and respond differently to environmental changes, ensuring the stability and engineering robustness of long-term delivery quality in fixed installation scenarios. Attached Figure Description

[0025] Figure 1 A flowchart illustrating a speech intelligibility enhancement method based on sound field noise cognition provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a method for adaptive fine-tuning a probability estimation model provided in an embodiment of this application. Figure 3 A schematic flowchart illustrating a method for generating a hybrid gain mask provided in an embodiment of this application; Figure 4 This is a system architecture diagram of a speech intelligibility enhancement system based on sound field noise cognition, provided for an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application.

[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0028] Example 1 like Figure 1 As shown, this embodiment provides a speech intelligibility enhancement method based on sound field noise cognition. This method achieves refined enhancement of speech signals in complex acoustic environments by establishing a dual-channel decoupled cognitive model of reverberation and noise. Specifically, the method includes the following steps S10 to S50.

[0029] Step S10: Obtain the room acoustic prior information and the current input signal for the target sound field. Specifically, the room acoustic prior information is a structured description of the acoustic characteristics of the target space. As a global control variable for the entire enhancement process, it determines how subsequent algorithms distinguish and process different types of acoustic components. In this embodiment, the room acoustic prior information includes an early reflection spectrum and a late reverberation model. The early reflection spectrum consists of a set of reflection peaks composed of multiple time delays, amplitudes, and directions, representing deterministic reflection components that positively contribute to speech intelligibility and spatial perception. The late reverberation model includes frequency-related sound pressure attenuation coefficients or reverberation times, representing the exponential decay rate of reverberation energy. The current input signal is typically acquired by a distributed microphone array fixedly installed on the ceiling or wall. Due to the long pickup distance, this signal often contains low-energy direct sound, high-energy reverberation components, and irrelevant background noise.

[0030] Step S20: Based on prior room acoustic information, the reverberation energy spectrum of the current input signal is estimated through the first cognitive channel, and the current input signal is pre-whitened based on the reverberation energy spectrum to obtain a pre-whitened signal. Specifically, the first cognitive channel is a physically driven cognitive pathway whose core function is to deterministically track and estimate reverberation energy using the attenuation characteristics in the known prior room acoustic information, rather than relying on blind statistical assumptions. The pre-whitening process is not intended to completely eliminate reverberation to improve listening experience, but rather as a signal preprocessing method to subtract the reverberation energy component that conforms to the exponential decay law from the spectrum of the current input signal. The physical significance of this step is to remove the deterministic interference of "reverberation" from the observed signal, making the remaining pre-whitened signal statistically approximate a combination of "clean speech plus background noise". In this way, this embodiment creates a clean observation domain uncontaminated by reverberation tails for subsequent background noise estimation, fundamentally avoiding the noise overestimation problem caused by reverberation residue in traditional methods.

[0031] Step S30 involves estimating the background noise of the pre-whitened signal using the second cognitive channel to obtain the background noise power spectrum. Specifically, the second cognitive channel is a statistically driven cognitive pathway specifically designed to estimate non-stationary background noise in the pre-whitened signal domain. Since step S20 has already removed the reverberation component with long tail characteristics, the minimum statistics in the pre-whitened signal are no longer affected by the rise in reverberation energy, and can accurately reflect the lower limit level of random background noise such as air conditioning, traffic, or crowd noise. In this embodiment, the second cognitive channel can employ the Minimum Controlled Recursive Average (IMCRA) algorithm or a variant thereof, which tracks the noise power spectrum by smoothly searching for local minima on the time-frequency plane and recursively updating them. This serial cascaded design—first removing reverberation through the first cognitive channel, and then estimating noise through the second cognitive channel—constitutes the core mechanism of the "reverberation-noise decoupling" in this application, ensuring that two distinctly different types of interference are quantified independently and accurately without interfering with each other.

[0032] Step S40 involves generating simulated acoustic samples based on prior room acoustic information and adaptively fine-tuning a pre-deployed probability estimation model based on these samples. The fine-tuned model outputs the probability of clean speech presence and the probability of late reverberation dominance. Specifically, to overcome the problem of insufficient generalization ability of general deep learning models in specific room acoustic environments, this embodiment utilizes the prior room acoustic information obtained in step S10 to synthesize simulated acoustic samples in real time on-site. The synthesis process involves convolving clean speech samples with early reflection spectra to simulate multipath propagation, then superimposing exponentially decaying tails conforming to the late reverberation model, and finally mixing in background noise samples. Based on these simulated samples with precise component labels, fast adaptation algorithms such as meta-learning are used to fine-tune the pre-deployed probability estimation model online, transforming it into a dedicated model adapted to the current sound field within seconds. The fine-tuned model outputs two parallel probability values: the probability of clean speech presence represents the likelihood of the direct sound and early reflection components being present in the current time-frequency unit, which are crucial for speech intelligibility; the probability of late reverberation dominance represents the likelihood that the late reverberation component is dominant. This two-dimensional soft-decision output provides a refined basis for the subsequent selective retention of beneficial acoustic components, and the entire process does not require the collection of real user voice, effectively ensuring privacy and security.

[0033] Step S50: A hybrid gain mask is generated based on the probability of clean speech presence and the probability of late reverberation dominance. The current input signal is then enhanced using the hybrid gain mask and the background noise power spectrum to obtain an enhanced signal. Specifically, the generation of the hybrid gain mask is no longer a simple binary switch, but a continuously weighted function based on the aforementioned dual-probability output. The system can dynamically calculate the gain value of each time-frequency unit according to the relative magnitudes of the probability of clean speech presence and the probability of late reverberation dominance, thereby achieving complete preservation of speech components, selective suppression of late reverberation, and effective suppression of background noise. Subsequently, this hybrid gain mask is applied to the original current input signal (rather than the pre-whitened signal) and combined with the background noise power spectrum obtained in step S30 for final enhancement. This processing method utilizes the accurate noise estimation obtained in the pre-whitening stage while preserving the natural acoustic details in the original signal, avoiding the artificial traces that may be introduced by directly reconstructing speech on the pre-whitened signal. The final enhanced signal output improves clarity while maintaining good naturalness and listening comfort.

[0034] Through the coordinated operation of steps S10 to S50 described above, this embodiment constructs a complete "perception-decoupling-enhancement" closed loop. Room acoustic prior information is used throughout, serving as both the physical constraint of the first cognitive channel and the source of simulated sample generation. Pre-whitening processing acts as a bridge connecting the two cognitive channels, achieving physical separation of reverberation and noise at the signal level. The dual-probability model based on simulation fine-tuning transforms this physical separation into intelligent decision-making at the perception level. This architectural design solves the problem of dry and fragmented sound caused by the confusion between reverberation and noise in existing technologies.

[0035] Example 2 Based on Example 1, this example further refines the specific implementation of reverberation energy spectrum estimation and pre-whitening processing in step S20, aiming to establish a physical parameter-driven causal chain for signal processing and ensure the accuracy and robustness of reverberation cognition.

[0036] As one implementation method, estimating the reverberation energy spectrum of the current input signal through a first cognitive channel includes: calculating the reverberation energy spectrum using a recursive formula based on prior information about the room's acoustics. : ; in, For the current frame, For frequency points, This is the reverberation energy spectrum of the previous frame. This represents the frequency-dependent sound pressure attenuation coefficient from the room's acoustic prior information. The number of frame shift sampling points. The system sampling rate, For the current input signal at frequency point The amplitude spectrum at that location. Specifically, this recursive formula is not a simple mathematical fitting model, but a precise discretized description of the physical process of sound energy propagation within a closed space. The first term in the formula... This represents the reverberation energy remaining from the previous frame to the current frame, which follows a strict exponential decay law; the coefficient "2" in the exponential term originates from the square relationship between sound pressure and energy, i.e., sound pressure decays according to... Attenuation, and since the energy is proportional to the square of the sound pressure, therefore according to... Attenuation. The second term in the formula. This characterizes the new injection contribution of the current frame's input signal to the reverberation energy field. The key lies in the smoothing factor. Instead of relying on fixed values ​​obtained through empirical parameter tuning, the sound pressure attenuation coefficient is directly measured from prior acoustic information of the room. The parameters are precisely calculated from the system sampling parameters. This direct mapping of physical parameters allows the algorithm to adaptively meet the boundary conditions of different acoustic environments: in a strongly sound-absorbing room, With a relatively large reverberation level, the smoothing factor approaches 0, and the reverberation estimation can quickly follow changes in the current signal; in a long reverberation room, When the reverberation energy is relatively small, the smoothing factor approaches 1, and the reverberation energy exhibits a slow decay and trailing characteristic; while in the silent segment, when At this point, the reverberation energy automatically degenerates into a purely exponential decay process. In this way, the first cognitive channel achieves deterministic tracking of the reverberation component, avoiding the estimation lag or divergence problems that traditional blind estimation algorithms are prone to in non-stationary speech segments.

[0037] It should be noted that the reverberation energy spectrum of the previous frame... The state is stored within the recursive algorithm itself. In the first frame after system startup, Initialized to zero or the current frame signal energy The result calculated for each subsequent frame All of these are stored in the status register and used as the "reverberation energy spectrum of the previous frame" for calculation in the next frame. Frequency-dependent sound pressure attenuation coefficient. Directly extracted from the prior acoustic information of the room obtained in Phase 1: By exponentially fitting the envelope of each frequency band of the room impulse response, the decay rate of the sound pressure amplitude at each frequency point over time is obtained, in neperts per second. Frame shift sampling points. and system sampling rate Determined by the processing parameters of the short-time Fourier transform. It is usually set to a standard audio sampling rate such as 16000Hz or 48000Hz. It is usually set to 25% to 50% of the frame length (e.g., a frame shift of 128 or 256 points when the frame length is 512 points), and both determine the inter-frame physical time step. The time step directly affects the accuracy of calculating the attenuation of reverberation energy between adjacent frames.

[0038] Further, as one implementation method, pre-whitening processing of the current input signal based on the reverberation energy spectrum includes: calculating the transfer function based on the amplitude spectrum and reverberation energy spectrum of the current input signal, and constructing a time-varying pre-whitening filter based on the transfer function; wherein, the transfer function is determined based on the difference between the amplitude spectrum and the reverberation energy spectrum, a preset safety factor, and a minimum gain lower limit; the current input signal is filtered by the time-varying pre-whitening filter to obtain a pre-whitened signal. Specifically, the core purpose of the pre-whitening processing is to extract the reverberation component that conforms to the above-mentioned exponential decay law from the observed signal, thereby providing a clean observation domain with statistical characteristics approximating "clean speech plus background noise" for subsequent background noise estimation. When constructing the transfer function, this embodiment does not adopt the ideal inverse filter form, but introduces a dual constraint mechanism to enhance engineering robustness. First, a preset safety factor (e.g., 0.95) is introduced to weight the estimated value of the reverberation energy spectrum, retaining a certain processing margin to prevent the direct sound or early reflection sound from being mistakenly deleted due to overestimation of reverberation energy, thereby avoiding distortion or a dry listening experience in the speech signal. Secondly, setting a minimum gain lower limit (e.g., -20dB) limits the maximum attenuation of the transfer function in a specific time-frequency unit, preventing excessively deep spectral dips in frequency bands with extremely low signal-to-mix ratios. This not only avoids music noise artifacts introduced by depth filtering, but also prevents subsequent noise estimation algorithms from becoming numerically unstable due to weak input signals.

[0039] It should be noted that the safety factor and minimum gain lower limit are determined based on the trade-off between tolerance for reverberation estimation errors and signal fidelity in engineering practice. The safety factor is a multiplicative factor slightly less than 1, which prevents the pre-whitening filter from excessively deducting energy from the signal amplitude spectrum due to small deviations or instantaneous fluctuations in the reverberation energy spectrum estimation, thereby damaging the subtle components of direct sound or early reflections. This factor is determined through offline optimization using a large number of samples collected in typical rooms (such as conference rooms and lecture halls) with the same type of microphone array before actual deployment. A performance curve is plotted with the purity of noise estimation on the vertical axis and the retention of speech components on the horizontal axis, and the coefficient value that achieves the best balance between the two is selected. The minimum gain lower limit is an absolute value constraint, which prevents the filter gain from being calculated to approach zero or negative infinity in extreme time-frequency units where the reverberation energy is much greater than the direct sound energy (such as the tail of speech attenuation or long-distance pickup scenarios), causing the signal of that unit to be completely eliminated, resulting in an unnatural "sound black hole" effect. The lower limit is determined based on the human ear's perception threshold for sound interruption and the actual dynamic range of speech: a lower limit of -20dB means that even at this time-frequency unit, 1% of the original signal energy still passes through, preserving the auditory sense of sound field continuity. At the same time, this energy is low enough not to substantially pollute the background noise estimation of the subsequent minimum control recursive averaging algorithm. In practical implementation, these two parameters can be fixed as system configuration items before leaving the factory, or they can be adaptively fine-tuned during the system initialization stage based on measured prior information about room acoustics (such as the magnitude of reverberation time): the longer the reverberation time, the safety factor can be slightly reduced (e.g., from 0.95 to 0.90) to obtain a stronger reverberation stripping effect; the minimum gain lower limit can be slightly increased (e.g., from -20dB to -18dB) to ensure auditory continuity under strong reverberation.

[0040] By combining the physical-driven reverberation estimation with the controlled pre-whitening mechanism, this embodiment effectively solves the coupling problem between reverberation and noise at the signal level. Without pre-whitening, directly applying noise estimation algorithms such as minimum-controlled recursive averaging to the original reverberant signal can easily lead to misjudging high-energy late-stage reverberation as a significant increase in the background noise threshold due to the long-term correlation and non-stationary decay characteristics of the reverberation tail. This results in a severe overestimation of the noise power spectrum, leading to excessive suppression in subsequent enhancement processing, causing speech breaks or underwater sound effects. In contrast, this embodiment precisely quantizes and removes reverberation through the first cognitive channel, enabling the second cognitive channel to accurately capture the true background noise level without reverberation interference, laying a solid foundation for the performance of the entire speech clarity enhancement system.

[0041] Example 3 Building upon Example 1, this example further refines the specific implementation of the on-site adaptive fine-tuning and dual-probability output of the probability estimation model in step S40. As one implementation method, the pre-deployed probability estimation model is adaptively fine-tuned based on simulated acoustic samples, such as... Figure 2 As shown, the implementation includes: Step S31, acquiring clean speech samples and background noise samples; Step S32, based on the early reflection spectrum and late reverberation model in the room acoustic prior information, convolving and superimposing the reverberation tails on the clean speech samples, and combining them with the background noise samples to generate simulated acoustic samples; Step S33, based on the simulated acoustic samples, performing multiple batches of small-sample task sampling and gradient updates on the weights of the probability estimation model through a meta-learning algorithm to obtain the adapted model weights; the probability estimation model is a deep neural network model; Step S34, fine-tuning the probability estimation model through the adapted model weights to obtain the fine-tuned probability estimation model. Specifically, the core of this implementation lies in using the room acoustic prior information acquired in step S10 to construct high-fidelity training data in real time on-site, thereby solving the problem of insufficient generalization ability of general deep learning models in specific acoustic environments, while completely avoiding the privacy risks brought about by collecting real user voices. In the process of generating simulated acoustic samples, the system first randomly selects samples from the built-in clean speech library and background noise library as base materials. Next, using the time delay, amplitude, and direction triplet parameters recorded in the early reflection spectrum, multipath convolution is performed on the clean speech samples to accurately simulate the early arrival structure formed by sound waves after reflection from interfaces such as walls and ceilings in the target sound field. Subsequently, based on the frequency-related sound pressure attenuation coefficient in the late reverberation model, an exponentially decaying reverberation tail tone that conforms to the physical characteristics of the room is generated and superimposed on the convolved signal to restore the realistic reverberation tail effect. Finally, the processed reverberant speech is mixed with background noise samples at a preset signal-to-noise ratio to generate simulated acoustic samples containing thousands of pairs of "noisy reverberant input-component labels". This synthesis strategy based on physical priors ensures that the generated samples are highly consistent with the real acoustic environment of the current sound field in terms of statistical characteristics, and the entire process does not require microphones to collect any on-site human voices, thus guaranteeing data privacy and security from the source.

[0042] Furthermore, after obtaining simulated acoustic samples, this embodiment employs a meta-learning algorithm to rapidly adapt the probabilistic estimation model. Unlike traditional de novo training or full fine-tuning, meta-learning algorithms (such as lightweight variants of Reptile or MAML) aim to teach the model "how to learn quickly." Specifically, the system divides the simulated acoustic samples into multiple small-sample task batches and performs gradient updates in only a few steps (e.g., 5 steps) based on a pre-deployed deep neural network model (such as a convolutional recurrent network, CRN). This mechanism enables the model to capture the unique acoustic fingerprint features of the current room within seconds, rapidly transforming from a general speech enhancement model into an "acoustic expert" adapted to a specific sound field. Compared to conventional model training that relies on large amounts of real data, this method significantly reduces the time cost and computing power requirements for on-site deployment, making it possible to complete model personalization within the short preparation window before the meeting begins.

[0043] As one implementation method, the method outputs the probability of clean speech presence and the probability of late reverberation dominance through a fine-tuned probability estimation model. This includes: extracting spectral and spatial features from the current input signal to obtain signal features; inputting the signal features into the fine-tuned probability estimation model; and outputting the probability of clean speech presence and the probability of late reverberation dominance within a preset numerical range through two parallel Sigmoid activation units set in the output layer of the fine-tuned probability estimation model. The probability of clean speech presence represents the probability of the presence of direct sound and early reflection components. The sum of the probability of clean speech presence, the probability of late reverberation dominance, and the probability of noise presence determined based on the background noise power spectrum satisfies a preset normalization constraint. Specifically, this implementation method provides a refined decision-making basis for subsequent gain mask generation through a two-dimensional soft-decision output mechanism. The fine-tuned probability estimation model receives the spectral features (such as logarithmic amplitude spectrum) and spatial features (such as multi-channel phase difference or beamforming response) of the current input signal. After network inference, the two parallel Sigmoid activation units in the output layer map out two independent probability values. The probability of clean speech presence represents the likelihood that the current time-frequency unit contains beneficial acoustic components such as direct sound and early reflections, which are crucial for maintaining speech intelligibility and spatial localization. The probability of late reverberation dominance represents the likelihood that the current time-frequency unit is dominated by late reverberation components, which are typically the main sources of interference leading to speech blurring and decreased clarity. It's important to note that these two probability values ​​are not mutually exclusive binary states, but rather continuously varying soft metrics. This allows the system to simultaneously perceive the presence of speech components and the degree of reverberation masking within the same time-frequency unit, thus avoiding the misjudgment or truncation issues that traditional hard-threshold speech activity detection (VAD) is prone to when processing reverberant speech.

[0044] Furthermore, to ensure the physical rationality and numerical stability of the probability output, this embodiment introduces a normalization constraint. Specifically, the sum of the probability of pure speech, the probability of late reverberation dominance, and the probability of noise presence indirectly derived from the background noise power spectrum obtained in step S30, should be approximately equal to 1 in any time-frequency unit. This constraint constructs a complete probability space, meaning that the currently observed signal energy is completely decomposed into a linear combination of speech, reverberation, and noise components. For example, in the initial consonant burst of a speaker's sentence, the probability of pure speech may be as high as 0.9 or higher, while the probability of late reverberation dominance and the probability of noise presence are both close to 0; in the long trailing vowel at the end of a sentence, as the speech energy decays, the probability of late reverberation dominance may gradually rise to above 0.7, and the probability of pure speech may decrease accordingly; in the pure silence segment, the probability of noise presence dominates. This dynamic probability distribution characteristic enables the subsequently generated hybrid gain mask to adaptively adjust according to the real-time proportion of each component: maintaining high gain in the speech-dominant region to preserve details, applying moderate suppression in the reverberation-dominant region to improve clarity, and performing deep attenuation in the noise-dominant region to suppress background noise, thereby achieving refined perception and selective enhancement of complex acoustic environments.

[0045] Example 4 Based on Example 1, this example further refines the specific implementation of hybrid gain mask generation and enhancement processing in step S50, aiming to elevate signal processing from simple physical noise reduction to a perceptual enhancement level that conforms to the characteristics of human hearing, and solve the problems of speech dryness and lack of naturalness that are easily caused by traditional enhancement algorithms.

[0046] As one implementation method, a hybrid gain mask is generated based on the probability of the presence of clean speech and the probability of late reverberation dominance, such as... Figure 3As shown, the implementation includes: Step S41, obtaining the signal-to-reverberation ratio (SRR) estimate of the current time-frequency unit; Step S42, performing exponential weighted calculations on the probability of clean speech presence and the probability of late reverberation dominance to obtain the initial gain; Step S43, adaptively adjusting the exponent of the late reverberation dominance probability based on the SRR estimate, recalculating the initial gain to obtain the adjusted gain; Step S44, performing asymmetric recursive smoothing on the adjusted gain based on the time axis to obtain the mixing gain mask; wherein, during the execution of the asymmetric recursive smoothing, the start-up time of gain increase is shorter than the release time of gain decrease. Specifically, this implementation introduces the signal-to-reverberation ratio (SRR) as a dynamic control variable to address the problem of poor adaptability of fixed parameters under different acoustic conditions. When the signal-to-mixing ratio (SMR) estimate of the current time-frequency unit is low, it indicates that the late reverberation energy dominates relative to the direct sound, significantly enhancing the masking effect on the human ear. In this case, the system automatically increases the exponential weight corresponding to the late reverberation dominance probability (e.g., from the default 1.5 to 2.5), thereby increasing the suppression of reverberation components and preventing a muddy sound. Conversely, when the SMR is high, the reverberation masking effect weakens, and the system reduces the exponential weight to retain more spatial information. This adaptive mechanism ensures that the algorithm maintains the optimal clarity-naturalness balance in both near-field and far-field sound pickup or in rooms with different reverberation intensities.

[0047] Furthermore, asymmetric recursive smoothing is a crucial step in ensuring the naturalness of the sound. In traditional symmetrical smoothing, rapid changes in gain often lead to speech envelope distortion, producing artifacts similar to "underwater sound effects" or mechanical sounds. This embodiment employs an asymmetric strategy of "fast start, slow release": the attack time for gain increase is set to a short value (e.g., 5 to 10 milliseconds) to ensure that when the speech begins (especially transient components such as voiceless consonants / s / and / t / ), the gain can quickly respond and rise to the target value, fully preserving the transient details that are crucial for intelligibility; while the release time for gain decrease is set to a longer value (e.g., 50 to 100 milliseconds), allowing the gain to slowly fall back during speech decay or pauses, allowing some beneficial early reflections and natural reverberation tails to be preserved, forming a smooth auditory transition. This processing method effectively avoids the truncated speech or swallowing of breath sounds caused by abrupt gain changes, making the enhanced speech both clear and full.

[0048] As one implementation method, the current input signal is enhanced using a hybrid gain mask and background noise power spectrum, including: obtaining a language acoustic feature enhancement template matching the current language; the language acoustic feature enhancement template includes preset weights for key phoneme frequency bands in the current language; enhancing the current input signal based on the hybrid gain mask to obtain an initial enhanced signal; calculating a global masking threshold for the initial enhanced signal based on the background noise power spectrum and a psychoacoustic model; calculating the amplitude difference between the global masking threshold and the initial enhanced signal, and calculating a sharpness compensation gain based on the language acoustic feature enhancement template; performing sharpness compensation on the initial enhanced signal based on the sharpness compensation gain, and applying amplitude limiting constraints through a preset listening comfort threshold to obtain the enhanced signal. Specifically, this implementation method elevates speech enhancement from general spectrum recovery to a refined compensation stage combining linguistics and psychoacoustics. The language acoustic feature enhancement template is prior knowledge stored in the form of frequency band weight vectors, which identifies the key frequency bands that contribute the most to sharpness in a specific language. This allows subsequent compensation gains to be targeted, prioritizing the enhancement of weak links that determine semantic understanding, rather than blindly increasing the energy across the entire frequency band.

[0049] It should be noted that the construction of the language acoustic feature enhancement template was completed using an offline analysis and online invocation approach. Specifically, firstly, for each target language (such as Chinese, English, Japanese, etc.), a large-scale, multi-speaker clean speech database was collected, covering different genders, ages, speech rates, and dialect variations. Then, speech analysis tools were used to perform phoneme alignment and spectral analysis on each speech item. The speech analysis tools consist of a forced alignment tool based on a Hidden Markov Model or Deep Neural Network and a spectral statistical analysis module based on Short-Time Fourier Transform. The forced alignment tool uses a pre-trained acoustic model and pronunciation dictionary to accurately align the speech signal with the text at the frame level, while the spectral statistical analysis module statistically analyzes the energy distribution characteristics of each phoneme category according to the critical frequency band. Special attention is paid to key phonemes that contribute most to clarity, such as the high-frequency bursts of voiceless consonants ( / s / , / sh / , / f / , / t / , etc.), the mid-frequency formant transition region of voiced consonants, and the fundamental frequency and harmonic bands of tone information in tonal languages ​​(such as Chinese). Based on this, the relative importance weight of each frequency band in different phoneme categories is calculated to form the frequency band weight vector of the language. This vector shows obvious peaks in the consonant core frequency band (such as 2-6kHz) and the low-frequency transition frequency band unique to tonal languages ​​(such as 200-500Hz). Finally, the weight vectors of each language are stored as template files and pre-loaded into the system firmware.

[0050] Building upon this, the calculation of the global masking threshold incorporates psychoacoustic models (such as the publicly published Johnston (1988) psychoacoustic model or the psychoacoustic model I described in Appendix D of ISO / IEC 11172-3). This threshold characterizes the minimum sound level that the human ear can just barely perceive in the presence of residual background noise. Only when the amplitude of the initial enhancement signal in a certain time-frequency unit is lower than this masking threshold is it considered that the component has been completely masked by noise and cannot be perceived by the human ear, at which point the compensation mechanism is triggered. The magnitude of the compensation gain is equal to the difference between the masking threshold and the signal amplitude multiplied by the weight of the corresponding frequency band of the speech template. This "on-demand compensation" strategy avoids superimposing excess energy in frequency bands where the signal-to-noise ratio is already sufficient, thereby minimizing the risk of artificial traces and background noise amplification. During model operation, the power spectrum of the speech signal and background noise after hybrid gain masking is first divided into sub-bands, typically using a critical frequency band (Bark scale) that matches the frequency resolution of the human ear, dividing 0Hz to the upper limit of the sampling frequency into 24 or 25 Bark bands. For each Bark band, the signal energy and noise energy are calculated separately. Then, based on the tone purity (distinguishing between tone-like and noise-like current frames by spectral flatness measurement), the corresponding masking spread function is selected. The masking spread function for tone signals has a steeper slope (approximately +25dB / Bark), while the masking spread function for noise signals has a gentler slope (approximately +10dB / Bark). The energy of each sub-band is then spread across adjacent critical frequency bands, and the resulting superpositions yield the masking spread spectrum. Next, a frequency band-related relative masking offset (e.g., 2-6dB) is subtracted from the masking spread spectrum, and this is compared frequency-by-frequency with the absolute silence threshold (i.e., the minimum audible sound pressure level at each frequency, for example, approximately 2-3dB SPL at 1kHz, -3dB SPL at 4kHz, and 12dB SPL at 8kHz). The maximum value of the two is taken to form the final global masking threshold curve.

[0051] It should be noted that the global masking threshold changes dynamically frame by frame, and its specific value depends on the current spectral distribution of the signal and noise. Based on typical conference scenarios (residual air conditioning noise approximately 35-45 dB SPL, normal conversation approximately 60-65 dB SPL), a reference range can be given: In the 1-4 kHz mid-frequency band where speech activity is concentrated, the masking threshold is typically around 25-40 dB SPL; in the low-frequency band below 200 Hz where noise energy is relatively high, the masking threshold may rise to 40-50 dB SPL; in the high-frequency band (above 6 kHz), due to the natural attenuation of speech energy and the increase in the absolute silence threshold, the masking threshold is usually dominated by the absolute silence threshold, around 10-20 dB SPL. During operation, the system calculates the difference between the current amplitude of the enhanced speech and the global masking threshold (i.e., the audibility difference) for each time-frequency unit. When the speech amplitude is lower than the masking threshold, this difference is the gain reference amount that needs to be compensated. Subsequently, the final clarity compensation gain is generated by combining the frequency band weights of the speech acoustic feature enhancement template, thereby ensuring that key phonemes masked by noise can be clearly perceived by the human ear.

[0052] Finally, to prevent the compensation process from introducing new auditory discomfort, this embodiment sets a listening comfort threshold as a hard limiting constraint. Especially in the consonant burst range or high-frequency sibilance region, excessive instantaneous compensation gain can easily produce harsh spikes or metallic sounds. The system monitors the short-time loudness or peak level of the compensated signal in real time. Once the predicted value exceeds the preset comfort threshold, the compensation gain is dynamically compressed or clipped for protection. This constraint ensures that even under extremely harsh acoustic environments with strong compensation, the output speech always remains pleasing and comfortable to the ear, achieving a unity between technical specifications and subjective experience.

[0053] It should be noted that the setting of the listening comfort threshold is determined based on a combination of human auditory perception characteristics and subjective listening evaluation experiments. Its core objective is to prevent the intelligibility compensation gain from overshooting during consonant bursts (such as the instantaneous energy release of voiceless consonants like / s / , / t / , and / k / ), which could lead to a harsh sound or even auditory discomfort. The specific setting method is as follows: First, a short analysis window (typically 2-5 milliseconds) is set within which the instantaneous loudness of the speech signal after clarity compensation is calculated. Loudness calculation uses the Zwicker loudness model based on the ISO 532 standard or a simplified version, converting the physical sound pressure level into perceived loudness levels (in sonos or squares) to more closely approximate human hearing. Then, through analysis of a large number of normal speech samples, the natural loudness fluctuation range of consonant bursts under uncompensated conditions is statistically analyzed. Generally, instantaneous loudness fluctuations in normal speech within 10-20 squares (or an equivalent loudness level of approximately 10-15 squares) are considered natural and comfortable. Based on this, the listening comfort threshold is typically set as an upper limit for permissible short-term loudness increments, typically 15-20 squares (or an equivalent loudness level increment of approximately 12-18 dB). When the system predicts that the compensation gain will cause the instantaneous loudness increment to exceed the threshold, it triggers a limiting constraint, smoothly reducing the gain to within the threshold. This ensures that the enhanced speech improves clarity while remaining within the range of human ear comfort, avoiding clipping or harshness commonly found in digital audio. This threshold can be fine-tuned according to the actual application scenario (such as hearing aids, conferencing systems, and in-vehicle communication): the upper limit of the threshold can be appropriately relaxed in scenarios requiring extremely high clarity (such as hearing aids); while the lower limit of the threshold can be appropriately tightened in scenarios prioritizing long-term listening comfort (such as remote conferencing).

[0054] Example 5 In this embodiment, the specific implementation of obtaining the room acoustic prior information of the target sound field in step S10 is further refined. As one implementation method, obtaining the room acoustic prior information of the target sound field includes: transmitting a nonlinear logarithmic sweep frequency signal containing a sinusoidal frequency modulation term to the target sound field, and performing multi-channel synchronous acquisition through a sub-band time-division strategy to obtain the sweep frequency response signal; constructing a reference signal containing an inverse harmonic distortion model, and deconvolving the sweep frequency response signal based on the reference signal to obtain the room impulse response; sparsely reconstructing the room impulse response in the continuous time delay domain based on the atomic norm minimization method, separating the early reflection spectrum and the late reverberation model, and labeling them as room acoustic prior information; wherein, the early reflection spectrum consists of multiple time-delay-amplitude-direction reflection peaks, and the late reverberation model includes frequency-related reverberation time or sound pressure attenuation coefficients. Specifically, this implementation method constitutes the physical foundation of the entire speech clarity enhancement system, and its core purpose is to robustly extract structured acoustic fingerprints from complex real acoustic environments, providing accurate prior constraints for the dual-channel cognition and probabilistic model fine-tuning in the aforementioned embodiment. Although this step is listed first in the method flow description, in actual system deployment, it is usually executed during system initialization, recalibration triggered by environmental changes, or periodic maintenance, and belongs to the offline or quasi-offline processing module.

[0055] First, regarding the transmission and acquisition of the measurement signal, this embodiment employs a specially designed nonlinear logarithmic sweep frequency signal. Unlike traditional linear or pure logarithmic sweep frequencies, the instantaneous frequency change rate of this signal is sinusoidally modulated, and its mathematical form can be expressed as follows: Instantaneous phase The derivative of the instantaneous frequency Includes a sinusoidal modulation term .here, For signal amplitude, The modulation index, This is the modulation frequency. The physical mechanism of introducing this nonlinear modulation term is to combat the harmonic distortion that is prevalent in loudspeakers in practical engineering. In conventional frequency sweep measurements, the second and third harmonic distortions of the loudspeaker will produce false reflection peaks in the deconvolutioned room impulse response. These false peaks are easily misjudged as real early reflections, leading to errors in subsequent acoustic modeling. However, through sinusoidal frequency modulation design, these harmonic distortion energies can be focused in the time domain to specific, predictable locations, or orthogonally separated from the main frequency sweep signal in the frequency domain, thereby eliminating the interference of false reflection peaks at the source and ensuring the purity of the early reflection spectrum extraction.

[0056] Furthermore, to overcome the impact of background noise on measurement accuracy in complex sound fields, this embodiment employs a sub-band time-division strategy and a repetitive averaging mechanism. Specifically, the system divides the full-band sweep signal into multiple sub-bands using an octave band filter bank and transmits them in a time-division manner according to the signal-to-noise ratio characteristics of each sub-band. For low-frequency sub-bands (e.g., 100Hz to 200Hz), which are easily masked by steady-state low-frequency noise from air conditioning units, ventilation ducts, etc., the system automatically controls the speaker to transmit repeatedly multiple times (e.g., 4 times). After the microphone array synchronously acquires multi-channel response signals, it performs time-domain alignment and coherent averaging on these repeatedly acquired signals. According to signal processing theory, the energy of coherent signals increases linearly with the number of superpositions, while the energy of uncorrelated noise only increases with the square root of the number of superpositions. Therefore, 4 repetitive averagings can theoretically improve the signal-to-noise ratio of the low-frequency band by about 6dB. This mechanism effectively ensures the smoothness and accuracy of the room impulse response tail decay curve, thereby ensuring the sound pressure attenuation coefficient in the late reverberation model. The estimation accuracy is improved, avoiding the problem of overestimation of reverberation time caused by low-frequency noise pollution.

[0057] After obtaining the high signal-to-noise ratio swept frequency response signal, this embodiment does not directly use the ideal swept frequency signal as a reference for deconvolution. Instead, it constructs a generalized reference signal that includes an inverse model of harmonic distortion. This reference signal is in the form of... ,in and These represent the typical second and third harmonic distortion coefficients of the same model of loudspeaker, pre-calibrated through offline testing. Using the adaptive least mean square (LMS) algorithm, the swept frequency response signal is deconvolved with this generalized reference signal, essentially incorporating the loudspeaker's nonlinear characteristics into the inverse model of system identification. Compared to simple linear deconvolution, this approach more thoroughly eliminates the distortion introduced by the loudspeaker itself, resulting in a room impulse response (RIR) that more accurately reflects the acoustic transmission characteristics of the space itself, rather than a hybrid characteristic of "loudspeaker + space".

[0058] Finally, for the structured decomposition of the room impulse response, this embodiment employs the Atomic Norm Minimization (ANM) method. Unlike traditional matched pursuit or cepstral analysis based on discrete grids, ANM optimizes directly in the continuous time-delay domain, fundamentally avoiding parameter estimation bias caused by grid mismatch. This method models the room impulse response as the sum of two parts: an early reflection component composed of a finite number of sparse atoms, and a late reverberation component that conforms to exponential decay statistics. By solving a convex optimization problem, ANM can accurately separate the early reflection spectrum and the late reverberation model. The early reflection spectrum is composed of... It consists of the strongest reflection peaks, each reflection peak with a time delay Amplitude and incident direction The triplet provides a precise description. These parameters directly correspond to the multipath convolution conditions required for generating simulated acoustic samples in Example 3, enabling the synthesized data to realistically reproduce the spatial auditory characteristics of the target sound field. The late reverberation model outputs frequency-dependent sound pressure attenuation coefficients. This coefficient characterizes the exponential decay rate of reverberant sound energy over time and is the core physical parameter for calculating the reverberant energy spectrum in the recursive formula of Example 2. Through the above series of noise reduction, distortion reduction, and high-precision reconstruction methods, this embodiment provides a solid and reliable physical prior for the entire speech enhancement system, making subsequent algorithm processing no longer unfounded, but rather a precise control based on a deep understanding of the target sound field.

[0059] Example 6 like Figure 4As shown, this embodiment provides a speech intelligibility enhancement system 60 based on sound field noise cognition. System 60 is a physical device for implementing any of the methods in embodiments 1 to 5. Specifically, system 60 includes a distributed speaker array 61, a microphone array 62, a processor 63, and a memory 64. The distributed speaker array 61 and microphone array 62 are typically existing fixed-installation audio equipment within the target sound field (such as a conference room, lecture hall, command center, etc.), rather than being independent measurement hardware specifically added for this application. This architecture design fully utilizes the existing sound reinforcement and pickup infrastructure deployed in intelligent engineering, significantly reducing the system's hardware cost and construction complexity.

[0060] The memory 64 stores a computer program. When the processor 63 executes the computer program, it coordinates with the distributed speaker array 61 and microphone array 62 to implement the speech intelligibility enhancement method based on sound field noise cognition as described in the foregoing embodiments. Specifically, the processor 63, as the core control and computing unit of the system, establishes a data interaction connection with the distributed speaker array 61 and microphone array 62 through an audio transmission link. This audio transmission link can be a wired digital audio bus (e.g., AVB, Dante, or AES67 protocol), an analog audio line, or a wireless audio transmission protocol. In a preferred embodiment, the processor 63 is integrated into a central control host or a dedicated digital signal processor (DSP) device, and performs low-latency, high-precision clock synchronization and data transmission with multiple speaker units and microphone units distributed on the ceiling or walls through an AVB network audio bus. This synchronization mechanism is crucial for ensuring the time alignment of the frequency sweep signal transmission and response acquisition during the room acoustic prior information acquisition stage, directly determining the physical accuracy of subsequent reverberation energy spectrum estimation and pre-whitening processing.

[0061] During system operation, the processor 63's coordinated scheduling of hardware resources is mainly reflected in two stages. In the initialization or recalibration stage, the processor 63 reads preset nonlinear logarithmic sweep frequency signal parameters from the memory 64 and drives each unit in the distributed speaker array 61 to transmit measurement signals sequentially or in a time-division manner via the audio transmission link. Simultaneously, the processor 63 precisely controls the microphone array 62 to initiate multi-channel synchronous acquisition at the same time reference, transmitting the received sweep frequency response signals back to the processor memory. Subsequently, the processor 63 calls the algorithm module in the memory 64 to perform operations such as deconvolution, atomic norm minimization, and sparse reconstruction, generating and updating the room acoustic prior information stored in a specific area of ​​the memory. In the real-time speech enhancement stage, the processor 63 continuously receives the current input signal stream acquired by the microphone array 62, utilizes the stored room acoustic prior information and probabilistic estimation model weights, and performs dual-channel cognition, probabilistic model inference, and hybrid gain mask generation steps in parallel on on-chip or external DSP computing resources, ultimately outputting the enhanced speech signal.

[0062] Through the aforementioned system architecture, this embodiment translates abstract signal processing methods into concrete physical products. Since the system directly reuses the existing distributed loudspeakers and microphone array 62 within the target sound field as sensing and execution terminals, there is no need to deploy expensive professional acoustic measurement instruments or independent sensor networks. This not only significantly reduces hardware procurement and installation costs but also avoids the problems of decoration damage or aesthetic degradation caused by adding equipment. Furthermore, because all data processing is completed on the local processor, and the fine-tuning of the probability model relies entirely on synthesized simulated samples rather than real recordings, the system inherently possesses privacy protection characteristics at the physical architecture level, making it particularly suitable for deployment in highly secure, confidential locations or high-end business environments.

[0063] Example 7 Building upon Examples 1 to 6, this example further provides a self-healing maintenance mechanism for long-term system operation. As one implementation method, the method further includes: continuously calculating multidimensional internal features in the background, including reverberation prediction bias, noise estimation mutation rate, probabilistic chaos, and signal stationarity; inputting the multidimensional internal features into a pre-deployed isolated forest anomaly detection model for anomaly scoring; and triggering a hierarchical adaptive recalibration strategy when the anomaly score consistently exceeds a preset threshold. The hierarchical adaptive recalibration strategy includes: maintaining the current state for minor drifts; recalibrating and updating the parameters of the late-stage reverberation model based on maximum likelihood estimation for significant steady-state changes; and re-triggering the full acquisition of room acoustic prior information and online adaptive fine-tuning of the probabilistic estimation model after meeting idle conditions for non-steady-state drastic changes.

[0064] Specifically, the aforementioned multidimensional internal features constitute the "internal receptors" that enable the system to perceive its own health status and environmental compatibility. The original design purpose is to solve the problems of lag and unreliability in quality monitoring of traditional speech enhancement systems that rely on external labels or subjective scores. Among them, reverberation prediction bias refers to the residual statistics between the reverberation energy attenuation curve predicted by the first cognitive channel based on the current room acoustic prior information and the actual observed signal envelope attenuation. This feature directly reflects whether the sound pressure attenuation coefficient in the late reverberation model is still accurate. Noise estimation mutation rate refers to the rate of change of the background noise power spectrum output by the second cognitive channel between consecutive frames. Under normal circumstances, the background noise should be relatively stable. If the mutation rate increases abnormally, it often indicates that the nature of the noise source has changed or the noise estimation algorithm has fallen into a local minimum trap. Probabilistic chaos refers to the proportion of time-frequency units in which the probability of the existence of clean speech and the probability of late reverberation dominance output by the probabilistic estimation model are in the uncertain range of 0.3 to 0.7 for a long time. This feature characterizes the decrease in the model's confidence in the current sound field components. It usually occurs when the acoustic environment changes gradually, causing the model's prior knowledge to fail. Signal stationarity is measured by calculating the spectral flux of the input signal and is used to distinguish between structural changes in the environment and transient speech interference. These four dimensions, from the physical model matching degree, statistical estimation stability, neural network confidence, and signal characteristics, construct a complete system state observation space in four orthogonal directions.

[0065] Furthermore, this embodiment employs the Isolation Forest algorithm as the core engine for anomaly detection, rather than traditional fixed threshold or multivariate Gaussian models. This is because the normal state distribution of an acoustic environment is often non-Gaussian and multimodal, and in practical deployments, it is almost impossible to collect negative samples covering all failure modes. Isolation Forest isolates data points by randomly cutting them in the feature space. Normal data points are difficult to isolate due to their dense clustering, while anomalous data points are easily isolated due to their sparse and dispersed nature, making it naturally suitable for unsupervised anomaly detection tasks. In specific implementation, during the "healthy operation period" after initialization, the system collects the aforementioned four-dimensional feature data to train the Isolation Forest model, enabling it to learn the normal operating state boundaries of the system under the current specific sound field. When the real-time calculated feature vector falls outside this boundary, i.e., the anomaly score exceeds a preset threshold (e.g., 0.6) and persists for a certain duration (e.g., 3 seconds) to filter out transient disturbances, the system determines that a substantial change has occurred in the environment.

[0066] To address detected anomalies, this embodiment employs a refined, tiered adaptive recalibration strategy to balance maintenance accuracy and user experience. For minor drifts, such as parameter fluctuations caused by small numbers of people entering or leaving the system or slight changes in temperature and humidity, the system maintains the current state without intervention, relying on the algorithm's robustness to absorb errors and avoid auditory flicker caused by frequent adjustments. For significant steady-state changes, such as the opening or closing of curtains or the laying of carpets leading to uniform changes in sound absorption characteristics, the system triggers a lightweight parameter update: using the currently occurring speech signal as the excitation source, it rapidly corrects the sound pressure attenuation coefficient σ(f) in the late reverberation model online using the maximum likelihood estimation (MLE) method. This process involves only iterative optimization of a single parameter, typically taking hundreds of milliseconds to less than one second, completely imperceptible to call services, yet promptly correcting systematic biases in reverberation energy estimation. For drastic, non-steady-state changes, such as the opening of a movable partition wall causing a doubling of space volume, or the moving in of large furniture causing reconstruction of the reflection structure, the original early reflection spectrum and late reverberation model are completely invalidated, and local parameter repair cannot restore performance. The system will mark the event and enter a waiting state until it detects that the idle conditions are met (e.g., no voice activity detection for 5 consecutive minutes and stable background noise). Only then will it automatically re-trigger the process of acquiring all room acoustic prior information as described in Example 5 and the online adaptive fine-tuning of the probability estimation model as described in Example 3. This hierarchical mechanism ensures that the system can cope with daily environmental disturbances and completely rebuild acoustic cognition after major changes, achieving autonomous maintenance and performance preservation throughout its entire lifecycle.

[0067] This embodiment constructs a self-supervised closed loop that does not rely on external feedback, giving the voice enhancement system the ability to operate stably in complex and ever-changing engineering environments for a long time, fundamentally solving the problem of performance degradation caused by environmental drift of fixed-installation equipment.

[0068] Example 8 This embodiment provides a full lifecycle application scenario verification for a smart conference room, aiming to connect the scattered technical features in the aforementioned embodiments into a complete solution, and verify the collaborative effect and technical advantages of each module in a real engineering environment through specific scenario parameters.

[0069] During the pre-meeting initialization phase, the system automatically acquires prior acoustic information about the room and adapts the probability estimation model to the actual situation. Specifically, 15 minutes before the scheduled start time of the meeting, the central control unit triggers the initialization process according to preset instructions, sending a synchronized playback command to the distributed ceiling speaker array via the AVB network audio bus. The speakers sequentially emit a nonlinear logarithmic sweep signal with a sound pressure level of approximately 40 dB(A). This sound pressure level setting ensures that the measured signal has a sufficient signal-to-noise ratio to overcome ambient noise while avoiding auditory interference to those who may enter the venue early due to excessive volume. The sweep signal includes a sinusoidal frequency modulation term, which can effectively tolerate the harmonic distortion of the speakers themselves and prevent false reflection peaks after deconvolution. For the low-frequency sub-band of 100Hz to 200Hz, which is susceptible to masking by steady-state noise from air conditioning units, the system automatically controls the signal of this sub-band to be repeatedly emitted 4 times. After synchronous acquisition by the microphone array, coherent averaging is performed, which theoretically improves the signal-to-noise ratio of this frequency band by about 6 dB, ensuring the estimation accuracy of the low-frequency attenuation coefficient in the late reverberation model. Subsequently, the processor deconvolves the swept frequency response signal using a generalized reference signal incorporating a harmonic distortion inverse model, and sparsely reconstructs the early reflection spectrum and late reverberation model in the continuous time-delay domain based on the atomic norm minimization method. The entire process takes approximately 5 seconds. Next, the system synthesizes simulated acoustic samples on-site using acquired prior room acoustic information. The Reptile meta-learning algorithm adaptively fine-tunes the pre-deployed probability estimation model, completing model adaptation within seconds with only 5 gradient updates. This transforms the general model into a specialized model adapted to the acoustic characteristics of the current conference room, all without requiring the collection of any real human voices, thus completely avoiding the risk of privacy breaches.

[0070] During the real-time enhancement phase of the meeting, the system continuously runs a dual-channel cognitive and perceptual enhancement process. When participants speak, the first cognitive channel calculates the reverberation energy spectrum frame by frame using a recursive formula based on the sound pressure attenuation coefficient in the late reverberation model, and uses a time-varying pre-whitening filter with a safety factor of 0.95 and a minimum gain limit of -20dB to remove the reverberation tail. The second cognitive channel runs the IMCRA algorithm on the pre-whitened signal to accurately estimate the background noise power spectrum. The fine-tuned probability estimation model outputs the probability of clean speech presence and the probability of late reverberation dominance in real time, both of which satisfy the normalization constraint with the probability of noise presence. Based on the dual probability output and the signal-to-mixing ratio estimate, the system generates a mixing gain mask, in which an exponent of 1.0 is applied to the probability of clean speech presence and an exponent of 1.5 is applied to the probability of late reverberation dominance, and the reverberation exponent is automatically increased to 2.5 in the low signal-to-mixing ratio section to accelerate the decay. At the same time, an asymmetric recursive smoothing strategy with an onset time of 5 milliseconds and a release time of 50 milliseconds is adopted to fully preserve the transient details of consonants while maintaining a natural reverberation transition. For Chinese speech scenarios, the system loads a language acoustic feature enhancement template, assigning high weights to the voiceless consonant fricative region from 2kHz to 4kHz and the tone transition point around 250Hz. Combined with a global masking threshold calculated by a psychoacoustic model, compensation is applied only to key frequency bands masked by noise. Furthermore, a listening comfort threshold limiting constraint prevents harsh overshoot, ultimately outputting clear and natural enhanced speech. Simultaneously, a background self-supervised monitoring module continuously calculates four-dimensional internal features: reverberation prediction bias, noise estimation mutation rate, probabilistic chaos, and signal stationarity, and inputs these into an isolated forest model for anomaly scoring. During normal meetings, if the anomaly score remains below a preset threshold of 0.6, the system maintains stable operation without triggering any recalibration.

[0071] During the post-meeting environmental change recalibration phase, the system demonstrated its autonomous maintenance capabilities in handling unsteady and drastic changes. For example, when staff removed the movable soundproof wall connecting the adjacent room after the meeting, causing the space volume to instantly double, the backend monitoring module almost simultaneously detected a significant increase in reverberation prediction bias, drift in noise estimation statistics, and abnormal spectral flux. The abnormal score of the isolated forest model exceeded the threshold of 0.6 for more than 3 seconds, which the system identified as an "unsteady and drastic change." Since the meeting had ended and the room had returned to silence, the system automatically triggered a full recalibration process after detecting no voice activity for 5 consecutive minutes: retransmitting the nonlinear logarithmic sweep signal, recalculating the room impulse response and sound pressure attenuation coefficient (e.g., RT60 changed from 0.6 seconds to 1.2 seconds), and regenerating simulated samples based on the new acoustic prior information to fine-tune the probability estimation model online. After completion, the system sent a "Full acoustic environment optimization completed" log to the central control backend, making optimal preparations for the next meeting. This event-driven hierarchical response mechanism enables the system to autonomously adapt to structural changes in the physical space without human intervention, effectively solving the problem of performance degradation caused by environmental drift of fixed-installation equipment, and verifying the engineering robustness and intelligence level of the technical solution of this application throughout its entire life cycle.

[0072] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions within the technical scope disclosed in this application. For example, the above method can be applied to other enclosed acoustic spaces besides conference rooms, other equivalent sparse reconstruction algorithms can be used instead of atomic norm minimization, or other unsupervised anomaly detection models can be used instead of isolated forests. All such variations or substitutions should be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A speech intelligibility enhancement method based on sound field noise cognition, characterized in that, include: Acquire the room acoustic prior information and current input signal of the target sound field; Based on prior acoustic information of the room, the reverberation energy spectrum of the current input signal is estimated through the first cognitive channel, and the current input signal is pre-whitened based on the reverberation energy spectrum to obtain the pre-whitened signal; Background noise power spectrum is obtained by estimating the background noise of the pre-whitened signal through the second cognitive channel; Simulated acoustic samples are generated based on prior room acoustic information, and a pre-deployed probability estimation model is adaptively fine-tuned based on the simulated acoustic samples. The fine-tuned probability estimation model outputs the probability of clean speech and the probability of late reverberation dominance. A hybrid gain mask is generated based on the probability of the presence of clean speech and the probability of late reverberation dominance. The current input signal is then enhanced using the hybrid gain mask and the background noise power spectrum to obtain the enhanced signal. The adaptive fine-tuning of the pre-deployed probability estimation model based on simulated acoustic samples includes: Obtain clean speech samples and background noise samples; Based on the early reflection spectrum and late reverberation model in the prior information of room acoustics, the clean speech samples are convolved and superimposed with the reverberation tail, and combined with the background noise samples to generate simulated acoustic samples. Based on simulated acoustic samples, the weights of the probability estimation model are sampled and updated in multiple batches using a meta-learning algorithm to obtain the adapted model weights; the probability estimation model is a deep neural network model. The probability estimation model is fine-tuned by adapting the model weights to obtain the fine-tuned probability estimation model. The probability of clean speech presence and the probability of late reverberation dominance are output by the fine-tuned probability estimation model, including: The current input signal is subjected to spectral and spatial feature extraction to obtain signal features; The signal features are input into the fine-tuned probability estimation model. The output layer of the fine-tuned probability estimation model has two parallel Sigmoid activation units, which output the probability of pure speech presence and the probability of late reverberation dominance within a preset numerical range, respectively. The probability of pure speech presence represents the probability of the presence of direct sound and early reflection sound components. The sum of the probability of pure speech presence, the probability of late reverberation dominance, and the probability of noise presence determined based on the background noise power spectrum satisfies the preset normalization constraint. The generation of the hybrid gain mask based on the probability of pure speech presence and the probability of late reverberation dominance includes: Obtain the signal-to-mixing ratio estimate for the current time-frequency cell; The initial gain is obtained by exponentially weighting the probability of pure speech presence and the probability of late reverberation dominance. The exponent of the late reverberation dominance probability is adaptively adjusted based on the signal-to-mixing ratio estimate, and the initial gain is recalculated to obtain the adjusted gain. A hybrid gain mask is obtained by performing asymmetric recursive smoothing on the adjusted gain based on the time axis; wherein, during the execution of the asymmetric recursive smoothing, the start time of gain increase is shorter than the release time of gain decrease.

2. The speech intelligibility enhancement method based on sound field noise cognition according to claim 1, characterized in that, The pre-whitening process of the current input signal based on the reverberation energy spectrum includes: The transfer function is calculated based on the amplitude spectrum and reverberation energy spectrum of the current input signal, and a time-varying pre-whitening filter is constructed based on the transfer function; wherein, the transfer function is determined based on the difference between the amplitude spectrum and the reverberation energy spectrum, a preset safety factor, and a minimum gain lower limit; The current input signal is filtered by a time-varying pre-whitening filter to obtain a pre-whitening signal.

3. The speech intelligibility enhancement method based on sound field noise cognition according to claim 2, characterized in that, The step of estimating the reverberation energy spectrum of the current input signal through the first cognitive channel includes: Based on prior acoustic information of the room, the reverberation energy spectrum is calculated using a recursive formula. : ; in, For the current frame, For frequency points, This is the reverberation energy spectrum of the previous frame. This represents the frequency-dependent sound pressure attenuation coefficient from the room's acoustic prior information. The number of frame shift sampling points. The system sampling rate, This represents the amplitude spectrum of the current input signal at frequency point f. The background noise estimation of the pre-whitened signal through the second cognitive channel includes: By using a minimum-controlled recursive averaging algorithm, the minimum value of each frequency point is smoothly searched on the spectrum of the pre-whitened signal and recursively smoothed to obtain the background noise power spectrum.

4. The speech intelligibility enhancement method based on sound field noise cognition according to claim 1, characterized in that, The enhancement process of the current input signal by using a hybrid gain mask and background noise power spectrum includes: Obtain a language acoustic feature enhancement template that matches the current language; the language acoustic feature enhancement template includes preset weights for the frequency bands of key phonemes in the current language; The current input signal is enhanced using a hybrid gain mask to obtain the initial enhanced signal. The global masking threshold of the initial enhancement signal is calculated based on the background noise power spectrum and a psychoacoustic model. Calculate the amplitude difference between the global masking threshold and the initial enhanced signal, and calculate the sharpness compensation gain based on the speech acoustic feature enhancement template; The initial enhanced signal is compensated for intelligibility based on the intelligibility compensation gain, and then the amplitude is limited by a preset listening comfort threshold to obtain the enhanced signal.

5. The speech intelligibility enhancement method based on sound field noise cognition according to claim 1, characterized in that, The acquisition of room acoustic prior information for the target sound field includes: A nonlinear logarithmic sweep frequency signal containing a sinusoidal frequency modulation term is emitted into the target sound field, and multi-channel synchronous acquisition is performed through a sub-band time-division strategy to obtain the sweep frequency response signal; A reference signal containing an inverse model of harmonic distortion is constructed, and the swept frequency response signal is deconvolved based on the reference signal to obtain the room impulse response; Based on the atomic norm minimization method, the room impulse response is sparsely reconstructed in the continuous time delay domain, separating the early reflection spectrum and the late reverberation model, and labeled as the room acoustic prior information; the early reflection spectrum consists of multiple time delay-amplitude-direction reflection peaks, and the late reverberation model includes frequency-related reverberation time or sound pressure attenuation coefficient.

6. The speech intelligibility enhancement method based on sound field noise cognition according to claim 5, characterized in that, Also includes: The system continuously calculates multidimensional internal features in the background, including reverberation prediction bias, noise estimation abruptness rate, probabilistic chaos, and signal stationarity. Multidimensional internal features are input into a pre-deployed isolated forest anomaly detection model for anomaly scoring. When the anomaly score continues to be higher than a preset threshold, a hierarchical adaptive recalibration strategy is triggered. The hierarchical adaptive recalibration strategy includes: maintaining the current state for small drifts; and recalibrating and updating the parameters of the late reverberation model based on maximum likelihood estimation for significant steady-state changes. In response to unsteady and drastic changes, after the idle condition is met, the full acquisition of the room's acoustic prior information and the online adaptive fine-tuning of the probability estimation model are retried.

7. A speech intelligibility enhancement system based on sound field noise cognition, characterized in that, Includes a distributed speaker array, microphone array, processor, and memory; The memory stores a computer program. When the processor executes the computer program, it coordinates the distributed speaker array and the microphone array to implement a speech intelligibility enhancement method based on sound field noise cognition as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Single-channel speech enhancement method

    CN115132215A

  • Voice noise reduction method based on reasoning optimization and Bluetooth earphone

    CN121506162A