A method and system for speech enhancement of a loudspeaker
By acquiring multi-dimensional speech features and visual information from the megaphone microphone signal, and combining this with multi-sensor data, the system dynamically adjusts adaptive update strategies and noise suppression processing to solve the problems of speech clarity and echo cancellation in megaphone systems under complex acoustic environments, thereby improving the clarity and listening comfort of the amplified signal.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN JUNSHIAN TECH CO LTD
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-05
AI Technical Summary
Existing loudspeaker systems struggle to adapt effectively to background noise and echoes in complex acoustic environments, leading to decreased speech clarity, misinterpretation of human voices as echoes by adaptive filters, and failure of echo cancellation modules, all of which negatively impact communication efficiency and user experience.
By acquiring multi-dimensional speech features from the loudspeaker microphone signal, human voice confidence is generated, and an adaptive update strategy for acoustic echo cancellation processing is adjusted. Fluctuating noise is identified and suppressed. Combined with visual information and multi-sensor data, the adaptive update strategy and noise suppression processing are dynamically adjusted.
It enables intelligent adjustment of echo cancellation and noise suppression in complex acoustic environments, improving the clarity and listening comfort of the amplified signal, and enhancing communication efficiency and user experience.
Smart Images

Figure CN122157680A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and more specifically, to a speech enhancement method and system for a loudspeaker. Background Technology
[0002] In modern meetings, classrooms, and outdoor events, loudspeaker systems play a crucial role in ensuring that speakers' voices are clearly and loudly conveyed to the audience. However, in practical applications, loudspeakers often face challenges from complex environments, such as ambient noise and acoustic echoes. These factors can significantly reduce speech intelligibility, making it difficult for listeners to understand. Traditional loudspeaker equipment often struggles to cope with these varied acoustic environments, failing to effectively adapt to background noise and echoes, thus impacting communication efficiency and the overall experience.
[0003] In a public address system, when the acoustic reflection path changes significantly due to adjustments in the conference room layout, particularly the introduction of short-duration, high-energy echoes from highly reflective surfaces (such as glass walls), and if distant, low-energy, intermittent human speech is also present, and the system's energy threshold for judging speech activity cannot accurately distinguish between these valid human voices and echoes, the adaptive filter coefficient set will be incorrectly updated, misrepresenting valid human voices as echoes. Furthermore, if environmental noise (such as fluctuating air conditioning noise) is further superimposed on the microphone pickup signal, and the system's internal signal processing chain uses a fixed sequence of echo cancellation followed by noise suppression, the acoustic echo cancellation module will completely fail. It will not only fail to converge to the true echo path but will also incorrectly represent human voices and environmental noise into the filter, ultimately outputting severely distorted residual echoes and annoying low-frequency noise instead of clear amplified speech, severely impacting listening experience and communication efficiency.
[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0005] This application discloses a speech enhancement method and system for loudspeakers, aiming to solve the problem that existing loudspeaker systems are difficult to effectively adapt to background noise and echo effects in complex acoustic environments, resulting in a decrease in speech clarity.
[0006] The technical solution of this application is as follows: In a first aspect, this application discloses a method for enhancing the speech of a loudspeaker, including: The system acquires the sound signal picked up by the microphone of the loudspeaker, performs frame segmentation on the sound signal, extracts multi-dimensional speech features from each frame of the sound signal, outputs human voice confidence based on the multi-dimensional speech features and a preset classification rule, and generates human voice presence information based on the human voice confidence. The multi-dimensional speech features include at least two of the following: time domain features, frequency domain features and / or cepstral features. Acquire a loudspeaker reference signal, perform acoustic echo cancellation processing on the sound signal based on the loudspeaker reference signal to obtain an echo cancellation output signal, and when the presence of human voice information indicates the presence of human voice, adjust the adaptive update strategy of acoustic echo cancellation processing according to the confidence level of human voice. The adaptive update strategy includes freezing the update of adaptive filter coefficients and / or reducing the update intensity of the adaptive filter. Identify fluctuating noise in the echo-cancelled output signal that meets preset judgment conditions, preset intensity and / or dominant frequency changes over time, and determine the main noise frequency band of the echo-cancelled output signal based on the identification results of the fluctuating noise; Noise suppression processing is performed on the main noise frequency band, including at least narrowband suppression and / or adaptive gain suppression, to obtain a noise-suppressed output signal; Acquire the state information of the acoustic echo cancellation process, which includes at least the residual echo representation quantity; based on the human voice presence information, state information, and the identification results of fluctuating noise, adjust the update parameters of the adaptive update strategy used for subsequent processing and / or adjust the suppression intensity of the noise suppression process, and output the amplified speech signal after speech enhancement based on the noise suppression output signal.
[0007] Secondly, this application also discloses a voice enhancement system for a loudspeaker, comprising: The human voice presence information generation module is used to acquire the sound signal picked up by the microphone of the loudspeaker, perform frame-by-frame processing on the sound signal, extract multi-dimensional speech features for each frame of the sound signal, output human voice confidence based on the multi-dimensional speech features and according to the preset classification rules, and generate human voice presence information based on the human voice confidence. The multi-dimensional speech features include at least two of the following: time domain features, frequency domain features and / or cepstral features. The echo cancellation strategy adjustment module is used to acquire the loudspeaker reference signal, perform acoustic echo cancellation processing on the sound signal based on the loudspeaker reference signal to obtain the echo cancellation output signal, and adjust the adaptive update strategy of the acoustic echo cancellation processing according to the human voice confidence when the human voice presence information indicates the presence of human voice. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter. The noise identification module is used to identify fluctuating noise in the echo cancellation output signal that meets preset judgment conditions, preset intensity and / or dominant frequency changes over time, and to determine the main noise frequency band of the echo cancellation output signal based on the identification results of the fluctuating noise. The noise suppression processing module is used to perform noise suppression processing on the main noise frequency band. The noise suppression processing includes at least narrowband suppression and / or adaptive gain suppression to obtain a noise-suppressed output signal. The status information acquisition module is used to acquire the status information of acoustic echo cancellation processing. The status information includes at least the residual echo characterization quantity. The processing flow coordination module is used to adjust the update parameters of the adaptive update strategy for subsequent processing and / or adjust the suppression intensity of noise suppression processing based on the recognition results of human voice presence information, state information and fluctuating noise, and output the amplified speech signal after speech enhancement based on the noise suppression output signal.
[0008] Beneficial Effects: This application, by comprehensively considering the presence of human voices, the state of acoustic echo cancellation, and fluctuating noise, achieves intelligent and dynamic adjustment of echo cancellation and noise suppression strategies. It effectively solves problems in existing loudspeaker systems such as decreased speech intelligibility in complex acoustic environments, echo cancellation module failure, and misinterpretation of low-energy human voices. This method can significantly improve the clarity, naturalness, and listening comfort of amplified signals, overcoming the shortcomings of traditional loudspeaker equipment in dealing with ever-changing acoustic environments, thereby improving communication efficiency and user experience. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating a method for enhancing the voice of a loudspeaker provided in this application.
[0010] Figure 2 A flowchart of a speech enhancement system for a loudspeaker provided in this application.
[0011] In the diagram: 1. Human voice presence information generation module; 2. Echo cancellation strategy adjustment module; 3. Noise recognition module; 4. Noise suppression processing module; 5. Status information acquisition module; 6. Processing flow coordination module. Detailed Implementation
[0012] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0013] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0014] Traditional loudspeaker systems often struggle to guarantee speech clarity and intelligibility in complex acoustic environments, such as those with ambient noise and acoustic echoes. This is especially true when conference room layouts change, low-energy voices coexist with high-energy echoes, or there is fluctuating noise interference. Traditional speech enhancement methods may suffer from problems such as incorrect adaptive filter updates, misinterpreting voices as echoes, and poor noise suppression, leading to a significant deterioration in the quality of the amplified signal and impacting communication efficiency and user experience.
[0015] Reference Figure 1 In response, this application proposes a speech enhancement method for a loudspeaker, comprising: S1000: Acquires the sound signal picked up by the microphone of the loudspeaker, performs frame-by-frame processing on the sound signal, extracts multi-dimensional speech features for each frame of the sound signal, outputs human voice confidence based on the multi-dimensional speech features and according to the preset classification rules, and generates human voice presence information based on the human voice confidence. The multi-dimensional speech features include at least two of the following: time domain features, frequency domain features and / or cepstral features. S2000: Acquire loudspeaker reference signal, perform acoustic echo cancellation processing on sound signal based on loudspeaker reference signal to obtain echo cancellation output signal, and adjust the adaptive update strategy of acoustic echo cancellation processing according to human voice confidence when human voice presence information indicates the presence of human voice. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter. S3000: Identifies fluctuating noise in the echo cancellation output signal that meets preset judgment conditions, preset intensity and / or dominant frequency changes with time, and determines the main noise frequency band of the echo cancellation output signal based on the identification results of the fluctuating noise; S4000: Performs noise suppression processing on the main noise frequency band, including at least narrowband suppression and / or adaptive gain suppression, to obtain a noise-suppressed output signal; S5000: Acquires the state information of acoustic echo cancellation processing, which includes at least the residual echo representation quantity; adjusts the update parameters of the adaptive update strategy for subsequent processing and / or adjusts the suppression intensity of noise suppression processing based on the human voice presence information, state information and the identification results of fluctuating noise, and outputs the amplified speech signal after speech enhancement based on the noise suppression output signal.
[0016] Specifically, a sound signal refers to the raw audio signal picked up from the environment by the microphone of a loudspeaker, which contains various components such as human voice, echo, and environmental noise; framing the sound signal involves dividing the continuous sound signal into several short frames to facilitate subsequent feature extraction and analysis.
[0017] Multidimensional speech features refer to a variety of parameters extracted from each frame of sound signal that can characterize speech characteristics. For example, time-domain features may include short-time energy, zero-crossing rate, etc.; frequency-domain features may include spectral amplitude, spectral centroid, spectral bandwidth, etc.; cepstral features may include Mel-frequency cepstral coefficients (MFCC), etc. The combined use of these features helps to describe speech information more comprehensively and accurately.
[0018] Human voice confidence is a value calculated based on multi-dimensional speech features using preset classification rules (such as machine learning models or statistical models). It is used to quantify the probability of human voices in the current frame of sound signal. The higher the confidence, the greater the probability of human voices. Human voice presence information is a binary information determined based on the human voice confidence, indicating whether human voices are present. Subsequent processing uses the presence / absence of human voices as a gating condition, without extending it to continuous quantity or multi-state information.
[0019] The loudspeaker reference signal refers to the original audio signal played by the loudspeaker, which serves as the reference for acoustic echo cancellation processing.
[0020] Acoustic echo cancellation is a technique designed to eliminate acoustic echoes caused by sound played from a speaker in a microphone-picked signal. Its core principle is to model the speaker reference signal using an adaptive filter to estimate and eliminate echo components. The adaptive update strategy refers to the method and intensity of adaptive filter coefficient updates. For example, freezing adaptive filter coefficient updates means stopping the update of filter coefficients under specific conditions to avoid misinterpreting human voices as echoes. Reducing the update intensity of the adaptive filter involves decreasing the step size during updates, making the filter converge more slowly to improve stability.
[0021] Fluctuating noise refers to noise in the echo cancellation output signal whose intensity, dominant frequency, or spectral distribution changes irregularly over time, such as air conditioning airflow noise or fan noise. Identifying this type of noise is crucial for subsequent accurate suppression. The main noise band refers to the frequency range in which fluctuating noise is mainly distributed. It serves as the key band for noise suppression processing and is used to guide the allocation of frequency band resources for narrowband suppression and / or adaptive gain suppression, rather than exclusively limiting the processing to this band.
[0022] Noise suppression is a technique designed to reduce or eliminate noise components in echo cancellation output signals to improve speech clarity; narrowband suppression refers to suppressing noise within a specific narrow frequency range; adaptive gain suppression dynamically adjusts the gain of different frequency bands based on the real-time characteristics of the noise. For example, methods based on spectral subtraction or Wiener filtering can be used to estimate the gain and adjust the suppression intensity of short-time spectra.
[0023] State information refers to parameters generated during acoustic echo cancellation processing that describe the current operating state of the system. For example, residual echo characterization can quantify the energy or intensity of residual echoes that still exist after echo cancellation. Update parameters refer to specific values used to adjust adaptive update strategies, such as the step size factor of an adaptive filter. Suppression strength refers to the degree of noise suppression processing.
[0024] The speech enhancement method for a loudspeaker in this application first acquires the sound signal picked up by the microphone of the loudspeaker. The sound signal can be picked up by a single microphone or a microphone array. For example, an omnidirectional microphone can be used to pick up the overall sound in the conference room, or multiple directional microphones can be used to capture sound from different directions. The acquired sound signal is then sent to a processing module for frame segmentation processing. Frame segmentation processing can divide the continuous audio stream into short frames, for example, each frame is 20 milliseconds, and there can be 50% overlap between frames.
[0025] Multi-dimensional speech features are extracted from each frame of audio signal, including time-domain features (such as short-time energy and zero-crossing rate), frequency-domain features (such as spectral amplitude, spectral centroid, and spectral bandwidth), and cepstral features (such as Mel-frequency cepstral coefficients, MFCC). For example, the short-time Fourier transform (STFT) of each frame can be calculated to obtain spectral information, and MFCC features can be extracted based on this. Based on these multi-dimensional speech features, a human voice confidence score is output according to a preset classification rule, which can be a classifier based on a support vector machine (SVM) or a neural network. Human voice presence information is generated based on the output human voice confidence score; for example, when the human voice confidence score is higher than a certain threshold, it is determined that a human voice is present.
[0026] Next, the speaker reference signal is obtained. The speaker reference signal is the original audio signal played by the loudspeaker and can be directly obtained from the audio output interface of the loudspeaker. Based on the speaker reference signal, acoustic echo cancellation processing is performed on the sound signal to obtain the echo-cancelled output signal. The acoustic echo cancellation processing can use an adaptive filtering algorithm, such as the Normalized Least Mean Square (NLMS) algorithm or the Recursive Least Squares (RLS) algorithm. By adjusting the adaptive filter coefficients, the output of the speaker reference signal after filtering is made to approximate the echo component in the microphone pickup signal, and the estimated echo is subtracted from the microphone pickup signal.
[0027] When the presence of human voices indicates the presence of human voices, the adaptive update strategy of the acoustic echo cancellation processing is adjusted according to the confidence level of the human voices. For example, when the confidence level of human voices is high, the update of the adaptive filter coefficients is frozen, or the update intensity of the adaptive filter is reduced (such as reducing the step size factor of the NLMS algorithm) to improve the stability when human voices are present.
[0028] Subsequently, fluctuating noise in the echo cancellation output signal that meets preset judgment conditions, preset intensity, and / or dominant frequency changes over time is identified. The identification of fluctuating noise is achieved by analyzing the short-time spectrum, the amplitude of energy fluctuation over time, and the drift characteristics of the dominant frequency. For example, when the energy in a certain frequency band fluctuates significantly and irregularly in a short period of time, it is identified as fluctuating noise. Based on the identification results of fluctuating noise, the main noise frequency band of the echo cancellation output signal is determined. For example, air conditioning airflow noise may be concentrated in the low-frequency region.
[0029] Noise suppression processing is performed on the main noise frequency bands to obtain a noise-suppressed output signal. The noise suppression processing includes narrowband suppression and / or adaptive gain suppression, such as designing narrowband filters to suppress the identified main noise frequency bands, or dynamically adjusting the gain of different frequency bands according to the real-time characteristics of the noise. When low-frequency fluctuating noise is identified, a larger suppression gain can be applied to the low-frequency region, and the frequency band gain can be dynamically estimated and updated using methods based on spectral subtraction or Wiener filtering.
[0030] Finally, the state information of the acoustic echo cancellation process is obtained. The state information includes at least a residual echo characterization quantity, which is used to quantify the level of residual echo that still exists after echo cancellation. For example, the cross-correlation or energy ratio between the echo cancellation output signal and the loudspeaker reference signal can be calculated within a preset time window and / or a preset frequency band to form a reproducible residual echo characterization quantity.
[0031] Based on the aforementioned adaptive update strategy adjustment that freezes or reduces update intensity based on human voice confidence, the update parameters (such as step size factor) of the adaptive update strategy used for subsequent processing are further adjusted according to the recognition results of human voice presence information, state information, and fluctuating noise. For example, when the human voice presence information indicates the presence of human voice and the residual echo representation is high, the update parameters are tuned more cautiously to avoid accidentally damaging human voice. When the fluctuating noise recognition result indicates the presence of strong noise, the suppression intensity of the noise suppression process is increased. And the amplified speech signal after speech enhancement is output based on the noise suppression output signal.
[0032] In another embodiment of this application, after outputting the confidence level of human voice based on multi-dimensional speech features and a preset classification rule, the method further includes: S1100: Acquire video streams from pre-deployed cameras in the conference area; S1101: Perform visual analysis on the video stream to generate visual voice confidence, wherein the visual analysis includes at least detecting facial pose and detecting lip movements; S1102: The confidence of human voice and the confidence of visual human voice are fused to obtain the confidence of fused human voice, and when the confidence of human voice is in a preset fuzzy range, the confidence of fused human voice is improved according to the confidence of visual human voice. S1103: When the confidence level of the fused human voice meets the preset confidence threshold, the confidence level of the fused human voice is used to replace the human voice confidence level to adjust the adaptive update strategy, so that the adaptive update strategy at least performs the freeze of adaptive filter coefficient updates and / or reduces the update intensity of the adaptive filter.
[0033] Specifically, after acquiring the sound signal picked up by the microphone of the loudspeaker and performing frame segmentation processing, extracting multi-dimensional speech features from each frame of the sound signal, and outputting the voice confidence score based on the multi-dimensional speech features through preset classification rules, this method further introduces visual information: First, it acquires the video stream collected by the camera pre-deployed in the conference area; the camera can be one or more, and its deployment position should ensure that it can effectively capture the face and lip-reading information of the conference participants.
[0034] Subsequently, visual analysis is performed on the video stream to generate visual voice confidence; the visual analysis includes at least detecting facial pose and detecting lip movements, wherein detecting facial pose is used to determine whether there is a person facing the microphone or communicating, and detecting lip movements is used to indicate whether there is vocalization behavior, and the visual voice confidence is quantified accordingly.
[0035] Furthermore, the voice confidence score (based on audio features) is fused with the visual voice confidence score to obtain a fused voice confidence score. This fusion can employ weighted averaging, Bayesian fusion, or other multimodal fusion algorithms. As a preferred implementation, when the voice confidence score is within a preset fuzzy range, the fused voice confidence score is boosted based on the visual voice confidence score. The preset fuzzy range refers to the intermediate range where the voice confidence score neither explicitly indicates the presence of a voice nor explicitly indicates the absence of a voice, such as between 0.4 and 0.6, to improve the reliability of the fusion judgment when the audio judgment is uncertain.
[0036] Finally, when the confidence level of the fused human voice meets the preset confidence threshold, the confidence level of the fused human voice is used to replace the human voice confidence level to adjust the adaptive update strategy. The preset confidence threshold is usually higher than the upper limit of the fuzzy interval, such as 0.7 or 0.8. When the confidence level of the fused human voice reaches this threshold, the adaptive update strategy performs at least freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter: freezing the adaptive filter coefficient update means stopping or significantly slowing down the adjustment of the filter coefficients during the presence of human voice to avoid the human voice signal being misjudged as an echo and canceled; reducing the update intensity of the adaptive filter means reducing the step size of the filter coefficient adjustment to reduce the risk of false cancellation while maintaining a certain degree of adaptability.
[0037] This application's solution overcomes the limitations of relying solely on audio features for voice detection by introducing visual information. When the voice confidence level is in an ambiguous range, the system may still perform adaptive filter updates when a voice is present, mistakenly treating the voice as an echo and causing speech distortion; or it may incorrectly freeze updates when a voice is absent, resulting in slower echo cancellation convergence. This solution provides independent evidence by acquiring video streams from the conference area and performing visual analysis (e.g., detecting facial pose and lip movements). When the audio voice confidence level is ambiguous but the visual voice confidence level is high, the fused voice confidence level is improved, making the adjustment basis of the adaptive update strategy more reliable.
[0038] In some preferred embodiments, the following specific example illustrates the situation: Imagine a meeting scenario where a loudspeaker is in operation, and the microphone picks up a sound signal containing human voices, echoes, and background noise. At a certain moment, due to the presence of background noise (such as air conditioning noise or keyboard typing), the system calculates a human voice confidence score of 0.55 based on multi-dimensional speech features, which falls within a preset ambiguity range (e.g., 0.4 to 0.6). In this case, it is difficult to determine whether someone is speaking based solely on the audio information.
[0039] The proposed solution simultaneously acquires the video stream of the meeting area; through visual analysis of the video stream, the system detects that a participant in the meeting area is making lip movements and that their face is facing the microphone, thereby generating a visual voice confidence score of 0.8.
[0040] Since the confidence level of the audio voice is in the fuzzy range, the system will improve the confidence level of the fused voice based on the confidence level of the visual voice. For example, through weighted fusion, the confidence level of the fused voice is improved to 0.75, which meets the preset confidence level threshold (e.g., 0.7).
[0041] Therefore, the system will replace the original audio voice confidence with a fused voice confidence of 0.75 to adjust the adaptive update strategy for acoustic echo cancellation processing. Since the fused voice confidence indicates the presence of voice, the adaptive update strategy will be adjusted to freeze the adaptive filter coefficient update to prevent the voice signal from being mistaken for an echo and cancelled.
[0042] In this way, even when the audio information is unclear, the system can more accurately determine the presence of human voices and take the correct adaptive update strategy, protecting the integrity of human voices, avoiding speech distortion, and ensuring the effectiveness of echo cancellation.
[0043] In another embodiment of this application, it is further proposed that the microphone includes at least two microphone channels, and S4000 includes: S4100: Acquires structural vibration signals from vibration sensors pre-deployed at the lightweight structure of the conference room and air pressure fluctuation signals from air pressure sensors pre-deployed in the conference room and / or air conditioning outlet area. S4101: Identify the occurrence of ultra-low frequency excitation based on air pressure fluctuation signals, identify the occurrence of structural resonance based on structural vibration signals, and confirm the existence of the physical cause of secondary noise when ultra-low frequency excitation and structural resonance are simultaneously satisfied. S4102: Perform multi-channel spectral analysis on the audio signal to obtain abnormal fluctuation characteristics, and calculate the spatial correlation of abnormal fluctuation characteristics between different microphone channels to determine spatial inconsistency anomalies. Spatial inconsistency anomalies are defined as at least two microphone channels having a spatial consistency index lower than a preset threshold within at least one frequency band corresponding to the abnormal fluctuation characteristics. The spatial consistency index includes at least the inter-channel correlation coefficient and / or amplitude squared coherence. S4103: The results of confirming the existence of the physical cause of secondary noise are fused with the spatial inconsistency anomaly to confirm the existence of secondary noise, and when the existence of secondary noise is confirmed, the noise frequency band corresponding to the secondary noise is determined. S4104: Using the noise frequency band as at least a part of the main noise frequency band, perform inverse filtering on the echo cancellation output signal within the main noise frequency band, and dynamically adjust the center frequency and / or bandwidth of multiple narrowband filters to suppress the spectral region corresponding to secondary noise, so as to obtain a noise-suppressed output signal. S4105: Calculate the auditory comfort characterization quantity for the noise suppression output signal and compare it with the preset comfort threshold to obtain the comfort evaluation result; S4106: When the comfort assessment results indicate that the preset comfort conditions are not met, adjust the parameters of the inverse filtering process and / or the suppression parameters of the narrowband filter.
[0044] Specifically, to more comprehensively perceive and identify secondary noise, the microphone of the loudspeaker can include at least two microphone channels for spatial analysis; vibration sensors can be deployed on lightweight structures in the conference room (such as walls, ceilings, tabletops, etc.) to pick up structural vibration signals caused by external vibration sources (such as air conditioner outdoor units, traffic noise, floor vibrations, etc.); and barometric pressure sensors can be deployed in the conference room or air conditioning vent area to detect minute barometric pressure fluctuation signals caused by airflow disturbances or low-frequency sound waves.
[0045] Among them, ultra-low frequency excitation can be understood as sound waves or vibrations with extremely low frequencies (usually below 20Hz), which may be difficult for the human ear to perceive directly but will affect the structure; structural resonance refers to the structure producing large-amplitude vibrations when the external excitation frequency matches the natural frequency of the structure; when ultra-low frequency excitation and structural resonance are satisfied at the same time, the physical cause of secondary noise can be confirmed, serving as physical evidence for the identification of secondary noise.
[0046] In practical applications, multi-channel spectral analysis of audio signals can reveal abnormal fluctuation characteristics, such as abnormal energy peaks or persistent fluctuations within a specific frequency band. By calculating the spatial correlation of abnormal fluctuation characteristics between different microphone channels, spatial inconsistency anomalies can be identified. Spatial inconsistency anomalies refer to the low spatial consistency of audio signals picked up by at least two microphone channels within one or more frequency bands, such as a channel correlation coefficient or amplitude squared coherence lower than a preset threshold. This spatial inconsistency may be related to the location or propagation path of the noise source being different from that of the target speech.
[0047] Furthermore, by fusing the confirmation results of the physical causes of secondary noise (from vibration sensors and barometric pressure sensors) with the judgment results of spatial inconsistency anomalies, the existence of secondary noise can be more accurately confirmed; when the existence of secondary noise is confirmed, its corresponding noise frequency band is determined.
[0048] In a preferred embodiment, the determined secondary noise frequency band is taken as at least a part of the main noise frequency band, and the echo cancellation output signal is subjected to inverse filtering within this frequency band. The inverse filtering is designed to generate a filtered output that cancels out the secondary noise components, and the center frequency and / or bandwidth of multiple narrowband filters can be dynamically adjusted to accurately suppress the spectral region corresponding to the secondary noise.
[0049] In addition, to ensure the listening comfort of the enhanced speech, a listening comfort characteristic is calculated for the noise-suppressed output signal, such as based on psychoacoustic parameters like loudness, sharpness, and roughness. This characteristic is compared with a preset comfort threshold to obtain a comfort evaluation result. When the evaluation result indicates that the preset comfort conditions are not met, the parameters of the inverse filtering process (such as filter order and coefficients) and / or the suppression parameters of the narrowband filter (such as suppression depth and bandwidth) are adjusted to effectively suppress noise while maximizing the naturalness and listening comfort of the speech.
[0050] This application addresses the problems of inaccurate identification, incomplete suppression, and potential impact on listening comfort in traditional speech enhancement methods when dealing with complex secondary noise. It achieves secondary noise identification and localization by acquiring physical cause information through multimodal sensors and combining it with spatial analysis using multi-channel microphones, thus improving the robustness of noise identification. It employs inverse filtering and dynamic narrowband filtering to finely suppress secondary noise, reducing damage to the target speech. Furthermore, it introduces a listening comfort assessment and feedback mechanism, enabling noise suppression processing to adaptively optimize based on actual listening experience, providing clear, natural, and comfortable speech output even in noisy environments, thereby improving the overall speech enhancement performance of the loudspeaker and the user experience.
[0051] In another embodiment of this application, it is further proposed that, based on the identification results of human voice presence information, state information, and fluctuating noise, the update parameters of the adaptive update strategy used for subsequent processing and / or the suppression intensity of the noise suppression processing are adjusted, specifically including: S5100: Acquire multiple performance indicators corresponding to the echo cancellation output signal and / or noise suppression output signal. The performance indicators include speech clarity, echo suppression level, noise suppression level, and speech naturalness. The performance indicators are calculated based on the echo cancellation output signal and / or noise suppression output signal according to preset calculation rules. S5101: Determine the real-time environmental state information based on the human voice presence information, the residual echo characterization quantity in the state information, and the identification results of fluctuation noise; S5102: Based on performance indicators and environmental status information, dynamically adjust the weights and / or priorities of speech clarity indicators, echo suppression level, noise suppression level, and speech naturalness indicators to obtain adjustment parameters; S5103: Perform multi-objective optimization decision based on the adjustment parameters to determine the processing flow of the sound signal and the combination of related parameters. The related parameters include at least the update parameters of the adaptive update strategy and / or the suppression intensity of the noise suppression processing. The processing flow includes at least the synergistic method of acoustic echo cancellation processing and noise suppression processing.
[0052] Among them, several performance indicators refer to various metrics used to quantify the output quality of the speech enhancement system. For example, the speech intelligibility indicator measures the clarity and / or intelligibility of the processed speech; the echo suppression level is used to evaluate the thoroughness of echo cancellation; the noise suppression level reflects the degree to which background noise is reduced; and the speech naturalness indicator focuses on whether the processed speech sounds natural and distortion-free. These indicators are obtained by analyzing the echo cancellation output signal and / or noise suppression output signal based on preset calculation rules.
[0053] Specifically, real-time environmental state information is a comprehensive description of the current acoustic environment characteristics. It combines the judgment of whether human voices are present, the quantitative information of residual echoes in acoustic echo cancellation processing, and the identification results of fluctuating noise. For example, environmental state information may include scenarios such as active human voices with strong echoes, or no human voices but fluctuating noise.
[0054] Dynamically adjusting weights and / or priorities means that the system will assign different levels of importance to the target that needs to be optimized most at the moment, depending on the environmental conditions. For example, in an environment with active human voices and strong echoes, the priority of echo suppression level will be increased; while in an environment with weak human voices and high background noise, the priority of speech intelligibility index and noise suppression level will be increased.
[0055] Multi-objective optimization decision-making seeks the optimal combination of parameters and processing flow under the guidance of the dynamically adjusted weights and / or priorities mentioned above, in order to achieve a balance among multiple conflicting performance objectives. Relevant parameters may include, for example, the step size of the adaptive filter, the gain factor of noise suppression, etc., while the coordination method of the processing flow can refer to whether the acoustic echo cancellation processing and noise suppression processing are executed serially, in parallel, or interact with each other through some feedback mechanism.
[0056] In some preferred embodiments, this application is implemented as follows: Assuming a meeting scenario with a loudspeaker in operation, the system continuously acquires sound signals picked up by the microphone and processes them frame by frame, extracting time-domain, frequency-domain, and cepstral features. It outputs voice confidence and generates voice presence information. Simultaneously, it acquires the speaker reference signal, performs acoustic echo cancellation processing to obtain an echo-cancelled output signal, identifies fluctuating noise in the echo-cancelled output signal, determines the main noise frequency band, and performs noise suppression processing on the main noise frequency band to obtain a noise-suppressed output signal. During parameter adjustment, multiple performance metrics are calculated in real time, such as speech clarity using PESQ (Perceptual Evaluation of Speech Quality), echo suppression level using ERLE (Echo Return Loss Enhancement), noise suppression level using SNR (Signal-to-Noise Ratio) gain, and MOS (Mean Opinion Metrics)... The system uses a predictive model to evaluate the speech naturalness index. Then, based on the presence of human voices, the status information of acoustic echo cancellation processing (e.g., whether the residual echo representation is high), and the identification results of fluctuating noise (e.g., whether there is air conditioning noise or projector fan noise), it determines real-time environmental status information, such as active human voices, strong echoes, and continuous noise. Based on performance indicators and environmental status information, the system dynamically adjusts the weights and / or priorities of various performance objectives. For example, under conditions of active human voices, strong echoes, and continuous noise, it increases the weights of echo suppression level and speech clarity index, while appropriately reducing the weight of the speech naturalness index to allow for stronger processing. Subsequently, a multi-objective optimization decision algorithm is executed to search for the optimal parameter combination. For example, the step size factor of the adaptive update strategy is appropriately reduced to avoid interference from human voices on echo cancellation, while the gain factor of noise suppression processing is increased to suppress air conditioning noise. Acoustic echo cancellation processing and noise suppression processing are set to work in parallel to maximize overall performance. Parallel collaboration refers to concurrent computation or pipelined execution of modules, while noise suppression is still based on the synchronously / delay-aligned echo cancellation output signal.
[0057] In another embodiment of this application, an adaptive update strategy for acoustic echo cancellation processing is further proposed, which adjusts the acoustic echo cancellation processing based on the voice confidence level when the voice presence information indicates the presence of voice. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter, including: S2100: Periodically plays a preset test signal through a speaker and picks up the reflected sound of the preset test signal in the acoustic environment through the microphone of the amplifier to obtain the reflected sound signal; S2101: Analyze the reflected acoustic signal to evaluate the acoustic reflection path, obtain the path evaluation result, and predict the evolution trend of the acoustic reflection path based on the path evaluation result over a continuous preset time period. S2102: Based on the evolution trend and the presence information of human voices, adjust the adaptive update strategy of the acoustic echo cancellation processing. The adjustment includes at least the following: when the change in the predicted acoustic reflection path exceeds the preset change threshold and the presence information of human voices indicates that there are no human voices, increase the update intensity of the adaptive update strategy to accelerate convergence; when the presence information of human voices indicates that there are human voices, freeze the update of the adaptive filter coefficients and / or reduce the update intensity of the adaptive filter. S2103: When adjusting the adaptive update strategy, a protective gain is applied to the echo cancellation output signal in the human voice dominant frequency region based on the human voice dominant frequency region determined by the multi-dimensional speech features. S2104: Based on the residual echo characterization in the state information, evaluate the residual echo level, and when the residual echo level is higher than the preset echo threshold, perform re-adaptive control on the acoustic echo cancellation processing based on the reflected sound signal. The re-adaptive control includes at least re-initializing the adaptive filter coefficients and / or further increasing the update strength of the adaptive update strategy when the human voice presence information indicates that there is no human voice.
[0058] Specifically, a preset test signal is periodically played through a loudspeaker, and the reflected sound of the preset test signal in the acoustic environment is picked up by the microphone of the loudspeaker to obtain the reflected sound signal. The preset test signal can be a short-duration, wideband signal, such as white noise, pink noise, or a swept-frequency signal, the purpose of which is to detect the impulse response of the acoustic environment. Periodic playback means repeating the test signal at certain intervals in order to continuously monitor changes in the acoustic environment. The reflected sound signal is the signal containing the ambient echo picked up by the microphone when the test signal is played.
[0059] Furthermore, the reflected acoustic signals are analyzed to evaluate the acoustic reflection path, obtaining path evaluation results. Based on the path evaluation results over a continuous preset time period, the evolution trend of the acoustic reflection path is predicted. The acoustic reflection path refers to the path of sound emitted by the loudspeaker after reflection through the room to the microphone. By analyzing the reflected acoustic signals, such as calculating their impulse response or frequency domain transfer function, the current acoustic reflection path can be evaluated. The path evaluation results can be a quantitative description of the characteristics of the acoustic reflection path, such as echo delay, attenuation, and frequency response characteristics. Based on the path evaluation results over a continuous preset time period, time series analysis, Kalman filtering, and other methods can be used to predict the future trend of the acoustic reflection path, such as whether it tends to stabilize, changes slowly, or changes drastically.
[0060] Specifically, based on the evolution trend and the presence of human voices, the adaptive update strategy for acoustic echo cancellation processing is adjusted. The adjustment includes at least the following: when the change in the predicted acoustic reflection path exceeds a preset change threshold, and the presence of human voices indicates the absence of human voices, increasing the update intensity of the adaptive update strategy to accelerate convergence; when the presence of human voices indicates the presence of human voices, freezing the update of the adaptive filter coefficients and / or reducing the update intensity of the adaptive filter; the preset change threshold is used to determine the drastic degree of change in the acoustic reflection path; when the environment changes drastically and there are no human voices, the step size factor of the adaptive filter can be increased or a more aggressive update algorithm can be adopted to enable the filter to quickly adapt to the new acoustic environment; when human voices are present, to avoid distortion of the human voice signal, the update of the adaptive filter coefficients is frozen or the update intensity is reduced.
[0061] In practical applications, when adjusting the adaptive update strategy, a protective gain is applied to the echo cancellation output signal in the human voice dominant frequency region determined by multi-dimensional speech features. The human voice dominant frequency region refers to the frequency range in which human voice energy is mainly concentrated, which can usually be obtained by spectral analysis of multi-dimensional speech features. The purpose of applying the protective gain is to prevent residual echoes or noise from excessively suppressing human voice when human voice is present, especially when the adaptive filter update intensity decreases or freezes, thereby protecting the integrity and clarity of human voice.
[0062] Furthermore, based on the residual echo representation in the state information, the residual echo level is evaluated, and when the residual echo level is higher than a preset echo threshold, re-adaptive control is performed on the acoustic echo cancellation processing based on the reflected sound signal. The residual echo representation is the state information output by the acoustic echo cancellation processing module, used to quantify the residual echo energy or perception after echo cancellation. When the residual echo level is too high, it indicates that the current adaptive filter may not have effectively eliminated the echo, and re-adaptive control is required. Re-adaptive control includes at least re-initializing the adaptive filter coefficients and / or further increasing the update strength of the adaptive update strategy when the human voice presence information indicates that there is no human voice. Re-initialization allows the filter to start learning again from a known state, while increasing the update strength helps the filter converge to a new optimal state more quickly.
[0063] This application improves the echo cancellation performance and speech quality of loudspeaker speech enhancement methods in complex dynamic acoustic environments. By dynamically evaluating and predicting acoustic reflection paths, the system can respond to environmental changes earlier and more accurately, avoiding the convergence lag of adaptive filters caused by environmental changes in traditional methods. When human voices are present, protective gains are applied in the dominant frequency region of human voices to avoid excessive suppression of human voices, ensuring the naturalness and clarity of speech. At the same time, real-time monitoring of residual echo levels and re-adaptive control mechanisms ensure that even in extreme cases, insufficient echo cancellation can be corrected in a timely manner, maintaining high-quality speech enhancement effects. This enables loudspeakers to provide a stable, clear, and echo-free speech experience in scenarios such as meetings and teaching.
[0064] In another embodiment of this application, it is further proposed that when adjusting the adaptive update strategy of acoustic echo cancellation processing according to the evolution trend and human voice presence information, the adjustment also includes: S2105: When the evolution trend indicates that the change in the acoustic reflection path exceeds the preset change threshold, and the human voice presence information indicates the presence of low-energy human voice, the update parameters of the adaptive update strategy used for acoustic echo cancellation processing, the gain parameters of the protective gain, and the judgment parameters of the preset classification rule are adjusted in a coordinated manner. The update parameters include at least the step size factor of the adaptive filter. The low-energy human voice is the human voice whose human voice energy representation based on the short-time energy in the multi-dimensional speech features is lower than the preset human voice threshold. The human voice energy representation is the short-time energy calculated based on the multi-dimensional speech features and / or its logarithmic form. S2106: The coordinated adjustment includes at least: reducing the step size factor and / or increasing the protective gain of the dominant frequency region of human voices during the period when the presence information of human voices indicates the presence of low-energy human voices, and adjusting the confidence threshold and / or fuzzy interval boundary of the preset classification rule to improve the sensitivity of low-energy human voices.
[0065] Specifically, when the change in the acoustic reflection path exceeds a preset threshold and the system detects the presence of low-energy human voices, a collaborative adjustment mechanism will be activated. Low-energy human voices can be understood as those whose human voice energy representation, calculated based on multi-dimensional speech features extracted from the sound signal (such as short-time energy or its logarithmic form), is lower than a preset threshold. Such low-energy human voices typically refer to those speaking softly or from a distance from the microphone. The goal of collaborative adjustment is to maintain the effectiveness of echo cancellation while protecting low-energy human voices. The updated parameters of collaborative adjustment include at least the step size factor of the adaptive filter, which is a key parameter affecting the convergence speed and stability of the adaptive filter. The human voice energy representation is a representation calculated based on the short-time energy and / or its logarithmic form from the multi-dimensional speech features, used to quantify the energy level of the human voice.
[0066] Specific measures for coordinated adjustment include: First, reducing the step size factor of the adaptive filter. A reduced step size factor means a weaker update intensity of the adaptive filter coefficients, thereby reducing potential interference to human voices during the presence of low-energy human voices. Second, increasing the protective gain in the frequency region dominated by human voices. The protective gain aims to prevent human voices from being excessively attenuated during echo cancellation or noise suppression. Increasing this gain can provide stronger protection for low-energy human voices. In addition, adjusting the confidence threshold and / or fuzzy interval boundaries of the preset classification rules. The preset classification rules are used to output the confidence of human voices based on multi-dimensional speech features. Adjusting their thresholds and boundaries can make the system more sensitive to the determination of low-energy human voices and reduce the risk of them being misclassified as non-human voices.
[0067] This application improves the speech enhancement performance of loudspeakers in complex acoustic environments, especially suitable for scenarios with dynamically changing acoustic reflection paths and the presence of low-energy human voices. By synergistically adjusting the update parameters of the adaptive update strategy, the gain parameters of the protective gain, and the judgment parameters of the preset classification rules, the system can more finely balance echo cancellation, noise suppression, and human voice protection. Reducing the step size factor helps stabilize the adaptive filter and reduce the negative impact on human voices when low-energy human voices are present. Increasing the protective gain enhances the protection against low-energy human voices. Adjusting the threshold and boundary of the classification rules improves the detection capability of low-energy human voices and reduces the risk of missed detections or false positives. These synergistic measures work together to ensure that the output speech maintains high clarity and naturalness when low-energy human voices and dynamic environments coexist.
[0068] In another embodiment of this application, a gain parameter for coordinating the adjustment of the protective gain is further proposed, including: S2105-1: Based on multi-dimensional speech features, obtain the human voice spectrum feature representation quantity and acquire the environmental noise information represented by the recognition results of fluctuating noise. The environmental noise information includes at least the noise type, intensity, spectrum distribution and transient fluctuation characteristics. The human voice spectrum feature representation quantity is a spectrum feature vector representing the distribution of human voice energy with frequency. S2105-2: When the human voice presence information indicates the presence of low-energy human voice, and the environmental noise information indicates the presence of transient noise that overlaps with the spectral characteristics of human voice and has transient fluctuation characteristics, the gain parameter of the protective gain and / or the suppression parameter of the noise suppression processing shall be adjusted in the first frequency band. S2105-3: The first frequency band adjustment includes at least: increasing the protective gain and / or reducing the suppression intensity of noise suppression processing in the human voice-dominated frequency region; and increasing the suppression intensity of noise suppression processing in the non-human voice-dominated frequency region according to the transient fluctuation characteristics to suppress transient noise.
[0069] Specifically, the human voice spectral feature representation refers to the spectral feature vector calculated based on the short-time spectrum of multi-dimensional speech features, used to characterize the distribution of human voice energy at different frequencies. This vector can reflect the frequency components and energy intensity of the human voice at a specific moment. For example, it can be represented by the spectral envelope, Mel-frequency cepstral coefficients (MFCC), or other spectral-related features extracted after Fourier transforming the sound signal. Its purpose is to accurately characterize the frequency characteristics of the human voice so that it can be distinguished and matched with noise in the future. Among them, the environmental noise information is obtained based on the identification results of fluctuating noise, which includes noise type (such as fan noise, keyboard typing sound, environmental humming, etc.), intensity, spectral distribution, and transient fluctuation characteristics. "Sensitivity" refers to the rapid change in the energy or frequency components of noise within a short period of time, such as a sudden knocking sound or a brief friction sound. When the presence of human voice information indicates the presence of low-energy human voice, and the presence of transient noise information indicates the presence of transient noise that overlaps with the spectral characteristics of human voice and has transient fluctuation characteristics, a first frequency band adjustment is performed on the gain parameter of the protective gain and / or the suppression parameter of the noise suppression processing. The first frequency band adjustment includes increasing the protective gain and / or decreasing the suppression intensity of the noise suppression processing in the frequency region dominated by human voice to protect low-energy human voice and prevent speech distortion or loss. At the same time, in the frequency region not dominated by human voice, the suppression intensity of the noise suppression processing is increased according to the transient fluctuation characteristics to suppress transient noise and reduce its impact on human voice.
[0070] The solution proposed in this application, by introducing human voice spectral characteristic parameters and environmental noise information (especially transient fluctuation characteristics), can more accurately identify the relationship between low-energy human voice and overlapping transient noise, thereby performing refined first-band adjustment; in the human voice dominant frequency region, by increasing the protective gain or reducing the noise suppression intensity, the integrity of low-energy human voice is ensured and excessive suppression is avoided; in the non-human voice dominant frequency region, the noise suppression intensity is increased according to the characteristics of transient noise, effectively eliminating transient interference, and achieving a balance between human voice protection and noise suppression.
[0071] In some preferred embodiments, the following specific example illustrates the situation: Imagine a meeting scenario where a participant is speaking softly (generating low-energy vocals), while occasional keyboard clicks and slight airflow noise from an air conditioner vent (transient noise) occur in the environment. The spectra of these transient noises may overlap with certain frequency components of the low-energy vocals and exhibit rapidly changing transient fluctuations. First, the system extracts multi-dimensional speech features from the microphone-picked sound signal to obtain a vocal spectral feature representation, thus depicting the frequency distribution of the low-energy vocals. Simultaneously, based on the identification results of fluctuating noise, it acquires environmental noise information and identifies the type, intensity, spectral distribution, and transient fluctuation characteristics of the keyboard clicks and airflow noise. When the system determines the presence of a low-energy vocal and identifies a frequency feature representation that matches that vocal spectral feature... When there is overlapping transient noise, such as keyboard typing sounds mainly concentrated in the mid-to-high frequencies while low-energy human voices have energy distribution in the mid-to-low frequencies and some mid-to-high frequencies, the system performs the first crossover adjustment: in the frequency region dominated by human voices (e.g., 500Hz to 2000Hz), the protective gain is increased and / or the suppression intensity of the noise suppression process is reduced to preserve human voice details; in the frequency region not dominated by human voices (e.g., below 200Hz or above 2000Hz, and the mid-to-high frequency region where human voice energy is weak but transient noise is significant), the suppression intensity of the noise suppression process is increased according to the transient fluctuation characteristics, for example, applying stronger narrowband suppression to the transient high-frequency components of keyboard typing sounds, thereby effectively suppressing overlapping transient noise while protecting low-energy human voices.
[0072] In another embodiment of this application, it is further proposed that the preset classification rule also includes a threshold parameter for distinguishing between echoes and human voices, and the method further includes: S2200: Based on the loudspeaker reference signal and the echo cancellation output signal, calculate the echo similarity representation quantity and / or the echo energy ratio representation quantity. The echo similarity representation quantity is a scalar or frequency band vector that represents the correlation or coherence between the loudspeaker reference signal and the echo cancellation output signal. The echo energy ratio representation quantity is a scalar or frequency band vector that represents the proportion or ratio of the echo component energy related to the loudspeaker reference signal in the echo cancellation output signal. S2201: When the presence of human voice information indicates the presence of human voice, and the human voice energy representation is lower than the preset human voice threshold, and the echo similarity representation and / or echo energy ratio representation meet the preset echo judgment conditions based on threshold parameters, the threshold parameters on which the distinction between echo and human voice is based are adjusted in the second frequency band. The threshold parameters include the signal energy threshold and / or the similarity threshold corresponding to the echo similarity representation based on the speech activity judgment. S2202: The second frequency band adjustment includes at least: in the human voice-dominated frequency region, lowering the signal energy threshold and / or raising the similarity threshold; in the non-human voice-dominated frequency region, raising the signal energy threshold and / or lowering the similarity threshold; S2203: Based on the results of the second frequency band adjustment, adjust the adaptive update strategy of the acoustic echo cancellation processing and / or the gain parameters of the protective gain.
[0073] Specifically, the preset classification rules include threshold parameters for distinguishing between echoes and human voices. These threshold parameters are pre-set values used to determine whether the current signal component is a human voice or an echo during speech signal processing. For example, the threshold parameters include a signal energy threshold to determine whether the signal meets the minimum energy requirement for human voices, and a similarity threshold to determine whether the echo similarity between the signal and the speaker reference signal exceeds a certain threshold.
[0074] The echo similarity characterization is used to quantify the similarity or correlation between the loudspeaker reference signal and the echo cancellation output signal to identify echo components. For example, it can be obtained by calculating the cross-correlation coefficient, coherence function, or amplitude squared coherence in the frequency domain of the two signals. The echo energy ratio characterization is used to characterize the proportion of energy of the echo component related to the loudspeaker reference signal in the echo cancellation output signal to evaluate the echo intensity. For example, it can be obtained by calculating the ratio of the residual echo energy related to the loudspeaker reference signal in the echo cancellation output signal to the total energy of the echo cancellation output signal. The echo similarity characterization and / or echo energy ratio characterization can be scalar values or frequency band vectors reflecting different frequency band characteristics.
[0075] In practical applications, when the presence of human voice information indicates the presence of human voice, and the human voice energy representation is lower than a preset human voice threshold (i.e., low-energy human voice exists), and the calculated echo similarity representation and / or echo energy ratio representation meet the preset echo judgment conditions based on the current threshold parameters, a second frequency band adjustment is triggered on the threshold parameters on which the distinction between echo and human voice is based; the preset echo judgment conditions include, for example, considering the presence of significant echo when the echo similarity representation is higher than a certain threshold or the echo energy ratio representation is higher than a certain threshold.
[0076] The second crossband adjustment is used to differentiate the threshold parameters in different frequency regions to more accurately distinguish between low-energy human voices and echoes. Specifically, in the frequency region dominated by human voices, the signal energy threshold is lowered so that weaker human voices can still be identified, while the similarity threshold is increased to make the system have higher requirements for echo similarity to avoid misjudging low-energy human voices as echoes. In the frequency region not dominated by human voices, the signal energy threshold is increased to more strictly filter non-human voice components, while the similarity threshold is lowered to improve the sensitivity to echo recognition and strengthen echo suppression.
[0077] Therefore, based on the results of the second frequency band adjustment, the system adjusts the adaptive update strategy and / or the gain parameter of the protective gain in the acoustic echo cancellation processing. For example, in the human voice-dominated frequency region, the update intensity of the adaptive update strategy can be appropriately reduced to protect low-energy human voices, and the gain parameter of the protective gain can be increased to improve clarity. In the non-human voice-dominated frequency region, the update intensity of the adaptive update strategy can be increased to accelerate echo cancellation convergence, and the gain parameter of the protective gain can be reduced to avoid protecting non-human voice components.
[0078] The proposed solution introduces echo similarity and echo energy ratio representations to provide richer evidence of echo presence, enabling more reliable differentiation between low-energy human voices and residual echoes in low-energy human voice scenarios. By adjusting the threshold parameters in a second frequency band, the system lowers the threshold for human voice recognition and raises the threshold for echo determination in the human voice-dominant frequency region, thereby protecting low-energy human voices. Simultaneously, in the non-human voice-dominant frequency region, the system raises the threshold for non-human voice recognition and lowers the threshold for echo determination, thereby more actively suppressing echoes. Combined with the coordinated adjustment of the adaptive update strategy and protective gain, the system can dynamically optimize speech enhancement processing based on real-time signal characteristics.
[0079] In some preferred embodiments, assuming a meeting scenario where a participant is speaking softly with some residual echo, the system identifies the presence of low-energy human voices through multi-dimensional speech features and calculates echo similarity and echo energy ratio metrics based on the speaker reference signal and echo cancellation output signal. When these metrics indicate significant residual echo and the human voice energy is below a preset threshold, a second frequency band adjustment is triggered: in the dominant human voice frequency region (e.g., 500Hz to 3000Hz), the signal energy threshold used for speech activity judgment is lowered and the similarity threshold corresponding to the echo similarity metrics is raised to avoid misjudging low-energy human voices as echoes. In non-human voice dominant frequency regions (e.g., below 500Hz or above 3000Hz), the signal energy threshold is increased and the similarity threshold is decreased to more strictly suppress background noise and non-human voice echo components and enhance echo suppression. Based on the adjusted threshold parameters, the adaptive update strategy of the acoustic echo cancellation processing is adjusted accordingly. For example, the step size factor of the adaptive filter is reduced in the human voice dominant frequency region to protect human voice, while the step size factor is increased in the non-human voice dominant frequency region to accelerate echo convergence. At the same time, the gain parameter of the protective gain is increased or decreased accordingly in the corresponding frequency band, thereby effectively suppressing residual echo and outputting a clear and natural amplified signal while protecting low-energy human voice.
[0080] In another embodiment of this application, a speech enhancement method for a loudspeaker is further proposed, wherein the microphone includes at least two microphone channels, and the method further includes: S2204: Perform spatial analysis on the audio signals from at least two microphone channels to obtain directional information from at least two sound sources; S2205: For each sound source direction, perform spatial filtering to obtain the direction separation signal of the corresponding direction, wherein the direction separation signal is a signal obtained by weighting and synthesizing the sound signals of at least two microphone channels according to the spatial filtering weights corresponding to the sound source direction; S2206: Extract the human voice energy representation and the human voice spectrum feature representation obtained based on multi-dimensional speech features from the separated signals in each direction, and output the human voice confidence level in the corresponding direction based on the preset classification rules. S2207: Based on the confidence level of human voice in each direction and the human voice energy representation in the corresponding direction, adjust the signal energy threshold in the corresponding direction respectively, and use the adjusted signal energy threshold for the distinction between echo and human voice in the corresponding direction and / or for adjusting the adaptive update strategy in the corresponding direction. S2208: When the confidence of human voice in at least two directions meets the preset confidence threshold, apply protective gain to the dominant frequency region of human voice in the corresponding direction.
[0081] Specifically, the microphone includes at least two microphone channels, the purpose of which is to acquire sound field spatial information by utilizing the spatial sampling capabilities of multiple microphones, thereby distinguishing sound sources from different directions; wherein, spatial analysis of the sound signals of at least two microphone channels is to process the multi-channel sound signals to estimate the direction information of the sound sources, for example, by first using angle of arrival (DOA) estimation, independent component analysis (ICA) to obtain the direction of the sound sources, and then using the direction of the sound sources for subsequent spatial filtering processing, thereby identifying and locating the direction information of at least two sound sources in the environment, the direction of the sound sources may include the direction of human voices and / or the direction of noise.
[0082] Furthermore, spatial filtering is performed for each sound source direction to separate the sound source signal from the mixed microphone signal. Spatial filtering can employ techniques such as adaptive beamforming, fixed beamforming, or zero-point constrained beamforming. By weighting and synthesizing the sound signals from at least two microphone channels according to spatial filtering weights corresponding to the sound source direction, a direction separation signal for the corresponding direction is obtained. The direction separation signal is the output signal of the spatial filtering process, which mainly contains speech or noise components from the specific sound source direction.
[0083] Based on this, the voice energy representation and the voice spectrum feature representation obtained based on multi-dimensional speech features are extracted from the signals separated in each direction. The voice confidence of the corresponding direction is output according to the preset classification rules, so as to perform independent voice activity detection and feature extraction for each sound source direction.
[0084] Furthermore, based on the confidence level of human voice in each direction and the corresponding human voice energy representation, the signal energy threshold for each direction is adjusted. The signal energy threshold is a key parameter for judging speech activity and distinguishing between echo and human voice. Directional adjustment makes the detection in different directions more sensitive or robust. For example, the signal energy threshold is lowered for low-energy human voice directions to improve detection sensitivity, and the signal energy threshold is raised for noise-dominant directions to avoid misjudgment. The adjusted signal energy threshold is used for distinguishing between echo and human voice in the corresponding direction and / or for adjusting the adaptive update strategy in the corresponding direction.
[0085] As a preferred implementation, when the confidence of human voices in at least two directions meets the preset confidence threshold, protective gains are applied to the dominant frequency regions of human voices in the corresponding directions to prevent excessive suppression from causing speech distortion or loss in multi-speaker scenarios.
[0086] The proposed solution uses spatial analysis and spatial filtering of multiple microphone channels to separate the mixed signal into multiple directional signals. This allows the calculation of human voice energy representation, human voice spectral feature representation, and human voice confidence to be performed independently by direction. Based on this, the signal energy threshold, adaptive update strategy, and protective gain are adjusted in a directional manner, thereby improving the speech enhancement effect in complex environments with multiple sound sources.
[0087] In some preferred embodiments, the following specific example illustrates the situation: Imagine a large conference room equipped with a microphone array containing four microphone channels. Three attendees (A, B, and C) are positioned at different points on the microphone array. Meanwhile, the air conditioning vent (direction D) in the conference room generates continuous low-frequency noise. Attendee A is giving a presentation with a loud voice; attendee B occasionally interjects with a lower volume; and attendee C is whispering with the person next to them.
[0088] First, the microphone array picks up the sound signals in the conference room. The system performs spatial analysis on these multi-channel sound signals, such as using angle of arrival (DOA) estimation to determine directional information and identify at least three main sound source directions: the direction of participant A, the direction of participant B, the direction of participant C, and the direction of air conditioning noise D.
[0089] Next, spatial filtering is performed for each identified sound source direction to generate a direction separation signal for the corresponding direction. For example, a direction separation signal mainly containing the sound of participant A, a direction separation signal mainly containing the sound of participant B, and a direction separation signal mainly containing the sound of participant C can be generated. A direction separation signal for suppressing air conditioning noise in direction D can also be generated.
[0090] Then, for each direction of the separated signal, the voice energy representation and the voice spectrum feature representation are extracted respectively, and the voice confidence of the corresponding direction is output based on the preset classification rules. For example, the voice confidence of participant A direction is high, the voice confidence of participant B direction is medium to low (because its sound energy is low), the voice confidence of participant C direction is low (because it communicates in a soft voice), and the voice confidence of air conditioner noise D direction is very low.
[0091] Based on the confidence level of human voices in each direction and the corresponding energy representation of human voices in that direction, the system adjusts the signal energy threshold for each direction: for the direction of participant B, the signal energy threshold is lowered to improve the detection sensitivity of low-energy human voices and ensure that their interruptions are accurately identified; for the direction of air conditioning noise D, the signal energy threshold is raised to more effectively suppress noise in that direction; the adjusted signal energy threshold is used to distinguish between echoes and human voices in the corresponding direction and / or to adjust the adaptive update strategy for the corresponding direction.
[0092] Finally, since the voice confidence of participants A and B (or C) meets the preset confidence threshold, the system applies protective gain to the dominant voice frequency region in the direction of participant A and the direction of participant B (or C), respectively. This protects the voice of participant A when they are speaking and participant B interrupts softly, while suppressing noise in the direction of air conditioning noise D and avoiding affecting the voice, resulting in a clear, natural and robust voice enhancement effect.
[0093] Reference Figure 2 The specific embodiments of this application also disclose a voice enhancement system for a loudspeaker, including: The human voice presence information generation module 1 is used to acquire the sound signal picked up by the microphone of the loudspeaker, perform frame-by-frame processing on the sound signal, extract multi-dimensional speech features for each frame of the sound signal, output human voice confidence based on the multi-dimensional speech features and according to the preset classification rules, and generate human voice presence information based on the human voice confidence. The multi-dimensional speech features include at least two of the following: time domain features, frequency domain features and / or cepstral features. The echo cancellation strategy adjustment module 2 is used to acquire the loudspeaker reference signal, perform acoustic echo cancellation processing on the sound signal based on the loudspeaker reference signal to obtain the echo cancellation output signal, and adjust the adaptive update strategy of the acoustic echo cancellation processing according to the human voice confidence level when the human voice presence information indicates the presence of human voice. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter. The noise identification module 3 is used to identify fluctuating noise in the echo cancellation output signal that meets preset judgment conditions, preset intensity and / or dominant frequency changes with time, and to determine the main noise frequency band of the echo cancellation output signal based on the identification results of the fluctuating noise. The noise suppression processing module 4 is used to perform noise suppression processing on the main noise frequency band. The noise suppression processing includes at least narrowband suppression and / or adaptive gain suppression to obtain a noise-suppressed output signal. The status information acquisition module 5 is used to acquire the status information of acoustic echo cancellation processing. The status information includes at least the residual echo characterization quantity. The processing flow coordination module 6 is used to adjust the update parameters of the adaptive update strategy for subsequent processing and / or adjust the suppression intensity of the noise suppression processing based on the human voice presence information, state information and the recognition results of fluctuating noise, and output the amplified speech signal after speech enhancement based on the noise suppression output signal.
[0094] This system aims to address the issue of reduced speech intelligibility in traditional loudspeakers in complex acoustic environments, such as those with ambient noise and acoustic echoes. Through a modular design, the system intelligently senses human voices, dynamically adjusts echo cancellation strategies, accurately identifies and suppresses fluctuating noise, and collaboratively optimizes the entire speech enhancement processing flow. The modules work closely together to ensure that the loudspeaker outputs clear and natural speech signals in various complex scenarios, significantly improving user experience and communication efficiency.
[0095] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for enhancing the speech of a loudspeaker, characterized in that, include: Acquire the sound signal picked up by the microphone of the loudspeaker, perform frame segmentation on the sound signal, extract multi-dimensional speech features for each frame of the sound signal, output human voice confidence based on the multi-dimensional speech features and generate human voice presence information based on the human voice confidence, wherein the multi-dimensional speech features include at least two of the following: time domain features, frequency domain features and / or cepstral features. Acquire a loudspeaker reference signal, perform acoustic echo cancellation processing on the sound signal based on the loudspeaker reference signal to obtain an echo cancellation output signal, and when the human voice presence information indicates the presence of human voice, adjust the adaptive update strategy of the acoustic echo cancellation processing according to the human voice confidence level. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter. Identify fluctuating noise in the echo-cancelled output signal that meets preset judgment conditions, preset intensity and / or dominant frequency changes over time, and determine the main noise frequency band of the echo-cancelled output signal based on the identification result of the fluctuating noise; Noise suppression processing is performed on the main noise frequency band, and the noise suppression processing includes at least narrowband suppression and / or adaptive gain suppression to obtain a noise-suppressed output signal; Obtain the status information of the acoustic echo cancellation process, wherein the status information includes at least a residual echo characterization quantity; Based on the human voice presence information, the state information, and the identification result of the fluctuating noise, the update parameters of the adaptive update strategy used for subsequent processing and / or the suppression intensity of the noise suppression processing are adjusted, and the amplified speech signal after speech enhancement is output based on the noise suppression output signal.
2. The speech enhancement method for a loudspeaker according to claim 1, characterized in that, After outputting the voice confidence score based on the multi-dimensional voice features and a preset classification rule, the method further includes: Acquire video streams from pre-deployed cameras in the meeting area; Visual analysis is performed on the video stream to generate visual voice confidence, wherein the visual analysis includes at least detecting facial pose and detecting lip movements; The voice confidence score is fused with the visual voice confidence score to obtain a fused voice confidence score. When the voice confidence score is within a preset fuzzy range, the fused voice confidence score is increased based on the visual voice confidence score. When the fused human voice confidence score meets the preset confidence score threshold, the fused human voice confidence score is used to replace the human voice confidence score to adjust the adaptive update strategy, so that the adaptive update strategy at least performs frozen adaptive filter coefficient updates and / or reduces the update intensity of the adaptive filter.
3. The speech enhancement method for a loudspeaker according to claim 1, characterized in that, The microphone includes at least two microphone channels and performs noise suppression processing on the main noise frequency band. The noise suppression processing includes at least narrowband suppression and / or adaptive gain suppression to obtain a noise-suppressed output signal, including: Acquire structural vibration signals from vibration sensors pre-deployed at the lightweight structure of the conference room, and air pressure fluctuation signals from air pressure sensors pre-deployed in the conference room and / or air conditioning outlet area; The occurrence of ultra-low frequency excitation is identified based on the air pressure fluctuation signal, the occurrence of structural resonance is identified based on the structural vibration signal, and when the ultra-low frequency excitation and the structural resonance are simultaneously satisfied, the physical cause of the secondary noise is confirmed. Multi-channel spectral analysis is performed on the sound signal to obtain abnormal fluctuation characteristics, and the spatial correlation of the abnormal fluctuation characteristics between different microphone channels is calculated to determine spatial inconsistency anomaly. The spatial inconsistency anomaly is that the spatial consistency index between at least two microphone channels is lower than a preset threshold in at least one frequency band corresponding to the abnormal fluctuation characteristics. The spatial consistency index includes at least the inter-channel correlation coefficient and / or amplitude squared coherence. The results of confirming the existence of the physical cause of the secondary noise are fused with the spatial inconsistency anomaly to confirm the existence of the secondary noise, and when the existence of the secondary noise is confirmed, the noise frequency band corresponding to the secondary noise is determined. The noise frequency band is taken as at least a part of the main noise frequency band. The echo cancellation output signal is subjected to inverse filtering within the main noise frequency band. The center frequency and / or bandwidth of multiple narrowband filters are dynamically adjusted to suppress the spectral region corresponding to the secondary noise, so as to obtain a noise suppression output signal. The auditory comfort characterization quantity is calculated for the noise suppression output signal and compared with a preset comfort threshold to obtain the comfort evaluation result; When the comfort assessment result indicates that the preset comfort conditions are not met, adjust the parameters of the inverse filtering process and / or the suppression parameters of the narrowband filter.
4. The voice enhancement method for a loudspeaker according to claim 1, characterized in that, Based on the human voice presence information, the state information, and the identification result of the fluctuating noise, the update parameters of the adaptive update strategy used for subsequent processing and / or the suppression intensity of the noise suppression processing are adjusted, including: The system acquires multiple performance indicators corresponding to the echo cancellation output signal and / or the noise suppression output signal. The performance indicators include speech intelligibility, echo suppression level, noise suppression level, and speech naturalness. The performance indicators are calculated based on the echo cancellation output signal and / or the noise suppression output signal according to preset calculation rules. Based on the human voice presence information, the residual echo characterization quantity in the state information, and the identification result of the fluctuation noise, the real-time environmental state information is determined; Based on the performance indicators and the environmental state information, the weights and / or priorities of the speech intelligibility indicator, the echo suppression level, the noise suppression level, and the speech naturalness indicator are dynamically adjusted to obtain the adjustment parameters; Multi-objective optimization decision-making is performed based on the adjustment parameters to determine the processing flow of the sound signal and the combination of related parameters. The related parameters include at least the update parameters of the adaptive update strategy and / or the suppression intensity of the noise suppression processing. The processing flow includes at least the synergistic method of the acoustic echo cancellation processing and the noise suppression processing.
5. The voice enhancement method for a loudspeaker according to claim 1, characterized in that, When the presence of human voice information indicates the presence of human voice, the adaptive update strategy for acoustic echo cancellation processing is adjusted according to the confidence level of the human voice. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter, including: A preset test signal is played periodically through a speaker, and the reflected sound of the preset test signal in the acoustic environment is picked up through the microphone of the amplifier to obtain the reflected sound signal. The reflected acoustic signal is analyzed to evaluate the acoustic reflection path, the path evaluation result is obtained, and the evolution trend of the acoustic reflection path is predicted based on the path evaluation result over a continuous preset time period. Based on the evolution trend and the human voice presence information, the adaptive update strategy of the acoustic echo cancellation processing is adjusted. The adjustment includes at least the following: when the predicted change of the acoustic reflection path exceeds a preset change threshold and the human voice presence information indicates that there is no human voice, the update intensity of the adaptive update strategy is increased to accelerate convergence; when the human voice presence information indicates that there is human voice, the update of the adaptive filter coefficients is frozen and / or the update intensity of the adaptive filter is reduced. When adjusting the adaptive update strategy, a protective gain is applied to the echo cancellation output signal in the human voice dominant frequency region determined by the multi-dimensional speech features. Based on the residual echo characterization in the state information, the residual echo level is evaluated, and when the residual echo level is higher than a preset echo threshold, the acoustic echo cancellation process is re-adaptively controlled based on the reflected sound signal. The re-adaptive control includes at least re-initializing the adaptive filter coefficients and / or further increasing the update strength of the adaptive update strategy when the human voice presence information indicates that there is no human voice.
6. The voice enhancement method for a loudspeaker according to claim 5, characterized in that, When adjusting the adaptive update strategy of the acoustic echo cancellation process based on the evolution trend and the human voice presence information, the adjustment further includes: When the evolution trend indicates that the change in the acoustic reflection path exceeds a preset change threshold, and the human voice presence information indicates the presence of low-energy human voices, the update parameters of the adaptive update strategy used for the acoustic echo cancellation processing, the gain parameters of the protective gain, and the determination parameters of the preset classification rule are adjusted collaboratively. The update parameters include at least the step size factor of the adaptive filter. The low-energy human voice is a human voice whose human voice energy representation based on the short-time energy in the multi-dimensional speech features is lower than a preset human voice threshold. The human voice energy representation is the short-time energy and / or its logarithmic form calculated based on the multi-dimensional speech features. The coordinated adjustment includes at least: during the period when the human voice presence information indicates the presence of low-energy human voices, reducing the step size factor and / or increasing the protective gain of the dominant frequency region of the human voices, and adjusting the confidence threshold and / or fuzzy interval boundary of the preset classification rule to improve the sensitivity of low-energy human voices.
7. The voice enhancement method for a loudspeaker according to claim 6, characterized in that, The gain parameters of the protective gain are adjusted in a coordinated manner, including: Based on the multi-dimensional speech features, a human voice spectral feature representation quantity is obtained, and environmental noise information represented by the recognition result of the fluctuating noise is acquired. The environmental noise information includes at least noise type, intensity, spectral distribution and transient fluctuation characteristics. The human voice spectral feature representation quantity is a spectral feature vector representing the distribution of human voice energy with frequency. When the human voice presence information indicates the presence of the low-energy human voice, and the environmental noise information indicates the presence of transient noise that overlaps with the spectral characteristics of the human voice and has the transient fluctuation characteristics, the gain parameter of the protective gain and / or the suppression parameter of the noise suppression processing are adjusted in the first frequency band. The first frequency band adjustment includes at least: increasing the protective gain and / or decreasing the suppression intensity of the noise suppression process in the human voice-dominated frequency region; and increasing the suppression intensity of the noise suppression process according to the transient fluctuation characteristics in the non-human voice-dominated frequency region to suppress the transient noise.
8. The voice enhancement method for a loudspeaker according to claim 7, characterized in that, The preset classification rules also include threshold parameters for distinguishing between echoes and human voices, and the method further includes: Based on the loudspeaker reference signal and the echo cancellation output signal, calculate the echo similarity representation quantity and / or the echo energy ratio representation quantity. The echo similarity representation quantity is a scalar or frequency band vector representing the correlation or coherence between the loudspeaker reference signal and the echo cancellation output signal. The echo energy ratio representation quantity is a scalar or frequency band vector representing the proportion or ratio of the echo component energy related to the loudspeaker reference signal in the echo cancellation output signal. When the voice presence information indicates the presence of a voice, and the voice energy representation is lower than a preset voice threshold, and the echo similarity representation and / or the echo energy ratio representation meet the preset echo judgment condition based on the threshold parameter, the threshold parameter used to distinguish between echo and voice is adjusted in the second frequency band. The threshold parameter includes the signal energy threshold used to judge speech activity and / or the similarity threshold corresponding to the echo similarity representation. The second frequency band adjustment includes at least: in the human voice-dominated frequency region, lowering the signal energy threshold and / or raising the similarity threshold; in the non-human voice-dominated frequency region, raising the signal energy threshold and / or lowering the similarity threshold; Based on the results of the second frequency band adjustment, the adaptive update strategy of the acoustic echo cancellation process and / or the gain parameter of the protective gain are adjusted accordingly.
9. The voice enhancement method for a loudspeaker according to claim 8, characterized in that, The microphone includes at least two microphone channels, and the method further includes: Spatial analysis is performed on the sound signals from at least two microphone channels to obtain directional information from at least two sound sources; For each of the sound source directions, spatial filtering is performed to obtain a direction separation signal for the corresponding direction, wherein the direction separation signal is a signal obtained by weighting and synthesizing the sound signals of at least two microphone channels according to the spatial filtering weights corresponding to the sound source directions; For each direction-separated signal, extract the human voice energy representation and the human voice spectrum feature representation obtained based on the multi-dimensional speech features, and output the corresponding direction's human voice confidence based on the preset classification rules; Based on the confidence level of human voice in each direction and the human voice energy representation in the corresponding direction, the signal energy threshold in the corresponding direction is adjusted respectively, and the adjusted signal energy threshold is used to distinguish between echo and human voice in the corresponding direction and / or to adjust the adaptive update strategy in the corresponding direction. When the confidence level of human voices in at least two directions meets the preset confidence threshold, protective gains are applied to the dominant frequency regions of human voices in the corresponding directions.
10. A voice enhancement system for a loudspeaker, characterized in that, include: A human voice presence information generation module is used to acquire the sound signal picked up by the microphone of the loudspeaker, perform frame-by-frame processing on the sound signal, extract multi-dimensional speech features for each frame of the sound signal, output human voice confidence based on the multi-dimensional speech features and according to the preset classification rules, and generate human voice presence information based on the human voice confidence. The multi-dimensional speech features include at least two of the following: time domain features, frequency domain features and / or cepstral features. An echo cancellation strategy adjustment module is used to acquire a loudspeaker reference signal, perform acoustic echo cancellation processing on the sound signal based on the loudspeaker reference signal to obtain an echo cancellation output signal, and when the human voice presence information indicates the presence of human voice, adjust the adaptive update strategy of the acoustic echo cancellation processing according to the human voice confidence level. The adaptive update strategy includes freezing the adaptive filter coefficient update and / or reducing the update intensity of the adaptive filter. The noise identification module is used to identify fluctuating noise in the echo cancellation output signal that meets preset judgment conditions, preset intensity and / or dominant frequency changes over time, and to determine the main noise frequency band of the echo cancellation output signal based on the identification result of the fluctuating noise. A noise suppression processing module is used to perform noise suppression processing on the main noise frequency band. The noise suppression processing includes at least narrowband suppression and / or adaptive gain suppression to obtain a noise-suppressed output signal. The status information acquisition module is used to acquire the status information of the acoustic echo cancellation process, wherein the status information includes at least the residual echo characterization quantity. The processing flow coordination module is used to adjust the update parameters of the adaptive update strategy for subsequent processing and / or adjust the suppression intensity of the noise suppression processing based on the human voice presence information, the state information and the recognition result of the fluctuating noise, and output the amplified speech signal after speech enhancement based on the noise suppression output signal.