Adaptive noise cancellation method and system for a communications headset based on ambient sound detection

By using multi-dimensional environmental noise analysis and online reinforcement learning to optimize state decision mapping parameters, the problem of noise suppression and voice fidelity in call center headsets under complex acoustic scenarios was solved, achieving improved adaptive noise reduction and voice quality.

CN122269186APending Publication Date: 2026-06-23BEIJING ZHONGLIAN NORTH INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGLIAN NORTH INFORMATION TECH CO LTD
Filing Date
2026-05-11
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing noise reduction technologies in call center headsets are ill-suited to adapting to dynamically changing environmental noise, resulting in voice distortion and processing delays. They fail to achieve the optimal balance between noise suppression and voice fidelity in complex acoustic scenarios.

Method used

The system acquires the raw mixed signal through the built-in microphone of the headset, performs multi-dimensional environmental noise analysis, generates critical control commands using a state decision mapper, and combines online estimation and gain stabilization control to achieve perceptual domain spectrum enhancement processing and online reinforcement learning, thereby optimizing the state decision mapping parameters.

Benefits of technology

It improves the adaptive noise reduction capability of the headset in complex environments, enhances voice quality and user auditory comfort, and resolves the conflict between environmental noise suppression and call voice fidelity in real-time signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122269186A_ABST
    Figure CN122269186A_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive noise reduction method and system for a call center headset based on ambient sound detection, belonging to the field of communication acoustic signal processing technology. Key technical points include: acquiring the original mixed signal using the built-in microphone of the call center headset; obtaining critical control commands through multi-dimensional ambient noise analysis and based on the state decision mapping parameters of a state decision mapper; obtaining an intermediate audio signal based on the original mixed signal through online estimation of the critical stability point and gain stabilization control; obtaining an enhanced speech signal through perceptual domain spectral enhancement processing; and obtaining updated state decision mapping parameters based on historical critical control commands through online reinforcement learning. This invention generates commands through a mapper based on noise analysis and executes stabilization control and spectral enhancement, and optimizes parameters using reinforcement learning, achieving active stabilization control of the acoustic-electric feedback closed loop. This resolves the conflict between noise suppression and speech fidelity, improving adaptive noise reduction capability and speech quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication acoustic signal processing technology, and more specifically to an adaptive noise reduction method and system for call center headsets based on ambient sound detection. Background Technology

[0002] Call center headsets have become indispensable devices in various office, customer service, and professional communication scenarios. Users' demands for call clarity are increasing, making noise reduction technology a core function of call center headsets. Traditional noise reduction methods are mainly divided into passive and active noise reduction. Passive noise reduction relies on physical sound insulation structures and has limited effectiveness in suppressing low- and mid-frequency environmental noise. While active noise reduction can cancel some noise through anti-phase sound waves, it often uses fixed thresholds and single algorithms, making it difficult to adapt to dynamically changing environmental noise. Furthermore, it easily introduces voice distortion and processing delays during processing, affecting the naturalness and fluency of real-time calls. In recent years, although digital signal processing technology and machine learning have made progress in the field of acoustics, existing call center headset noise reduction solutions still lack multi-dimensional and accurate perception and intelligent decision-making capabilities for environmental noise. They cannot achieve the optimal balance between noise suppression and voice fidelity in complex acoustic scenarios, nor can they self-optimize based on user preferences and the usage environment. Therefore, existing technologies have shortcomings. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to provide an adaptive noise reduction method and system for call center headsets based on ambient sound detection. This method acquires the original mixed signal through the built-in microphone of the call center headset and performs multi-dimensional ambient noise analysis. A critical control command is generated based on a state decision mapper. Based on the critical control command and the original mixed signal, online estimation of the critical stability point and gain stabilization control are performed to obtain an intermediate audio signal. Then, an enhanced speech signal is obtained through perceptual domain spectral enhancement processing. Finally, the state decision mapper parameters are updated through online reinforcement learning using the enhanced speech signal and historical critical control commands. This method achieves active utilization and stable control of the acoustic-electric feedback closed-loop system, resolving the fundamental conflict between ambient noise suppression and call voice fidelity in real-time signal processing in call center scenarios. It improves the adaptability and naturalness of the noise reduction effect, enhancing the voice quality and user auditory comfort in call center communication.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] This invention provides an adaptive noise reduction method for call center headsets based on ambient sound detection, comprising:

[0006] The system acquires the raw mixed signal using the microphone built into the headset, performs multi-dimensional environmental noise analysis, and makes decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands.

[0007] Based on the critical control command and the original mixed signal, the intermediate audio signal is obtained through online estimation of the critical stable point and gain stabilization control.

[0008] Based on the intermediate audio signal, an enhanced speech signal is obtained through perceptual domain spectral enhancement processing;

[0009] Based on the enhanced voice signal and historical critical control commands, updated state decision mapping parameters are obtained through online reinforcement learning.

[0010] As a further improvement of the present invention, the method of acquiring the original mixed signal based on the microphone built into the call center headset, performing multi-dimensional environmental noise analysis, and making a decision based on the state decision mapping parameters of the state decision mapper to obtain critical control commands includes:

[0011] Based on the original mixed signal, the instantaneous impact intensity is obtained through short-time energy analysis and differential calculation;

[0012] Based on the original mixed signal and the preset characteristic frequencies of the headphone and ear canal system, a list of threat frequencies is obtained through power spectrum estimation.

[0013] Based on the original mixed signal, the equivalent spatial diffusion is obtained by calculating the time-varying statistical characteristics of the signal.

[0014] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, a critical control command is obtained by multi-feature fusion and decision-making through a state decision mapper. The critical control command includes target stability margin, active notch center frequency, and loop gain adjustment strategy.

[0015] As a further improvement of the present invention, the critical control command is obtained by multi-feature fusion and decision-making through a state decision mapper based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion. The critical control command includes target stability margin, active notch filter center frequency, and loop gain adjustment strategy, including:

[0016] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the target stability margin is obtained through the stability margin mapping rule in the state decision mapper.

[0017] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the active notch center frequency is obtained through the threat frequency selection rule in the state decision mapper.

[0018] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, a loop gain adjustment strategy is obtained through the gain strategy mapping rule in the state decision mapper.

[0019] Based on the target stability margin, active notch filter center frequency, and loop gain adjustment strategy, critical control commands are obtained through integration.

[0020] As a further improvement of the present invention, the step of obtaining an intermediate audio signal based on the critical control command combined with the original mixed signal through online estimation of the critical stability point and gain stabilization control includes:

[0021] Based on the original mixed signal, by injecting pseudo-random probe signals into the electroacoustic loop and performing closed-loop response cross-correlation analysis, the critical frequency and critical gain estimates are obtained.

[0022] Based on the critical gain estimate and critical control command, the safe operating gain is calculated through gain constraints.

[0023] Based on the critical frequency, safe operating gain, and critical control command, an intermediate audio signal is obtained through adaptive notch filtering and loop phase fine-tuning.

[0024] As a further improvement of the present invention, the step of obtaining estimates of the critical frequency and critical gain based on the original mixed signal by injecting a pseudo-random probe signal into the electroacoustic loop and performing closed-loop response cross-correlation analysis includes:

[0025] Based on the original mixed signal, the microphone received signal is obtained by injecting a pseudo-random self-noise probe signal;

[0026] Based on the self-noise probe signal and the microphone received signal, a cross-correlation function sequence is obtained by calculating the normalized cross-correlation function;

[0027] Based on the cross-correlation function sequence, the position and amplitude of the cross-correlation function peak are obtained through a peak detection algorithm;

[0028] Based on the position and amplitude of the peak value of the cross-correlation function, the critical frequency and the corresponding critical gain estimate are obtained through growth trend analysis.

[0029] As a further improvement of the present invention, the step of calculating the safe operating gain based on the critical gain estimate and the critical control command through gain constraints includes:

[0030] Based on the target stability margin in the critical control command, the stability margin compensation coefficient is calculated through gain constraints.

[0031] The safe operating gain is calculated based on the critical gain estimate and the stability margin compensation coefficient.

[0032] As a further improvement of the present invention, the step of obtaining an enhanced speech signal based on the intermediate audio signal through perceptual domain spectral enhancement processing includes:

[0033] Based on the intermediate audio signal, the power spectrum characteristics of the residual noise are obtained through speech activity detection and background noise spectrum estimation in the perceptual domain spectrum enhancement processing.

[0034] Based on the power spectrum characteristics of the residual noise, an enhanced speech signal is obtained through a non-uniform frequency response enhancement algorithm based on an auditory masking model in the perceptual domain spectral enhancement processing.

[0035] As a further improvement of the present invention, the step of obtaining updated state decision mapping parameters through online reinforcement learning based on the enhanced speech signal and historical critical control commands includes:

[0036] Based on the enhanced speech signal, a speech quality index is obtained through a perceptual speech quality estimation algorithm;

[0037] Stability indices are obtained through loop stability analysis based on the intermediate audio signal.

[0038] A comprehensive evaluation scalar is obtained by weighted fusion of the aforementioned speech quality and stability indicators.

[0039] Based on historical critical control commands, combined with the aforementioned speech quality index, stability index, and comprehensive evaluation scalar, the state decision mapper is optimized through online reinforcement learning based on stochastic gradient descent to obtain updated state decision mapping parameters.

[0040] As a further improvement of the present invention, the state decision mapper is optimized by combining historical critical control instructions with the speech quality index, stability index, and comprehensive evaluation scalar, and the updated state decision mapping parameters are obtained through online reinforcement learning based on stochastic gradient descent, including:

[0041] Based on historical critical control instructions, an environment state vector and action vector are constructed through online reinforcement learning based on stochastic gradient descent, and a reward scalar is constructed based on the comprehensive evaluation scalar.

[0042] Based on the environmental state vector, action vector, and reward scalar, the state decision mapper is optimized using a stochastic gradient descent strategy to obtain updated state decision mapping parameters.

[0043] This invention provides an adaptive noise reduction system for call center headsets based on ambient sound detection, the system comprising:

[0044] Voice decision system: Based on the microphone built into the call center headset, the system acquires the raw mixed signal, performs multi-dimensional environmental noise analysis, and makes decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands;

[0045] Voice-controlled stabilization system: Based on the critical control command and the original mixed signal, an intermediate audio signal is obtained through online estimation of the critical stabilization point and gain stabilization control;

[0046] Speech enhancement system: Based on the intermediate audio signal, an enhanced speech signal is obtained through perceptual domain spectral enhancement processing;

[0047] Policy learning system: Based on the enhanced speech signal and historical critical control commands, updated state decision mapping parameters are obtained through online reinforcement learning.

[0048] This invention performs multi-dimensional environmental noise analysis on the raw mixed signal collected by the built-in microphone of the call center headset. It generates critical control commands through a state decision mapper, and performs online estimation of the critical stability point and gain stabilization control on the raw mixed signal according to the critical control commands to obtain an intermediate audio signal. Then, it obtains an enhanced speech signal through perceptual domain spectral enhancement processing. Based on the enhanced speech signal and historical critical control commands, it continuously optimizes the state decision map parameters through online reinforcement learning, realizing the active utilization and stable control of the critical point of the acoustic-electric feedback closed-loop system. This invention solves the conflict between environmental noise suppression and call voice fidelity in real-time signal processing in call center scenarios, significantly improves the adaptability and speech naturalness of the noise reduction system in complex and changing environments, and effectively improves the voice quality and user auditory comfort of call center communication. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating the steps of the adaptive noise reduction method for call center headsets based on ambient sound detection according to the present invention.

[0050] Figure 2 A flowchart illustrating the steps involved in generating multi-dimensional environmental noise analysis and critical control commands.

[0051] Figure 3 A flowchart illustrating the steps of online estimation of critical stability point and gain stabilization control.

[0052] Figure 4 This is a schematic diagram of the adaptive noise reduction system for call center headsets based on ambient sound detection, as described in this invention. Detailed Implementation

[0053] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof.

[0054] The term "and / or" in the following text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0055] like Figure 1 As shown, the present invention provides an adaptive noise reduction method for call center headsets based on ambient sound detection, comprising:

[0056] The system acquires the raw mixed signal using the microphone built into the headset, performs multi-dimensional environmental noise analysis, and makes decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands.

[0057] Based on critical control commands combined with the original mixed signal, the intermediate audio signal is obtained through online estimation of the critical stable point and gain stabilization control.

[0058] An enhanced speech signal is obtained by perceptual domain spectral enhancement processing based on the intermediate audio signal.

[0059] Based on enhanced voice signals and historical critical control commands, updated state decision mapping parameters are obtained through online reinforcement learning.

[0060] The headset is a terminal device specifically designed for voice communication scenarios, featuring audio playback and ambient sound acquisition capabilities. It has at least one built-in microphone for acquiring ambient sound and near-field speech. This microphone is used to acquire real-time audio signals containing user speech, ambient noise, and potential acoustic feedback. The original mixed signal is a time-domain audio signal acquired in real-time by the built-in microphone, containing target speech signals, ambient noise signals, and device circuit noise. Multi-dimensional ambient noise analysis involves a signal processing procedure that transforms the original mixed signal into a time-frequency domain, extracts statistical features, and classifies noise types. This process extracts three-dimensional noise features oriented towards closed-loop stability, including instantaneous impact intensity, a threat frequency list, and equivalent noise. Spatial diffusion is used to quantify the impact of environmental noise on the stability of the acoustic-electric feedback closed-loop system. The state decision mapper is a lightweight, pre-trained nonlinear mapping model whose internal parameters can be optimized through online learning. It is used to fuse three-dimensional feature vectors and map them into critical control commands. The state decision mapping parameters are the core parameter set characterizing the mapping relationship of the state decision mapper, including noise feature weight coefficients, state transition thresholds, and initial values ​​of the base gain. The critical control commands are the set of commands calculated and output by the state decision mapper based on the current multi-dimensional environmental noise analysis results and the state decision mapping parameters. They are used to specify the operating mode and core control parameters of the noise reduction system, including the target stability margin, active notch filter center frequency, and loop gain. The adjustment strategy involves online estimation of the critical stability point, which is achieved by injecting pseudo-random probe signals into the electroacoustic loop and performing closed-loop response cross-correlation analysis to estimate the critical frequency and critical gain of the current system in real time. This is used to obtain the boundary conditions under which the system is about to become unstable. Gain stabilization control calculates the safe operating gain based on the critical gain estimate and the target stability margin, and combines adaptive notch filtering and loop phase fine-tuning to stabilize the system near the target stability margin, thereby suppressing environmental noise while ensuring system stability. The intermediate audio signal is the audio signal obtained after the original mixed signal has been processed by gain stabilization control, in which environmental noise has been significantly suppressed but some non-stationary noise remains. The perceptual domain spectral enhancement... The enhanced speech signal is a signal processing procedure that combines the characteristics of human auditory perception to perform spectral amplitude correction, residual noise suppression, and speech detail compensation on the intermediate audio signal. It is used to further suppress residual noise and highlight the target speech components without introducing auditory distortion. The enhanced speech signal is a high-definition target speech signal that meets the needs of voice communication, obtained by enhancing the intermediate audio signal through perceptual domain spectral enhancement. The online reinforcement learning is a machine learning process that uses the quality evaluation result of the enhanced speech signal as a reward signal and combines it with the environmental noise state corresponding to the historical critical control command to perform real-time iterative optimization of the state decision mapping parameters. It is used to enable the noise reduction system to adapt to changes in different communication environments and continuously optimize decision accuracy.The updated state decision mapping parameters are the optimized parameter set obtained by iteratively updating the initial state decision mapping parameters through an online reinforcement learning process. These parameters include the updated noise feature weight coefficients, the adaptively adjusted state transition threshold, and the optimized base gain value.

[0061] This embodiment generates critical control commands through multi-dimensional environmental noise analysis and state decision mapping, obtains intermediate audio signals through online estimation of critical stability points and gain stabilization control, obtains enhanced speech signals through perceptual domain spectrum enhancement processing, and optimizes state decision mapping parameters through online reinforcement learning. This achieves active utilization and stable control of the critical point of the acoustic-electric feedback closed-loop system, thereby solving the trade-off between noise suppression and speech fidelity in traditional noise reduction methods. It enables adaptive noise reduction of the headset in complex and changing environments, significantly improving the adaptability of noise reduction effect and speech naturalness, enhancing the voice quality and user auditory comfort of call communication, while reducing processing latency and enabling the system to have continuous self-optimization capabilities.

[0062] Furthermore, this embodiment provides a step-by-step approach to obtain critical control commands by acquiring the raw mixed signal using the microphone built into the headset, performing multi-dimensional environmental noise analysis, and making decisions based on the state decision mapping parameters of the state decision mapper. The steps include:

[0063] Based on the original mixed signal, the instantaneous impact intensity is obtained through short-time energy analysis and differential calculation;

[0064] Based on the original mixed signal combined with the preset characteristic frequencies of the headphone and ear canal system, a list of threat frequencies is obtained through high-resolution power spectrum estimation.

[0065] Based on the original mixed signal, the equivalent spatial diffusion is obtained by calculating the time-varying statistical characteristics of the signal.

[0066] Based on instantaneous impact intensity, threat frequency list and equivalent spatial diffusion, critical control commands are obtained by multi-feature fusion and decision-making through state decision mapper. The critical control commands include target stability margin, active notch center frequency and loop gain adjustment strategy.

[0067] Among them, short-time energy analysis is a signal analysis method that calculates the energy value frame by frame after the original mixed signal is segmented, and is used to quantify the energy magnitude of a single frame signal to reflect the temporal energy distribution characteristics of the signal; differential calculation is a numerical calculation method that solves the first-order and second-order differences of the short-time energy values ​​of consecutive frames, and is used to capture the temporal variation trend and abrupt change characteristics of signal energy; instantaneous impact intensity is an index obtained based on the significant peak quantization of the second-order difference, and is used to characterize the severity of abrupt changes in environmental noise energy in the temporal domain; the preset characteristic frequency of the headphone and ear canal system is obtained by pre-shipment measurement or adaptive identification, and reflects the open-loop frequency response gain peak frequency band of the physical path from the headphone speaker to the microphone. The frequency set is used to determine the dangerous frequencies in the environment noise that are likely to cause system howling. It is obtained by combining offline acoustic calibration of the headset and ear canal system with open-loop frequency response testing. High-resolution power spectrum estimation is a method for estimating the power spectral density of signal frames based on an improved autoregressive model. It is used to perform high-precision frequency domain analysis on the original mixed signal to obtain the spectral distribution characteristics of the signal. The threat frequency list is a set of frequencies that are likely to cause acoustic-electric closed-loop system howling after comparing the peak frequency band of the power spectrum with the preset characteristic frequencies of the headset and ear canal system. The preset characteristic frequencies of the headset and ear canal system are determined based on the physical acoustic path characteristics from the headset speaker to the microphone, through the calibration of the headset and ear canal system. The values ​​were obtained by offline acoustic calibration of the earphone and ear canal system combined with open-loop frequency response testing. In this embodiment, the values ​​were 1.5kHz and 3.2kHz. The signal time-varying statistical characteristics are the normalized autocorrelation function of the original mixed signal in the speech silence segment and the decay rate characteristics of the function outside of zero hysteresis, used to characterize the sound field distribution characteristics of environmental noise. The equivalent spatial diffusion is a quantitative index with a value range of [0,1] calculated based on the signal time-varying statistical characteristics, used to characterize the degree to which the sound field corresponding to environmental noise is close to the diffusion field. The closer the value is to 1, the closer the sound field is to the diffusion field. The critical control command is the set of commands output by the state decision mapper based on the multi-dimensional environmental noise analysis results. The operating mode is used to guide subsequent gain stabilization control; the target stability margin is the core control parameter output by the state decision mapper, which specifies the margin between the operating point and the instability critical point of the acoustic-electric closed-loop system. The value range is (0,1). This parameter is dynamically adjusted according to the instantaneous impact intensity. The stronger the impact, the larger the margin to ensure stability; the active notch filter center frequency is the first threat frequency selected from the threat frequency list, which is used to indicate the center frequency point where the notch filter needs to be inserted in the loop; the loop gain adjustment strategy is the gain adjustment method formulated according to the equivalent spatial diffusion, which is divided into two types: global gain adjustment and frequency-selective gain adjustment, to adapt to the system stability control requirements under different sound field distributions.

[0068] Specifically, such as Figure 2 As shown, the original mixed signal is first pre-emphasized, framed, and windowed to convert it into short-time stationary signal frames. ,in For frame index, For the intra-frame sampling point index, the pre-emphasis coefficient is... A first-order FIR filter, frame division using frame length For 512 sampling points, frame shift For a sampling method with 256 sampling points, a modified Hamming window is used. This modifies the window function edges by smoothing them to reduce inter-frame signal distortion, building upon the traditional Hamming window. For each signal frame... Perform short-time energy analysis using the formula The short-time energy values ​​of each frame were calculated. Then, differential calculation is performed on the short-time energy values ​​of consecutive frames, using the formula... , First-order differences are obtained respectively and second-order difference The instantaneous impact intensity was calculated based on the absolute peak value of the second-order difference combined with the energy ratio. The calculation formula is ,in , The weighting coefficients and The weighting coefficients are set based on the energy distribution characteristics of continuous background noise and sudden impact noise in a noisy environment. This formula is derived from the traditional impact intensity calculation formula combined with the proportion of energy change, thus improving the ability to identify weak impact noise. Then, the signal frame... An improved high-resolution power spectrum estimation is performed using an autoregressive model based on the variable-order Burg algorithm. The model order is adaptively adjusted according to the spectral complexity of the signal frame, ranging from order 8 to 24. The power spectrum of the signal frame is obtained through this model. ,in Angular frequency, representing the frequency components of the signal in the frequency domain, is used to analyze the power spectrum. The spectral peak frequency bands are compared with preset characteristic frequencies of the headphone and ear canal system, and a preset frequency deviation threshold is set. The deviation will be less than the preset frequency deviation threshold. The frequencies are labeled as threat frequencies, and all threat frequencies are combined to obtain a threat frequency list. ,in This represents the total number of threat frequencies identified in this study. For the first The threat frequency was then identified; subsequently, the speech silence segments were determined from the original mixed signal using the basic energy threshold method, and pure noise frames within the speech silence segments were extracted. The normalized autocorrelation function of the pure noise frames was then calculated. ,in The lag order represents the time delay between two signal sampling points; the normalized autocorrelation function is calculated as follows: In the formula The frame length of a pure noise frame. The first frame of pure noise 1 sampling point; and solve for the decay rate of the function outside of zero hysteresis. Through formula The equivalent spatial diffusivity was calculated. ,in The preset maximum attenuation rate is set to the attenuation rate of the white noise autocorrelation function, which is derived from the autocorrelation attenuation characteristics of the diffuse and directional sound fields. Finally, a three-dimensional feature vector is constructed. The input is fed into a state decision mapper for multi-feature fusion and decision-making. The state decision mapper is an improved 3-layer fully connected neural network, which adds a noise feature attention mechanism to the hidden layers of the traditional fully connected neural network. The model collects 3D feature vectors of environmental noise from multiple scenes and the corresponding optimal control commands as training samples, and uses a stochastic gradient descent algorithm with L2 regularization for training. The regularization coefficient is set to 0.001, the learning rate is set to 0.01, and the number of training iterations is set to 10,000. The model input is a 3D feature vector. The first hidden layer has 64 neurons, and the activation function is a modified ReLU function. A slope of 0.02 is added in the negative interval to improve the gradient propagation ability. The second hidden layer has 32 neurons, and the activation function is the Sigmoid function. The output layer has 3 neurons, which correspond to the decision results of the target stability margin, the active notch filter center frequency, and the loop gain adjustment strategy, respectively. The decision results output by this model are integrated into the critical control command.

[0069] For example, this embodiment assumes a typical office call scenario: the user wears a headset for daily work, and the ambient noise mainly includes continuous low-frequency noise from the air conditioning system, occasional keyboard typing from colleagues, and conversations from distant colleagues. First, the original mixed signal is acquired in real time using the headset's built-in microphone. Assuming the sampling rate of the original mixed signal is 16kHz, it undergoes pre-emphasis, framing, and windowing processing, with a frame length of... With 512 sampling points, frame shift With 256 sampling points and a smoothing coefficient of 0.05 for the improved Hamming window, 20 consecutive signal frames were analyzed, and the short-time energy values ​​of each frame were obtained through short-time energy analysis. ,in The first difference is obtained through difference calculation. and second-order difference Set the weighting coefficient as , Substituting into the instantaneous impact strength calculation formula, we get High-resolution power spectrum estimation was performed on the signal frame using an autoregressive model based on the variable-order Burg algorithm. The model order was adaptively adjusted to 16 to obtain the power spectrum. The preset characteristic frequencies for the headphones and ear canal system are 1.5kHz and 3.2kHz, respectively, and the preset frequency deviation threshold is set to... The resulting threat frequency list is as follows: The speech silence segment was determined by the basic energy threshold method, and pure noise frames were extracted. The decay rate of the normalized autocorrelation function was then calculated. Preset maximum decay rate The rate is 1000 / s, and the decay rate of the autocorrelation function of the pure noise frame is 150 / s. Substituting these values ​​into the formula for calculating the equivalent spatial spread, we get... Constructing three-dimensional feature vectors The input is fed into the state decision mapper of an improved 3-layer fully connected neural network, and the network calculates the target stability margin. Active notch center frequency The loop gain adjustment strategy is frequency-selective gain adjustment, and the final critical control command is obtained after integration. In this embodiment, the values ​​of sampling rate, frame length, frame shift, weighting coefficient, model order, and frequency deviation threshold are merely examples. Those skilled in the art can set them according to actual conditions, and this embodiment does not impose any limitations on them.

[0070] This embodiment improves the analysis accuracy of the original mixed signal through improved front-end processing, achieves precise quantification of impact noise through short-time energy analysis and differential calculation, accurately identifies the whistling threat frequency through improved high-resolution power spectrum estimation combined with system characteristic frequencies, effectively characterizes the sound field distribution characteristics through time-varying statistical feature calculation, and achieves intelligent fusion of multiple features and precise generation of control commands through a state decision mapper. This realizes multi-dimensional accurate perception of complex environmental noise and intelligent decision-making of noise reduction control commands, improving the comprehensiveness of environmental noise analysis and the accuracy and adaptability of control command decisions, laying a precise decision-making foundation for subsequent gain stabilization control.

[0071] Furthermore, this embodiment provides a method for obtaining critical control commands based on instantaneous impact intensity, a threat frequency list, and equivalent spatial diffusion through a state decision mapper via multi-feature fusion and decision-making. The critical control commands include target stability margin, active notch filter center frequency, and loop gain adjustment strategy. The steps include:

[0072] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the target stability margin is obtained through the stability margin mapping rule in the state decision mapper.

[0073] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the active notch wave center frequency is obtained through the threat frequency selection rule in the state decision mapper.

[0074] Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the loop gain adjustment strategy is obtained through the gain strategy mapping rule in the state decision mapper.

[0075] Based on the target stability margin, active notch filter center frequency, and loop gain adjustment strategy, critical control commands are obtained through integration.

[0076] Among them, the stability margin mapping rule is a target stability margin assignment rule formulated in the state decision mapper based primarily on instantaneous impact intensity and secondarily on equivalent spatial diffusion. It is used to dynamically match the stability margin of the acoustic-electric closed-loop system according to the impact characteristics and sound field distribution characteristics of environmental noise. The target stability margin is a core control parameter output by the stability margin mapping rule, used to specify the margin between the operating point and the instability critical point of the acoustic-electric closed-loop system, with a value range of (0,1). The threat frequency selection rule is a threat frequency screening rule formulated in the state decision mapper based on the spectral energy proportion of the threat frequency and the degree of deviation from the characteristic frequencies of the headphone and ear canal system. It is used to select from the threat frequency list. The core frequency most prone to howling is designated as the active notch center frequency; the active notch center frequency is the core frequency selected from the threat frequency list based on the threat frequency selection rule, which requires focused notch suppression; the gain strategy mapping rule is the loop gain adjustment method selection rule formulated in the state decision mapper based on the value range of equivalent spatial diffusion, used to match the appropriate gain adjustment method according to the sound field distribution characteristics, and formulated based on the system stability test results under different sound fields; the loop gain adjustment strategy is the gain adjustment method obtained according to the gain strategy mapping rule, including global gain adjustment and frequency-selective gain adjustment, used to adapt to the system stability control requirements under different sound field distributions.

[0077] Specifically, firstly, the instantaneous impact intensity With preset multi-level impact thresholds The comparison shows that these two threshold levels correspond to three impact levels, and are also considered in conjunction with the equivalent spatial diffusion. The value of is used to execute the stability margin mapping rule in the state decision mapper. This rule divides the instantaneous impact intensity into three levels: the first impact, the second impact, and the third impact. The first impact level corresponds to the instantaneous impact intensity. Less than The second impact level corresponds to the instantaneous impact intensity. Greater than or equal to and less than The third impact level corresponds to the instantaneous impact intensity. Greater than or equal to And match a base stability margin for each level. Then, based on the equivalent spatial diffusion Make corrections, the correction formula is as follows ,in As a basic stability margin for each impact level, To achieve the target stability margin, The correction factor, typically 0.2, is derived from the influence of sound field distribution on system stability. For the first impact level, the basic stability margin is minimized to maximize noise reduction, while for the third impact level, a larger basic stability margin is used to ensure system stability. Then, a threat frequency list is used. Each threat frequency Calculate the risk value of howling The threat frequency selection rule in the execution state decision mapper is used to calculate the howling risk value. ,in Threat frequency Spectral energy, , These are the characteristic frequencies of the headphone and ear canal system. The maximum frequency deviation threshold. The weighting coefficients and The weighting coefficient The settings are typically based on the degree of influence of spectral energy factors and frequency deviation factors on howling risk. Value greater than To highlight the energy-dominated characteristics of howling, this formula is an improvement on the traditional howling risk assessment method by combining the proportion of spectral energy and the degree of frequency deviation. The howling risk value is selected through a threat frequency selection rule. The highest threat frequency is the center frequency of the active notch filter. If the threat frequency list is empty, the active notch center frequency is set to empty; ultimately, the equivalent spatial diffusion is... Compared with the preset diffusion threshold A comparison is performed, and the gain policy mapping rule in the state decision mapper is executed, wherein the preset diffusion threshold is... The value was determined through multi-scene sound field characteristic tests to distinguish between diffuse and directional sound fields; in this embodiment, the value is set to 0.7. (When the equivalent spatial diffusion...) Greater than or equal to the preset diffusion threshold When the sound field is determined to be diffuse, a global gain adjustment strategy is selected, and the equivalent spatial diffusion is adjusted accordingly. Less than the preset diffusion threshold When the sound field is determined to be directional, a frequency-selective gain adjustment strategy is selected. The target stability margin obtained through the stability margin mapping rule, the active notch center frequency obtained through the threat frequency selection rule, and the loop gain adjustment strategy obtained through the gain strategy mapping rule are structurally integrated to form a critical control instruction set containing three core control parameters.

[0078] This embodiment, based on multi-dimensional features such as instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, achieves precise decision-making on the core parameters of critical control commands through customized stability margin mapping rules, threat frequency selection rules, and gain strategy mapping rules in the state decision mapper. Through structured integration, a complete critical control command is obtained, realizing dynamic and precise matching between noise reduction control parameters and complex environmental noise characteristics. This improves the pertinence and effectiveness of critical control commands and ensures the synergy between noise reduction control and stability control in the subsequent acoustic-electric closed-loop system.

[0079] Furthermore, this embodiment provides a step for obtaining an intermediate audio signal based on critical control commands combined with the original mixed signal, through online estimation of the critical stability point and gain stabilization control, including:

[0080] Based on the original mixed signal, by injecting pseudo-random probe signals into the electroacoustic loop and performing closed-loop response cross-correlation analysis, the estimated values ​​of critical frequency and critical gain are obtained.

[0081] Based on the critical gain estimate and critical control command, the safe operating gain is calculated through gain constraints.

[0082] Based on the critical frequency, safe operating gain, and critical control commands, the intermediate audio signal is obtained through adaptive notch filtering and loop phase fine-tuning techniques.

[0083] The electroacoustic loop is a closed acoustic-electric signal transmission circuit consisting of a digital signal processor for the headset, a speaker, an acoustic path in the ear canal, and a built-in microphone, connected in sequence. It is the physical basis for adaptive noise reduction. The pseudo-random probe signal is a pseudo-random self-noise sequence with an amplitude smaller than the preset environmental noise floor value and a flat spectrum. It is used to excite the electroacoustic loop to obtain loop response characteristics without affecting normal voice communication. The closed-loop response cross-correlation analysis is a signal analysis method that calculates the normalized cross-correlation function between the pseudo-random probe signal and the loop response signal collected by the microphone to analyze the loop response characteristics and extract critical parameters. It is used to accurately identify the critical state parameters of the electroacoustic loop that are about to become unstable. The critical frequency is the characteristic frequency at which the electroacoustic loop satisfies the condition that the loop gain is close to 1 and the phase satisfies the oscillation condition when it is close to instability. It is the frequency point at which the electroacoustic loop is most prone to howling. The critical gain estimate is the digital gain value required for the electroacoustic loop to reach the critical state of near instability at the critical frequency. It is a characterization of the loop's failure rate. The core parameters of the stable boundary are: gain constraint calculation, which is a numerical calculation process based on the critical gain estimate and the target stability margin in the critical control command, and solves the safe operating gain through a preset mathematical relationship, used to establish a balance constraint between loop stability and noise reduction effect; safe operating gain is the actual operating digital gain of the electroacoustic loop obtained by gain constraint calculation, ensuring that the electroacoustic loop always operates in a stable region with a target stability margin from the critical instability state; adaptive notch filtering is a digital filtering technology that dynamically adjusts the filtering parameters according to the active notch center frequency in the critical control command to deeply suppress specific threat frequencies in the electroacoustic loop, used to specifically eliminate the impact of howling threat frequencies on loop stability; loop phase fine-tuning technology is a signal processing technology that performs small-amplitude phase adjustment of the safe operating gain in the forward path of the electroacoustic loop near the active notch center frequency and the critical frequency, used to optimize the suppression effect of the electroacoustic loop on threat frequencies while ensuring phase flatness of the voice path.

[0084] Specifically, such as Figure 3 As shown, speech activity detection is first performed on the original mixed signal to identify speech intervals. Then, pseudo-random probe signals are injected only into the speaker output of the electroacoustic loop during these intervals. ,in For sampling point indexing, the amplitude of the pseudo-random probe signal is set to be less than 20dB below the preset ambient noise floor value, and the spectrum remains flat within the 20Hz-8kHz audio frequency band. Simultaneously, the microphone reception signal, including the probe signal loop response, is acquired via the built-in microphone. Then, the pseudo-random probe signal With microphone to receive signals Perform closed-loop response cross-correlation analysis using the formula The normalized cross-correlation function is calculated, where The number of signal sampling points. The lag order is... For the first The amplitude of the pseudo-random probe signal at each sampling point For the first The amplitude of the microphone received signal at each sampling point is analyzed; the growth trend of the function envelope at different frequency points is analyzed, and the critical frequency of the electroacoustic loop is extracted. and the corresponding critical gain estimate This formula directly applies the traditional normalized cross-correlation function formula and optimizes the range of sampling points by combining it with audio signal analysis scenarios; then, it extracts the target stability margin from the critical control command. Calculation based on gain constraint using the formula Gain safe working gain This formula is derived from the principle of loop stability constraints. It attenuates the critical gain estimate by using a target stability margin to ensure the loop operates in the stable region. Then, the digital gain value of the electroacoustic loop is set to the safe operating gain. A globally stabilizing main control loop is constructed, and the active notch filter center frequency is extracted from the critical control command. If the center frequency of the active notch filter is If present, an improved second-order adaptive notch filter is dynamically inserted into the forward path of the electroacoustic loop. This filter adds a frequency tracking module to the traditional second-order notch filter, enabling it to track the center frequency of the active notch filter in real time. A tiny offset, its center frequency locked at... The preset quality factor is set to 10, and then loop phase fine-tuning technology based on the minimum phase principle is used to improve the safe operating gain. exist and nearby Phase fine-tuning is performed within the frequency range, with a fine-tuning step size of [value missing]. This maximizes the noise suppression depth of the electroacoustic loop in this frequency band; among which The phase fine-tuning bandwidth is set to 50Hz in this embodiment. This bandwidth is determined based on the frequency response characteristics of the electroacoustic loop, covering the main fluctuation range between the critical frequency and the center frequency of the active notch filter. Finally, the output signal of the electroacoustic loop, after global gain stabilization, adaptive notch filtering, and loop phase fine-tuning, is reconstructed to obtain an intermediate audio signal that eliminates most environmental noise and has no feedback risk. .

[0085] For example, this embodiment assumes a data center voice communication scenario: a user wears a headset to conduct voice communication in the data center, where there is continuous low-frequency noise from the fan and high-frequency electromagnetic noise from the equipment, and the electroacoustic loop is prone to howling at certain high frequencies; firstly, the collected raw mixed signal Speech activity detection was performed to identify speech pauses as ,exist Injecting pseudo-random probe signals into the inward electroacoustic loop Synchronously acquire signals received by the microphone ,right and Closed-loop response cross-correlation analysis was performed to calculate the normalized cross-correlation function. After analyzing its envelope growth trend, the critical frequency was extracted. and critical gain estimate Then extract the target stability margin from the critical control command. The safe operating gain is obtained through gain constraint calculation. Then set the digital gain of the electroacoustic loop to... Extract the active notch center frequency from the critical control command. An improved second-order adaptive notch filter is inserted into the loop forward path to adjust the critical frequency. and active notch center frequency Phase fine-tuning is performed, and the intermediate audio signal is finally obtained through signal reconstruction. When the pseudo-random probe signal amplitude is 20dB below the ambient noise floor, the number of sampling points... The target stability margin is 1024. The critical gain estimate is 0.05. The center frequency of the active notch filter is 1.2. With a frequency of 3.2 kHz, a phase fine-tuning step size of 0.1 rad, and a quality factor of 10, the safe operating gain is calculated using gain constraints. A gain of 1.14 is applied to the electroacoustic loop, and an improved second-order adaptive notch filter with a center frequency of 3.2kHz is inserted. Phase fine-tuning is performed on the 3.2kHz frequency band and the ±50Hz frequency band near the critical frequency of 2.8kHz. After processing, the resulting intermediate audio signal achieves a noise suppression depth of 35dB in the 2.8kHz and 3.2kHz frequency bands without any feedback risk. The dual thresholds for voice activity detection are set to 0.01 and 0.03, respectively, and the spectral flatness error threshold for the pseudo-random probe signal is set to ±1dB. The values ​​for pseudo-random probe signal amplitude, number of sampling points, target stability margin, and phase fine-tuning step size in this embodiment are merely examples; those skilled in the art can set them according to actual conditions, and this embodiment does not impose any limitations.

[0086] This embodiment achieves precise extraction of critical parameters of the electroacoustic loop by injecting pseudo-random probe signals during speech intervals and combining closed-loop response cross-correlation analysis. A safe operating gain that balances stability and noise reduction is obtained through gain constraint calculation. Targeted suppression of threat frequencies is achieved through improved adaptive notch filtering. The noise suppression depth is optimized through loop phase fine-tuning technology based on the minimum phase principle. This enables precise control of the critical state of the electroacoustic loop and efficient suppression of environmental noise, improving the stability and targeted noise suppression of the noise reduction system while ensuring the fidelity of the user's voice path, providing a high-quality signal foundation for subsequent voice enhancement processing.

[0087] Furthermore, this embodiment provides a step for obtaining estimates of the critical frequency and critical gain based on the original mixed signal by injecting a pseudo-random probe signal into the electroacoustic loop and performing closed-loop response cross-correlation analysis, including:

[0088] Based on the original mixed signal, the microphone received signal is obtained by injecting pseudo-random self-noise probe signals;

[0089] Based on the self-noise probe signal and the microphone received signal, a cross-correlation function sequence is obtained by calculating the normalized cross-correlation function;

[0090] Based on the cross-correlation function sequence, the position and magnitude of the cross-correlation function peak are obtained through a peak detection algorithm;

[0091] Based on the position and amplitude of the peak value of the cross-correlation function, the critical frequency and the corresponding critical gain estimate are obtained through growth trend analysis.

[0092] The self-noise probe signal is a pseudo-random probe signal injected into the electroacoustic loop, with an amplitude smaller than the preset ambient noise floor value and a flat spectrum. The microphone received signal is a composite audio signal, including the loop response of the self-noise probe signal, ambient noise, and the original mixed signal, collected by the built-in microphone after the self-noise probe signal is injected into the electroacoustic loop. The normalized cross-correlation function is a cross-correlation function calculated after standardizing the self-noise probe signal and the microphone received signal, used to eliminate the influence of signal amplitude differences on the loop response characteristic analysis. The cross-correlation function sequence is a sequence formed by arranging the function values ​​obtained by calculating the normalized cross-correlation function for different lag orders according to the lag order, which is the core data characterizing the correlation between the self-noise probe signal and the microphone received signal. The peak detection algorithm is an improved extreme point detection algorithm, which adds amplitude to the traditional peak detection algorithm. The threshold and neighborhood verification module is used to accurately extract the position and amplitude information of effective peaks from the cross-correlation function sequence. The position and amplitude of the cross-correlation function peaks are the lag order positions corresponding to the effective peaks extracted from the cross-correlation function sequence by the peak detection algorithm and the normalized cross-correlation function values ​​at those positions. The position of the cross-correlation function peaks reflects the time delay of signal correlation, and the amplitude of the cross-correlation function peaks reflects the degree of signal correlation. The growth trend analysis is a method for fitting and analyzing the changing trend of the cross-correlation function peak amplitudes at different frequency points over time, used to identify the characteristic frequency and critical gain of the electroacoustic loop when it is close to instability. The critical frequency and the corresponding critical gain estimate are the characteristic frequencies identified by the growth trend analysis where the electroacoustic loop meets the oscillation conditions and the gain is close to 1, and the digital gain estimate that makes the electroacoustic loop reach a near-instability state at that frequency.

[0093] Specifically, the original mixed signal is first subjected to an improved dual-threshold speech activity detection (VAD). This algorithm adds spectral feature verification to the traditional dual-threshold VAD, improving the accuracy of speech interval recognition and identifying speech intervals in the original mixed signal. During these intervals, a self-noise probe signal is injected into the speaker output of the electroacoustic loop. ,in This serves as the sampling point index for the self-noise probe signal, an optimized implementation of a pseudo-random probe signal. It employs a Gaussian white noise sequence with lower amplitude and higher spectral flatness to further reduce interference with the user's auditory perception. The signal amplitude is set to be 20dB less than the preset ambient noise floor value, and the spectral flatness error does not exceed ±1dB within the range of 20Hz to 8kHz. Simultaneously, the response signal of the electroacoustic loop is synchronously acquired through the built-in microphone of the headset, resulting in a microphone received signal containing the loop response of the self-noise probe signal. Then, the self-noise probe signal was analyzed. With microphone to receive signals Perform frame synchronization processing to ensure that the time axes of the two signals are consistent, and then use the formula... Different lag orders were calculated. The corresponding normalized cross-correlation function value, where Given the number of sampling points for a single frame signal, the function values ​​of all lag orders are calculated according to... Arranging them in ascending order yields a sequence of cross-correlation functions. ,in To determine the maximum lag order, this formula is directly applied from the classical normalized cross-correlation function formula and adapted to the frame processing mode of audio signals; finally, an improved peak detection algorithm is used to process the cross-correlation function sequence. To process this, first set an amplitude threshold. Filter out values ​​greater than the amplitude threshold. The function values ​​are then subjected to 3-neighbor validation to determine the effective peak values ​​and extract their corresponding lag order positions. With amplitude Then, the lag order is converted to the corresponding frequency point, and the peak amplitude at each frequency point is... Time series fitting was performed using an exponential growth model. Fit its growth trend, where For time steps, , , The fitting coefficients are the growth coefficients obtained from the fitting process. The point with the highest frequency is determined as the critical frequency. Based on the amplitude growth trend at that frequency point and combined with the electroacoustic loop model, the calculation is performed to obtain the value that makes the loop function at that frequency. Critical gain estimate for reaching near-instability This exponential growth model is derived from the signal amplitude variation law of loop oscillation and is adapted to the signal characteristics of the critical state of the electroacoustic loop. The electroacoustic loop model is a mathematical model of a closed acoustic-electric signal transmission loop consisting of a digital signal processor, a loudspeaker, an acoustic transmission path in the ear canal, and a built-in microphone. It includes core parameters such as loop gain, phase response, and frequency response, and is used to establish a quantitative mapping relationship between the growth trend of the peak amplitude of the cross-correlation function and the actual loop gain.

[0094] This embodiment achieves precise injection of self-noise probe signals based on the speech intervals of the original mixed signal. The cross-correlation function sequence characterizing the loop response is calculated by normalized cross-correlation function. The effective peak features in the sequence are accurately extracted using an improved peak detection algorithm. The critical state parameters of the electroacoustic loop are identified by growth trend analysis using an exponential growth model. This achieves accurate and real-time extraction of the critical frequency and critical gain estimates of the electroacoustic loop, improving the accuracy and timeliness of critical parameter identification and providing a reliable parameter basis for subsequent gain stability control.

[0095] Furthermore, this embodiment provides a step for calculating the safe operating gain based on the critical gain estimate and critical control command through gain constraints, including:

[0096] Based on the target stability margin in the critical control command, the stability margin compensation coefficient is calculated through gain constraints.

[0097] The safe operating gain is calculated based on the critical gain estimate and the stability margin compensation coefficient.

[0098] Among them, the gain constraint is a numerical constraint rule based on the stability principle of the electroacoustic loop, which uses the target stability margin as a constraint condition to attenuate the critical gain estimate. It is used to ensure that the electroacoustic loop operates in the stable region while maximizing the noise reduction effect. The stability margin compensation coefficient is a coefficient calculated based on the target stability margin in the critical control command and used to attenuate the critical gain estimate. Its value is uniquely determined by the target stability margin and is an intermediate parameter connecting the target stability margin and the safe operating gain.

[0099] Specifically, the target stability margin, which characterizes the stability margin requirement of the electroacoustic loop, is first extracted from the critical control command. Based on gain constraints, through the formula The stability margin compensation coefficient can be directly calculated. This formula is derived from the principle of electroacoustic loop stability constraints. The larger the target stability margin, the smaller the stability margin compensation coefficient, and the higher the attenuation of the critical gain estimate. Then, the critical gain estimate of the electroacoustic loop is obtained. The critical gain estimate With stability margin compensation coefficient To perform multiplication, use the formula The safe operating gain of the electroacoustic loop was calculated. This formula is the core calculation formula for gain constraint, directly derived from the numerical relationship between the stability margin compensation coefficient and the critical gain estimate, and is suitable for the gain control scenario of the electroacoustic loop in this scheme; finally, the calculated safe operating gain is... Perform amplitude limiting and set a lower gain limit. Gain Limit If the safety working gain Less than Then take If the safety working gain Greater than Then take To avoid insufficient noise reduction due to too low a gain value or loop instability due to too high a gain value, this limiting rule is specifically designed for the hardware performance and noise reduction requirements of the call center headset in this solution.

[0100] This embodiment calculates the stability margin compensation coefficient based on the target stability margin in the critical control command through gain constraint. The safe operating gain is obtained by multiplying the compensation coefficient with the critical gain estimate and combining amplitude limiting processing. This achieves accurate and stable constraint of the electroacoustic loop operating gain, improves the rationality and reliability of gain control, and effectively balances the stability and noise suppression effect of the electroacoustic loop.

[0101] Furthermore, this embodiment provides a step of obtaining an enhanced speech signal based on an intermediate audio signal through perceptual domain spectral enhancement processing, including:

[0102] Based on the intermediate audio signal, the power spectrum characteristics of residual noise are obtained by speech activity detection and background noise spectrum estimation in the perceptual domain spectral enhancement processing.

[0103] Based on the power spectrum characteristics of residual noise, an enhanced speech signal is obtained through a non-uniform frequency response enhancement algorithm based on an auditory masking model in the perceptual domain spectral enhancement processing.

[0104] Among them, the speech activity detection uses an improved dual-threshold speech activity detection algorithm. Based on the traditional dual-threshold short-time energy and short-time zero-crossing rate, it introduces spectral entropy feature verification to accurately distinguish speech segments from noise segments in the intermediate audio signal. It calculates the spectral entropy of each frame and compares it with a preset entropy threshold to eliminate interference from sudden noise. The background noise spectrum estimation uses an improved algorithm based on minimum controlled recursive average (MCRA). It recursively smooths the noise power spectrum during speech intervals and introduces an adaptive update factor to track the power spectrum changes of residual noise in real time. The residual noise power spectrum... The feature is a spectral vector, obtained by estimating the background noise spectrum, which characterizes the energy distribution of residual background noise at each frequency point in the current output signal; the non-uniform frequency response enhancement algorithm based on the auditory masking model is a nonlinear spectrum processing algorithm that combines psychoacoustic masking effect, takes the residual noise power spectrum as input, and calculates the gain coefficient at each frequency point. It is used to selectively enhance the intermediate audio signal without introducing auditory distortion. The algorithm enhances speech segments with noise below the masking threshold by estimating the noise masking threshold, and maintains or slightly suppresses the noise-dominant frequency band, thereby highlighting the target speech components.

[0105] Specifically, firstly, the intermediate audio signal Frame segmentation is performed, where This is the global sampling point index for the intermediate audio signal, with the frame length set to... Point, frame shift Points, after adding a Hanning window, yield the signal frame. ,in For frame index, For each frame, an improved dual-threshold speech activity detection is performed, first calculating the frame short-time energy. and short-time zero crossing rate Simultaneously calculate spectral entropy ,in For the first The frame signal is in the first The frame spectrum normalization probability of each frequency point As a frequency index, this formula characterizes the signal type by quantifying the uniformity of the signal spectrum distribution. Speech signals have a spectrum concentrated in the formant region, resulting in a low spectral entropy value, while noise signals have a uniform spectrum distribution and a high spectral entropy value. A preset energy threshold is set. The preset zero-crossing rate threshold is And the preset spectral entropy threshold is Among them, the energy threshold Take twice the average energy of the silent segment, and use the zero-crossing rate threshold. The spectral entropy threshold is set to 1.5 times the average zero-crossing rate of the silent segment. 1.2 times the average spectral entropy of the silent segment is used to distinguish speech frames from noise frames; when the frame's short-time energy... Energy greater than the preset threshold And short-term zero crossing rate Less than the preset zero-crossing rate threshold And spectral entropy Less than the preset spectral entropy threshold If a frame is identified as a speech frame, it is considered a noise frame; otherwise, it is considered a noise frame. Within the noise frame interval, the improved minimum control recursive average (MCRA) algorithm is used to update the background noise power spectrum estimate. Power spectrum of each noise frame The recursive average is calculated using the following formula: ,in For the first Background noise power spectrum estimate after updating for each noise frame For the first The estimated background noise power spectrum of the previous noisy frame (i.e., the previous frame). For the first The original observed noise power spectrum of each noise frame, smoothing factor The system adaptively adjusts based on the speech presence probability of the current frame; when the speech presence probability exceeds a preset speech presence probability threshold... Updates are paused when the value approaches 1. The residual noise power spectrum characteristics of the current frame are obtained after updating multiple noisy frames. Then, based on the residual noise power spectrum characteristics, a non-uniform frequency response enhancement algorithm based on an auditory masking model is executed. First, the masking threshold for each frequency point is calculated according to the psychoacoustic model. The frequency domain is divided into 24 Barker sub-bands. Within each sub-band, the expanded masking energy is calculated based on the noise power spectrum, and the absolute hearing threshold is considered to obtain the masking threshold at each frequency point. Then calculate the short-time power spectrum of the current frame of the intermediate audio signal. And estimate the prior signal-to-noise ratio. and posterior signal-to-noise ratio The formula for calculating the gain function is: ,in Angular frequency, To prevent small constants from being divided by zero, This is an estimate of the residual noise power spectrum. This represents the short-time power spectrum of the intermediate audio signal. and To adjust the parameters, Adjusting parameters for masking effect, As a signal-to-noise ratio dependent adjustment parameter, this embodiment sets... , This parameter combination was determined through subjective testing of speech quality in multiple scenarios, achieving an optimal balance between noise suppression and speech fidelity. The first term reflects the masking effect, providing a smaller gain in frequency bands where noise easily masks speech, thus avoiding excessive noise enhancement. The second term is a signal-to-noise ratio (SNR) dependent term, where the gain approaches 1 in frequency bands where the SNR is greater than a preset first SNR threshold (i.e., the speech-dominant frequency band), and approaches 0 in frequency bands where the SNR is less than a preset second SNR threshold. The gain curve is smoothed and limited before being multiplied by the spectrum of the intermediate audio signal. Finally, the enhanced speech signal is obtained through inverse Fourier transform and overlapping addition. The effect threshold is determined through human hearing characteristic testing. In this embodiment, it is set to 15dB. When the masking threshold is greater than the effect threshold, noise can be naturally masked by the human ear without the need for additional gain enhancement. The first signal-to-noise ratio (SNR) threshold and the second SNR threshold are determined by collecting mixed samples of typical environmental noise and standard Chinese speech in a call scenario and calculating the speech objective intelligibility (STOI) and subjective hearing score (MOS) at different SNRs. In this embodiment, the values ​​are 10dB and -5dB, respectively. When the SNR is greater than 10dB, there is no obvious noise interference in subjective hearing, so it is set as the first SNR threshold. When the SNR is less than -5dB, the speech components are completely masked by noise, and subjective hearing cannot recognize effective speech, so it is set as the second SNR threshold.

[0106] For example, this embodiment assumes an office scenario: intermediate audio signal It contains residual low-frequency air conditioning noise and slight keyboard typing transients after critical stabilization control. Assume a sampling rate of 16kHz, a frame length of 512 points, and a frame shift of 256 points. First, for the... The frame is used for speech activity detection, and its short-time energy is calculated as follows: The short-term zero crossing rate Spectral entropy is According to the preset energy threshold Preset zero-crossing threshold Preset spectral entropy threshold The frame meets the criteria and is classified as a speech frame, therefore the background noise spectrum is not updated; the residual noise spectrum has already been updated in the previous three consecutive noise frames. Its energy is higher below 200Hz; then it enters the auditory masking model, and under the Bark subband division, the masking threshold of the subband near 200Hz is... The noise power spectrum at this frequency is 15 dB, resulting in a low signal-to-noise ratio; when calculating the gain, the first factor... However, after limiting, it is taken as 1.0; the second factor In the low-noise-dominant frequency band, the overall gain is approximately 0.3, attenuating low frequencies. In the 1kHz speech formant region, the corresponding noise spectrum is low while the signal-to-noise ratio is high, with the second factor close to 1 and a gain of approximately 1.0. Ultimately, the enhanced speech signal exhibits suppressed low-frequency noise and improved speech clarity. The values ​​of parameters such as frame length, threshold, smoothing factor, and masking model in this embodiment are merely examples and can be adjusted by those skilled in the art based on actual device performance and auditory characteristics. This embodiment does not impose any limitations on these adjustments.

[0107] This embodiment achieves high-precision time-frequency domain conversion of intermediate audio signals through an improved short-time Fourier transform, accurately extracts the power spectrum features of residual noise through improved speech activity detection and background noise spectrum estimation, achieves real-time tracking of residual noise power spectrum features through an improved minimum statistic recursive averaging algorithm, and achieves efficient suppression of residual noise and optimization of speech perception quality by combining a non-uniform frequency response enhancement algorithm with an auditory masking model. This realizes further suppression of residual noise and targeted enhancement of key speech frequency bands in call center scenarios, improving the clarity, intelligibility, and auditory comfort of the enhanced speech signal, perfectly meeting the core voice communication needs of call center headsets, while avoiding the electronic sound effects introduced by traditional post-processing, providing users with a more comfortable listening experience.

[0108] Furthermore, this embodiment provides a step for obtaining updated state decision mapping parameters through online reinforcement learning based on enhanced voice signals and historical critical control commands, including:

[0109] Speech quality indicators are obtained based on enhanced speech signals using a perceptual speech quality estimation algorithm.

[0110] Stability indices are obtained based on loop stability analysis of the intermediate audio signal.

[0111] A comprehensive evaluation scalar is obtained by weighted fusion of speech quality indicators and stability indicators;

[0112] Based on historical critical control commands, combined with speech quality indicators, stability indicators, and comprehensive evaluation scalars, the state decision mapper is optimized through online reinforcement learning based on stochastic gradient descent to obtain updated state decision mapping parameters.

[0113] The perceptual speech quality estimation algorithm is an improved perceptual speech quality assessment model. Based on traditional objective speech quality assessment methods, it simplifies the auditory transformation process and introduces frequency band energy proportion features to perform real-time quality scoring of enhanced speech signals. It outputs a speech quality index by calculating the spectral envelope difference and time-domain envelope stability between the enhanced speech and the reference signal. The reference signal is obtained from locally stored clean speech templates. The speech quality index is a value ranging from 0 to 1, used to quantify the clarity and naturalness of the enhanced speech signal; a higher value indicates higher speech quality. The loop stability analysis is a process of calculating the actual stability margin of the electroacoustic loop based on the residual energy of the probe signal. It obtains the stability index by analyzing the residual oscillation energy in the closed-loop response after each probe injection and comparing it with the oscillation energy under critical conditions. The stability index ranges from 0 to... The values ​​between 1 and 1 represent the relative distance of the current electroacoustic loop from the instability boundary; the closer the value is to 1, the more stable the loop. The comprehensive evaluation scalar is a scalar value obtained by weighted fusion of speech quality and stability indicators, used to balance the importance of noise reduction effect and system stability. This scalar serves as the reward signal for reinforcement learning. The online reinforcement learning based on stochastic gradient descent is a reinforcement learning algorithm that adopts a deep Q-network framework and introduces a priority experience replay mechanism. It is used to optimize the parameters of the state decision mapper based on historical experience data. It takes the environmental state vector as input and the action value function as output, and updates the network parameters by minimizing the temporal difference error. The updated state decision map parameters are a new set of parameters obtained through iterative optimization of reinforcement learning, including the connection weights, bias values, and attention mechanism-related parameters of each layer of neurons in the state decision mapper.

[0114] Specifically, the first step is to acquire the enhanced speech signal. The signal is compared with a locally cached clean speech template, which is obtained by inversely filtering the background noise spectrum collected when the user is silent. The two signals are then framed, and the Mel-frequency cepstral coefficients and their first-order differences for each frame are calculated. The Euclidean distance between the Mel-frequency cepstral coefficients of the two signals is then calculated, and the speech quality index is obtained by combining the inter-frame continuity. ,in The MFCC distance in the current frame. The preset maximum distance is used; simultaneously, the microphone received signal after the most recent probe injection is read from the system cache. Combine it with the injected probe signal Cross-correlation analysis was performed to calculate the residual energy. ,in To estimate the loop impulse response, the residual energy is compared with the residual energy threshold at the critical state. Stability indexes were obtained through comparison. , The adjustment factor is used to adjust the sensitivity of the stability index to changes in residual energy. It is determined through multi-scenario loop stability testing, and in this embodiment, it is set to 2.0. The closer the stability index is to 1, the more stable the loop. The critical state is the instant when the system has just experienced a howling. The residual energy threshold is the baseline value of the residual energy of the probe signal when the system is in the critical howling state. It is obtained by simulating the residual energy of the probe signal at the instant of howling in a laboratory environment and taking the average value. Then, the voice quality index is... and stability indicators The weighted fusion is used to obtain the comprehensive evaluation scalar. ,in , These are the weighting coefficients for the corresponding indicators. In this embodiment, it is assumed that... , The comprehensive evaluation scalar is used as the reward signal for reinforcement learning; next, training samples for online reinforcement learning are constructed, and the system records the environmental state vector at each decision in real time. ,in Index for discrete decision-making time points. For the first The instantaneous impact intensity at the moment of decision-making For the first A list of threat frequencies at each decision-making moment. For the first The equivalent spatial diffusion at each decision moment, and the threat frequency list need to be encoded as fixed-dimensional features, such as list length, first-frequency energy percentage, etc., corresponding to the critical control commands. As an action, and the reward obtained after performing that action. ,in For the first The target stability margin at each decision moment For the first The active notch center frequency at each decision moment (set to 0 when there is no threat). For the first The loop gain adjustment strategy identifier at each decision moment (0 indicates global gain adjustment, 1 indicates frequency-selective gain adjustment); the experience of each interaction. The data is stored in the experience replay pool. A deep Q-network is used as the reinforcement learning model. Its network structure is the same as that of the state decision mapper, containing three fully connected layers. The input layer has a dimension of 5, the hidden layers have dimensions of 64 and 32, and the output layer has a dimension equal to the action space size. For the first The environmental state vector at each decision moment For the first Critical control instructions (actions) executed at each decision moment. To perform the action The resulting comprehensive evaluation scalar (instant reward). To transfer to the first after the action is performed The environmental state vector at each decision moment; to accelerate convergence, a priority experience replay mechanism is introduced, which prioritizes the sampling of experience based on the magnitude of the temporal difference error, and samples a small batch of samples from the experience pool every fixed preset number of steps to calculate the target Q value. ,in For experience sample indexing. For the first The target Q value for an empirical sample For the first The instant reward corresponding to each experience sample This is a discount factor used to weigh the importance of immediate rewards against future rewards; in this example, it is set to 0.9. For the target network in the first Environmental state corresponding to each empirical sample All possible actions The corresponding maximum Q value; The parameters of the target network are copied from the current network at fixed preset steps; this is achieved by minimizing the mean squared error loss function. Update the current network parameters using a stochastic gradient descent optimizer. ,in Let the mean squared error loss function of the current network be denoted as . The size of the small batch of samples is 32 in this embodiment. For the first The target Q value for an empirical sample For the current network in the th Environmental state corresponding to each empirical sample Next action The corresponding predicted Q value, The trainable parameters of the current network are set, the learning rate of the stochastic gradient descent optimizer is set to 0.001, and the batch size is [not specified]. During training, the parameters of the state decision mapper are gradually updated, so that the critical control commands it outputs can obtain higher rewards. Finally, after each parameter update, the updated state decision map parameters are obtained.

[0115] For example, this embodiment assumes an office scenario where the system has accumulated some experience data after running for a period of time. Let's say at a certain moment... The environmental condition is: instantaneous impact intensity Threat Frequency List Equivalent spatial diffusion The state decision mapper outputs the action as follows: The preset weighting coefficient is After performing this action, an enhanced speech signal is obtained, which is then used by a perceptual speech quality estimation algorithm to obtain a speech quality index. The loop stability analysis yielded the following results. Then, rewards are obtained through weighted fusion. ; this experience The system stores an experience pool, samples a batch of experiences from the pool, calculates its target Q-value, and updates the network parameters using SGD. Assuming that after multiple iterations, the network parameters are adjusted to output a better action in similar environments (e.g., appropriately increasing the stability margin to improve stability), the updated state decision mapping parameters are obtained after the parameter update for that frame. Further exemplified, the experience pool capacity is set to 10000, the batch size to 32, the learning rate to 0.001, the discount factor to 0.9, and the target network update steps to 100. After 1000 parameter updates, the output of the state decision mapper in similar environments is... The value was adjusted to 0.07, resulting in a subsequent reward of 0.90. The values ​​of parameters such as weight coefficient, experience pool capacity, and learning rate in this embodiment are merely examples; those skilled in the art can adjust them according to actual needs, and this embodiment does not impose any limitations on this.

[0116] This embodiment achieves accurate quantification of speech quality by fusing segmented signal-to-noise ratio with a simplified objective speech quality assessment perceptual speech quality estimation algorithm. It achieves real-time evaluation of system stability through loop stability analysis based on residual energy of probe signals. Weighted fusion yields a comprehensive evaluation scalar that considers both objectives. An improved online reinforcement learning framework based on stochastic gradient descent dynamically optimizes the state decision mapping parameters. This enables the noise reduction system to adaptively learn and continuously evolve in complex and changing communication environments, improving the decision accuracy of the state decision mapper and the environmental adaptability of the noise reduction strategy. It ensures that the headset maintains optimal noise reduction performance and speech quality in different scenarios, making the noise reduction strategy more tailored to individual user needs and environmental changes.

[0117] Furthermore, this embodiment provides a step for optimizing the state decision mapper using online reinforcement learning based on stochastic gradient descent, based on historical critical control commands combined with speech quality indicators, stability indicators, and a comprehensive evaluation scalar, to obtain updated state decision mapping parameters. The steps include:

[0118] Based on historical critical control instructions, an environment state vector and action vector are constructed through online reinforcement learning based on stochastic gradient descent, and a reward scalar is constructed based on a comprehensive evaluation scalar.

[0119] Based on the environment state vector, action vector, and reward scalar, the state decision mapper is optimized using a stochastic gradient descent strategy to obtain updated state decision mapping parameters.

[0120] The environment state vector is a vector composed of multi-dimensional noise features at the current moment, including instantaneous impact intensity, features encoded from the threat frequency list, and equivalent spatial diffusion, used to describe the impact of the current environment on the stability of the acoustic-electric closed-loop system; the action vector is the critical control command output by the state decision mapper under the current environment state, including target stability margin, active notch center frequency, and loop gain adjustment strategy identifier; the reward scalar is a comprehensive evaluation scalar, obtained by weighted fusion of speech quality index and stability index, used to evaluate the current action; the stochastic gradient descent strategy is a first-order optimization method used in deep reinforcement learning, which calculates the gradient of the loss function with respect to the network parameters and backpropagates to update the parameters, enabling the network to learn the optimal action value function.

[0121] Specifically, a series of time-series empirical data are first extracted from the historical database. Each empirical tuple includes an environmental state vector. Action vectors , reward scalar and the next environment state vector ; Environment state vector The construction method is to use the instantaneous impact strength Directly used as the first component; for the threat frequency list Encode the list and get its length. The energy percentage of the frequency corresponding to the highest energy in the list. And the minimum deviation between this frequency and the preset characteristic frequency. As the second to fourth components, the preset characteristic frequency, namely the characteristic frequency of the headset and ear canal system mentioned above, is obtained by offline acoustic calibration of the headset and ear canal system and combined with open-loop frequency response testing, based on the physical acoustic path characteristics from the headset speaker to the microphone; the equivalent spatial diffusion is also considered. As the fifth component, it ultimately forms a five-dimensional environment state vector. Action vector From the target stability margin Active notch center frequency and loop gain adjustment strategy identifier The constructed three-dimensional vector, where the active notch center frequency is 0 if it does not exist, the loop gain adjustment strategy identifier includes 0 or 1, and the reward scalar. This refers to the comprehensive evaluation scalar; the state decision mapper is used as the policy network, i.e., the Q-network, in reinforcement learning, and its input is the environment state vector. The output is the Q-value for each possible action. Since the action space is continuous and multidimensional, a deep deterministic policy gradient framework is adopted, including an actor network and a critic network. The actor network is responsible for outputting deterministic actions based on the state, and its structure is the same as the state-decision mapper. The critic network takes state-action pairs as input and outputs Q-values. During training, mini-batch samples are randomly sampled from the experience pool. Calculate the target value of the critic network. ,in For experience sample indexing. This is a discount factor used to weigh the importance of immediate rewards against future rewards; in this example, it is set to 0.95. For the target actor network, For the parameters of the target actor network, For the target critic network, Let be the parameters of the target critic network; the critic network parameters are updated by minimizing the critic loss function, and then the actor network parameters are updated by the policy gradient of the sampled samples, where the learning rates of the actor network and the critic network are set to 0.0001 and 0.001 respectively, the batch size is 64, and the soft update coefficient of the target network is . After multiple iterations and optimizations, the parameters of the actor network, i.e., the state decision mapper, are updated, and the updated state decision map parameters are finally output.

[0122] This embodiment constructs environmental state vectors and action vectors based on historical critical control commands, constructs a reward scalar based on a comprehensive evaluation scalar, and optimizes the state decision mapper through a deep deterministic policy gradient algorithm. This achieves a deep fusion of noise reduction control parameters and historical environmental experience, enabling efficient policy learning in a continuous action space. It improves the stability and convergence speed of reinforcement learning, enhances the optimization efficiency and generalization ability of the state decision mapper parameters, and ensures that the state decision mapper can output critical control commands that better meet environmental requirements. This allows the noise reduction system to quickly adapt to and optimize noise reduction strategies in complex dynamic environments, further improving the system's intelligence and personalization level.

[0123] Furthermore, such as Figure 4 As shown in the figure, this application provides an adaptive noise reduction system for call center headsets based on ambient sound detection. The system includes a voice decision-making system, a voice control stabilization system, a voice enhancement system, and a policy learning system.

[0124] The voice decision system acquires raw mixed signals based on the microphone built into the headset, performs multi-dimensional environmental noise analysis, and makes decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands.

[0125] The voice-controlled stabilization system, based on critical control commands combined with the original mixed signal, obtains the intermediate audio signal through online estimation of the critical stabilization point and gain stabilization control.

[0126] The speech enhancement system obtains an enhanced speech signal by performing perceptual domain spectral enhancement processing on the intermediate audio signal.

[0127] The strategy learning system, based on enhanced speech signals and historical critical control commands, obtains updated state decision mapping parameters through online reinforcement learning.

[0128] The voice decision-making system, voice-controlled stabilization system, voice enhancement system, and policy learning system are all located on the server. The server receives the raw mixed signal transmitted by the acquisition device and performs core signal analysis and intelligent decision-making on the server side. The acquisition device includes a call center headset and its built-in microphone, used to acquire the raw mixed signal containing user voice, environmental noise, and potential acoustic feedback in real time, and send the signal to the server for processing. The voice enhancement system can be deployed locally on the call center headset to perform final voice quality enhancement and output it to the user after the server returns the intermediate audio signal. After the acquisition device transmits the raw mixed signal to the server, the voice decision-making system first performs multi-dimensional environmental noise analysis, extracts instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, and generates critical control commands based on the current parameters of the state decision mapper, and then sends the commands to the voice-controlled stabilization system. The voice-controlled stabilization system, based on received critical control commands, performs online estimation of the critical stability point of the synchronously received raw mixed signal, calculates the safe operating gain, and processes the signal using adaptive notch filtering and loop phase fine-tuning techniques to output an intermediate audio signal. This intermediate audio signal is transmitted to a voice enhancement system deployed locally in the headset. The voice enhancement system uses perceptual domain spectral enhancement processing to obtain the final enhanced voice signal for the user to listen to. Simultaneously, the enhanced voice signal and related historical critical control commands are sent back to the policy learning system on the server. The policy learning system performs perceptual voice quality assessment on the enhanced voice signal and, combined with loop stability analysis, weights and fuses the voice quality and stability indices into a comprehensive evaluation scalar as a reward signal. This scalar optimizes the parameters of the state decision mapper through online reinforcement learning based on stochastic gradient descent and feeds the updated parameters back to the voice decision system, thus forming a complete closed-loop adaptive noise reduction system from environmental perception and intelligent control to effect evaluation and self-optimization.

[0129] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0131] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0132] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. An adaptive noise reduction method for call center headsets based on ambient sound detection, characterized in that, include: The system acquires the raw mixed signal using the microphone built into the headset, performs multi-dimensional environmental noise analysis, and makes decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands. Based on the critical control command and the original mixed signal, the intermediate audio signal is obtained through online estimation of the critical stable point and gain stabilization control. Based on the intermediate audio signal, an enhanced speech signal is obtained through perceptual domain spectral enhancement processing; Based on the enhanced voice signal and historical critical control commands, updated state decision mapping parameters are obtained through online reinforcement learning.

2. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 1, characterized in that, The process involves acquiring the raw mixed signal using the microphone built into the headset, performing multi-dimensional environmental noise analysis, and making decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands, including: Based on the original mixed signal, the instantaneous impact intensity is obtained through short-time energy analysis and differential calculation; Based on the original mixed signal and the preset characteristic frequencies of the headphone and ear canal system, a list of threat frequencies is obtained through power spectrum estimation. Based on the original mixed signal, the equivalent spatial diffusion is obtained by calculating the time-varying statistical characteristics of the signal. Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, a critical control command is obtained by multi-feature fusion and decision-making through a state decision mapper. The critical control command includes target stability margin, active notch center frequency, and loop gain adjustment strategy.

3. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 2, characterized in that, Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, a state decision mapper performs multi-feature fusion and decision-making to obtain critical control commands. These critical control commands include target stability margin, active notch filter center frequency, and loop gain adjustment strategy, including: Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the target stability margin is obtained through the stability margin mapping rule in the state decision mapper. Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, the active notch center frequency is obtained through the threat frequency selection rule in the state decision mapper. Based on the instantaneous impact intensity, threat frequency list, and equivalent spatial diffusion, a loop gain adjustment strategy is obtained through the gain strategy mapping rule in the state decision mapper. Based on the target stability margin, active notch filter center frequency, and loop gain adjustment strategy, critical control commands are obtained through integration.

4. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 1, characterized in that, The process of obtaining an intermediate audio signal based on the critical control command combined with the original mixed signal through online estimation of the critical stability point and gain stabilization control includes: Based on the original mixed signal, by injecting pseudo-random probe signals into the electroacoustic loop and performing closed-loop response cross-correlation analysis, the critical frequency and critical gain estimates are obtained. Based on the critical gain estimate and critical control command, the safe operating gain is calculated through gain constraints. Based on the critical frequency, safe operating gain, and critical control command, an intermediate audio signal is obtained through adaptive notch filtering and loop phase fine-tuning.

5. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 4, characterized in that, Based on the original mixed signal, by injecting a pseudo-random probe signal into the electroacoustic loop and performing closed-loop response cross-correlation analysis, the estimated values ​​of the critical frequency and critical gain are obtained, including: Based on the original mixed signal, the microphone received signal is obtained by injecting a pseudo-random self-noise probe signal; Based on the self-noise probe signal and the microphone received signal, a cross-correlation function sequence is obtained by calculating the normalized cross-correlation function; Based on the cross-correlation function sequence, the position and amplitude of the cross-correlation function peak are obtained through a peak detection algorithm; Based on the position and amplitude of the peak value of the cross-correlation function, the critical frequency and the corresponding critical gain estimate are obtained through growth trend analysis.

6. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 4, characterized in that, The process of calculating the safe operating gain based on the critical gain estimate and critical control command through gain constraints includes: Based on the target stability margin in the critical control command, the stability margin compensation coefficient is calculated through gain constraints. The safe operating gain is calculated based on the critical gain estimate and the stability margin compensation coefficient.

7. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 1, characterized in that, The process of obtaining an enhanced speech signal based on the intermediate audio signal through perceptual domain spectral enhancement includes: Based on the intermediate audio signal, the power spectrum characteristics of the residual noise are obtained through speech activity detection and background noise spectrum estimation in the perceptual domain spectrum enhancement processing. Based on the power spectrum characteristics of the residual noise, an enhanced speech signal is obtained through a non-uniform frequency response enhancement algorithm based on an auditory masking model in the perceptual domain spectral enhancement processing.

8. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 1, characterized in that, The updated state decision mapping parameters, obtained through online reinforcement learning based on the enhanced speech signal and historical critical control commands, include: Based on the enhanced speech signal, a speech quality index is obtained through a perceptual speech quality estimation algorithm; Stability indices are obtained through loop stability analysis based on the intermediate audio signal. A comprehensive evaluation scalar is obtained by weighted fusion of the aforementioned speech quality and stability indicators. Based on historical critical control commands, combined with the aforementioned speech quality index, stability index, and comprehensive evaluation scalar, the state decision mapper is optimized through online reinforcement learning based on stochastic gradient descent to obtain updated state decision mapping parameters.

9. The adaptive noise reduction method for call center headsets based on ambient sound detection according to claim 8, characterized in that, The state decision mapper is optimized using online reinforcement learning based on stochastic gradient descent, combining historical critical control commands with the speech quality index, stability index, and comprehensive evaluation scalar, to obtain updated state decision map parameters, including: Based on historical critical control instructions, an environment state vector and action vector are constructed through online reinforcement learning based on stochastic gradient descent, and a reward scalar is constructed based on the comprehensive evaluation scalar. Based on the environmental state vector, action vector, and reward scalar, the state decision mapper is optimized using a stochastic gradient descent strategy to obtain updated state decision mapping parameters.

10. An adaptive noise reduction system for call center headsets based on ambient sound detection, used to implement the adaptive noise reduction method for call center headsets based on ambient sound detection as described in any one of claims 1-9, characterized in that, The adaptive noise reduction system for call center headsets based on ambient sound detection includes: The voice decision system acquires raw mixed signals based on the microphone built into the headset, performs multi-dimensional environmental noise analysis, and makes decisions based on the state decision mapping parameters of the state decision mapper to obtain critical control commands. The voice-controlled stabilization system, based on the critical control command and the original mixed signal, obtains the intermediate audio signal through online estimation of the critical stabilization point and gain stabilization control; The speech enhancement system obtains an enhanced speech signal based on the intermediate audio signal through perceptual domain spectral enhancement processing; The policy learning system obtains updated state decision mapping parameters through online reinforcement learning based on the enhanced voice signal and historical critical control commands.