A causal reference-enhanced active noise reduction method for improving speech intelligibility and speech quality
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]本发明旨在解决现有主动噪声控制系统中语音与噪声不可区分的问题,避免语音被误抑制,同时满足实时系统的因果性与低延迟要求
[0028](1)因果约束裕度高,无额外算法延迟。
Smart Images

Figure CN122551815A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech signal processing and active noise control technology, specifically to a deep learning-based reference signal enhancement active noise control method, applicable to scenarios where noise and speech signals coexist and speech needs to be preserved while suppressing noise, including but not limited to applications such as headphones, in-vehicle systems, and wearable audio devices. Background Technology
[0002] With industrialization and urbanization, noise pollution has become increasingly prominent, adversely affecting human comfort and health. In some scenarios, passive noise control methods often have limited effectiveness due to limitations in material volume and weight. In contrast, active noise control (ANC), due to its ease of implementation, lightweight structure, and excellent suppression of low-frequency noise, has been widely used in headphones, automobiles, and aircraft.
[0003] However, noise also degrades the quality of spoken conversations, and neither passive nor active noise cancellation technologies can distinguish between noise and speech. Therefore, noise suppression inevitably affects the speech signal. For example, in active noise-canceling headphones, the reference microphone simultaneously captures ambient noise and speech signals, causing the ANC system to suppress both noise and speech components at the ear; combined with the passive sound insulation effect of the headphone shell, this significantly reduces the quality of spoken conversations.
[0004] In dialogue scenarios, besides disabling ANC, there are some hear-through solutions, such as using high-pass filtering and gain compensation to compensate for passive attenuation of the external microphone signal, thus achieving acoustically transparent hearing. However, these methods also transmit noise into the ear. To address this, some research combines hyperdirectional microphone array beamforming with multi-reference feedforward ANC to construct a hear-through headphone system with directional selectivity (V. Patel, J. Cheer, and S. Fontana, “Design and implementation of an active noise control headphone with directional hear-through capability,” IEEE Trans. Consumer Electronics, vol. 66, no. 1, pp. 32-40, 2019.). In this system, ambient sound is first attenuated by ANC, and then the beamformed signal is injected into the ear, achieving spatially selective auditory perception. However, this method fails when speech and noise come from the same direction, and its performance is highly dependent on the quality of beamforming.
[0005] The method of suppressing noise while preserving speech using active noise control techniques can be called "keep-speech ANC" (KSANC). Zhang et al. proposed a DeepANC method (H. Zhang and D. Wang, "Deep ANC: A deep learning approach to active noise control," Neural Networks, vol. 141, pp. 1-10, 2021.), which uses a convolutional recurrent network (CRN) to replace the control filter in the traditional ANC system, thereby achieving KSANC and improving speech intelligibility and speech quality in error signals. However, CRN-based methods inevitably introduce frame-level delays, making it difficult to meet the stringent causality requirements of practical applications (such as headphones).
[0006] Therefore, how to effectively suppress noise while preserving speech information while ensuring low system latency and causal constraints is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0007] The present invention aims to solve the problem of indistinguishability between speech and noise in existing active noise control systems, avoid speech being falsely suppressed, and at the same time meet the causality and low latency requirements of real-time systems.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0009] A causal reference-enhanced active noise reduction method for improving speech intelligibility and speech quality includes the following steps:
[0010] Step 1, acquire the reference microphone signal x v (n) and the error microphone signal e(n), where the reference microphone signal contains a noise component x(n) and a speech component v. r (n), that is:
[0011] x v (n)=x(n)+v r (n)
[0012] The error microphone signal includes the original noise d(n) at the error location and the original speech v. e (n) and secondary sound y(n), that is:
[0013] e(n)=d(n)+v e (n)+y(n)
[0014] Step 2, based on the error microphone signal e(n), the secondary control signal u(n), and the estimation of the secondary path (transfer function is...) Its time-domain impulse response is This yields an estimate of the original error signal:
[0015]
[0016] Step 3, the reference microphone signal x v (n) and the estimation of the original error signal The input is fed into a pre-trained neural network model that satisfies causal constraints to obtain an enhanced reference signal:
[0017]
[0018] Where g(·) represents a neural network, Represents network parameters;
[0019] Step 4, the enhanced reference signal The secondary control signal u(n) is generated by inputting the control filter w (with filter coefficients w) of the active noise control system:
[0020]
[0021] Step 5: The secondary control signal u(n) propagates through the secondary path (transfer function s, time-domain impulse response s) to generate secondary sound at the error microphone:
[0022] y(n)=s*u(n)
[0023] Its superposition with the original noise achieves noise suppression. Assuming the secondary path estimation is unbiased, the error signal can be further expressed as:
[0024]
[0025] During the training phase of the neural network model, a loss function L is used. enh Optimize neural network parameters Where the loss function L enh It includes a loss term for simultaneously optimizing noise suppression and speech preservation. During the application phase, the neural network parameters... Fixed. L enh The expression is:
[0026]
[0027] Compared with the prior art, the present invention has the following advantages:
[0028] (1) High causal constraint margin, no additional algorithm delay.
[0029] (2) It has good speech preservation effect and avoids secondary sound cancellation of speech components. It is superior to existing methods in both speech intelligibility (STOI) and speech quality (DNSMOS).
[0030] (3) It is robust and suitable for different noise types, incident directions and signal-to-noise ratios. Attached Figure Description
[0031] Figure 1 This is a schematic diagram illustrating the principle of the method of the present invention, using headphones as an example.
[0032] Figure 2 This is a block diagram of the deep neural network structure used in the method of this invention.
[0033] Figure 3 This is a schematic diagram of the measurement of the transfer function used in the embodiments of the present invention.
[0034] Figure 4 These are photographs of the transfer function measurement in an embodiment of the present invention. (a) shows the measurement environment, and (b) shows details related to the headphones.
[0035] Figure 5Here is a sample from an embodiment of the present invention: (a) is the time-domain waveform of the error signal, and (b) is the time-frequency diagram of the error signal. Detailed Implementation
[0036] The technical solution of the present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these examples are only used to illustrate the present invention and are not intended to limit the scope of the present invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0037] A causal reference-enhanced active noise reduction method for improving speech intelligibility and speech quality, such as Figure 1 As shown, it includes the following steps:
[0038] Step 1, acquire the reference microphone signal x v (n) and the error microphone signal e(n), where the reference microphone signal contains a noise component x(n) and a speech component v. r (n), that is:
[0039] x v (n)=x(n)+v r (n)
[0040] The error microphone signal includes the original noise d(n) at the error location and the original speech v. e (n) and secondary sound y(n), that is:
[0041] e(n)=d(n)+v c (n)+y(n)
[0042] Step 2, based on the error microphone signal e(n), the secondary control signal u(n), and the estimated secondary path (transfer function is...) Its time-domain impulse response is This yields an estimate of the original error signal:
[0043]
[0044] Step 3, the reference microphone signal x v (n) and the estimation of the original error signal The input is fed into a pre-trained neural network model that satisfies causal constraints to obtain an enhanced reference signal:
[0045]
[0046] Where g(·) represents a neural network, These represent the network parameters. In this embodiment, the neural network adopts a temporal modeling structure based on causal convolution, and is an overall multi-layer convolutional stacked network, as shown in the figure. Figure 2 As shown, the network input first undergoes feature extraction through a causal convolutional layer (Causal Conv), followed by deep feature modeling through several stacked convolutional modules. Each convolutional module may include multiple convolutional layers and a nonlinear activation function (tanh, σ). The convolutional layers may employ a dilated convolutional structure to expand the receptive field and achieve information transfer through residual connections or skip connections. In one specific implementation, the network may include multiple sets of stacked convolutional structures, each set containing several convolutional layers. The number of channels in each convolutional layer can be set to tens to hundreds, the kernel size can be set to several sampling points, and the stride can be set to 1 to ensure causality and no additional latency. The network output can undergo channel compression through multiple causal convolutions, ultimately outputting an enhanced reference signal. The specific parameters are described as follows: the kernel size of the input causal convolutional layer is 32, and the number of channels is 128; the convolutional stacking structure includes 4 modules, each module contains 12 convolutional layers, each layer has 128 channels, and the kernel size is 5; the output includes multiple causal convolutional layers, with the number of channels being 128, 64, 32, and 1 respectively, corresponding to kernel sizes of 5, 5, and 256.
[0047] Step 4, the enhanced reference signal The secondary control signal u(n) is generated by inputting the control filter w (with filter coefficients w) of the active noise control system:
[0048]
[0049] Step 5: The secondary control signal u(n) propagates through the secondary path (transfer function s, time-domain impulse response s) to generate secondary sound at the error microphone:
[0050] y(n)=s*u(n)
[0051] Its superposition with the original noise achieves noise suppression. Assuming the secondary path estimation is unbiased, the error signal can be further expressed as:
[0052]
[0053] During the training phase of the neural network model, a loss function L is used. enh Optimize neural network parameters In the application phase, neural network parameters Fixed. Where L enh The expression is:
[0054]
[0055] In this embodiment, to construct the acoustic model of the active noise control system, measurements are taken of each transmission path in the system. For example... Figure 3 As shown, multiple speakers are arranged at preset angles around the head model to simulate speech and noise signals, respectively. The speaker simulating the speech signal source is positioned 0.6m from the center of the head, while the speaker simulating the noise signal source is positioned 1.0m from the center of the head. Figure 4 Photographs of the measurement environment and related headphone details are provided. The reverberation time of the measurement environment is approximately 0.25 s. During the measurement process, the left and right channels of the headphones are treated as independent system channels; this embodiment selects one channel for analysis. By inputting test signals to speakers of simulated noise and speech signal sources located at different positions, the response signals to the reference and error microphones are acquired, thereby obtaining the corresponding impulse response function. The impulse response length is 256 sampling points, and the sampling rate is 8 kHz. Simultaneously, the secondary path between the headphone speaker and the error microphone is measured, and the resulting impulse response is truncated to 128 sampling points to reduce subsequent computational complexity. During data construction, the incident directions of the speech and noise signal sources can be selected within a preset angle range or randomly set to improve the system's adaptability.
[0056] In this embodiment, to train and evaluate the active noise control method, an experimental dataset including noise data and speech data is constructed. The noise data includes two types: synthetic noise and actual acquired noise. The synthetic noise is generated by dividing the 0-4000Hz frequency band into multiple sub-band signals with different bandwidths, and a corresponding band-limited noise signal is generated based on each sub-band to construct the training set, validation set, and test set. For each sub-band, multiple noise samples of preset duration (e.g., 10 seconds) can be generated. Furthermore, to verify the performance of the method in a real-world environment, various types of actual noise data, such as equipment operating noise, can be selected to test the method. Speech data can be selected from publicly available speech datasets and divided according to the speaker to construct the training set, validation set, and test set, with no overlap between speakers in the datasets. The speech signal can be segmented according to a preset duration to match the noise data. During data construction, the energy ratio of the speech signal and the noise signal is adjusted to ensure that the signal-to-noise ratio at the reference microphone meets a preset range. During the training phase, the signal-to-noise ratio (SNR) can be selected from multiple discrete values; during the testing phase, multiple different SNR conditions can be set to evaluate the performance of the method under different noise intensities.
[0057] In this embodiment, the neural network model is trained using supervised learning. During training, a gradient descent-based optimization algorithm, such as the Adam optimization algorithm, is used to update the network parameters. A preset learning rate is set during training and adjusted based on the number of training epochs or validation performance. The number of training epochs can be set to a certain number based on the model's convergence. For training data input, the network is trained using a batch input method, with each batch containing several training samples. The loss function can be an optimization objective function based on the error signal.
[0058] In this embodiment, the proposed method was verified through system experiments. In the experiments, the control filter of the active noise control system was adaptively updated using the recursive least squares (RLS) algorithm, with a forgetting factor set to 0.99999 and an initial covariance matrix set to 0.01I. System performance was evaluated using the speech intelligibility index STOI and the speech quality index DNSMOS, calculated based on the error signal in the last 3 seconds after the system reached steady state. In simulation experiments, the method was compared with unprocessed signals, the ideal speech-preserving active noise reduction method (IdealKSANC), the conventional active noise control method (Conventional ANC), DeepANC, and the method of this invention. IdealKSANC represents using a pure noise signal x(n) as the reference signal, only indicating the theoretical performance limit of KSANC, which is not feasible in practical applications. Conventional ANC corresponds to directly using the original reference signal x. v The traditional ANC method (n).
[0059] Table 1 shows the speech intelligibility (STOI) and speech quality (DNSMOS) of the error signals obtained using different methods at different signal-to-noise ratios.
[0060]
[0061] Experimental results show that, under different signal-to-noise ratio (SNR) conditions, the method of this invention outperforms the comparative methods in terms of speech intelligibility and speech quality. Specifically, as shown in Table 1, the method of this invention achieves the highest STOI and DNSMOS values under multiple SNR conditions. In contrast, traditional active noise control methods, due to the simultaneous suppression of speech components during noise reduction, have lower STOI values than the unprocessed signal, especially under medium-to-high SNR conditions where performance degrades significantly. Furthermore, the DeepANC method suffers from significant processing delays, making it difficult to meet the system's causality requirements, and its overall performance is inferior to that of the method of this invention.
[0062] Table 2 shows the average speech intelligibility (STOI) of the error signals obtained using different methods under real noise conditions.
[0063]
[0064] Table 3 shows the average speech quality (DNSMOS) of the error signals obtained using different methods under real noise conditions.
[0065]
[0066]
[0067] The results under real-world noise environments are shown in Tables 2 and 3. The method of this invention outperforms existing methods under various real-world noise types (including fan noise, engine noise, bearing noise, and gear noise). In high signal-to-noise ratio (SNR) scenarios, the performance of traditional active noise control methods is inferior to that of unprocessed methods because it simultaneously suppresses speech components and relatively weak noise. However, in low SNR scenarios, noise significantly reduces the STOI and DNSMOS of speech. In this case, traditional active noise control methods can suppress the main noise components, thus improving the corresponding STOI and DNSMOS. By simultaneously suppressing noise and preserving speech, the method of this invention consistently outperforms existing methods. Furthermore, experimental results show that the method of this invention has good robustness and generalization ability under different speech signals, different noise types, and different sound source directions. Figure 5 Taking typical fan noise with a signal-to-noise ratio of -5dB as an example, the time-domain waveform and time-frequency plot of the error signal are given. Figure 5 The six scenarios represent: original speech signal, superposition of noise and speech signal (unprocessed), Ideal KSANC (ideal speech-preserving active noise reduction), traditional active noise control (ANC), DeepANC, and the method of this invention (RSE-based KSANC). It can be seen that the method of this invention can effectively reduce noise components while preserving speech information, thereby significantly improving speech intelligibility and speech quality.
[0068] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A causal reference-enhanced active noise reduction method for improving speech intelligibility and speech quality, characterized in that, Includes the following steps: Step 1: Obtain the reference microphone signal and the error microphone signal, wherein the reference microphone signal includes noise components and speech components; Step 2: Based on the error microphone signal, the secondary control signal, and the estimation of the secondary path, obtain the estimate of the original error signal; Step 3: Input the estimated values of the reference microphone signal and the original error signal into a pre-trained neural network model that satisfies causal constraints to obtain the enhanced reference signal; wherein, the neural network model is used to suppress speech components in the input signal and retain noise components; Step 4: Input the enhanced reference signal into the control filter of the active noise control system to generate the secondary control signal; Step 5: The secondary control signal propagates through the secondary path to generate secondary sound at the error microphone. The secondary sound is superimposed on the original noise to achieve noise suppression.
2. The method of claim 1, wherein, The neural network model that satisfies the causal constraint is a temporal convolutional network, which adopts a causal convolutional structure and does not introduce additional processing delay.
3. The method according to claim 2, characterized in that, The temporal convolutional network includes multiple cascaded dilated causal convolutional modules. Each module has a residual connection structure and fuses features from multiple layers through skip connections.
4. The method of claim 3, wherein, The network output layer uses a multi-layer causal convolutional structure for channel dimensionality reduction.
5. The method of claim 1, wherein, The pre-trained neural network model mentioned in step 3 is obtained by training a joint optimization loss function based on the error signal. The joint optimization loss function includes a loss term for simultaneously optimizing noise suppression and speech preservation.