Noise reduction methods that preserve speech
By integrating a microphone sensor and blind source separation algorithm into the earcups, the transmission of voice signals is preserved while isolating noise, solving the problem that traditional earcups cannot communicate directly, and achieving independent voice preservation and noise isolation effects.
Patent Information
- Application Number
- CN202411014011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-07-26
AI Technical Summary
Traditional noise-canceling earmuffs isolate voice signals along with noise, preventing wearers from communicating directly with others. Existing methods rely on wireless communication devices for voice signal transmission, lacking independence.
The device employs earmuffs with physical sound insulation, combined with a microphone sensor, microprocessor module, audio driver circuit, and speaker. It utilizes convolutional blind source separation algorithm and speech signal discrimination method to separate and preserve noise and speech signals, and transmits speech signals directly inside the earmuffs.
While isolating noise, it effectively preserves voice signals, enabling wearers to communicate directly without relying on external devices, simplifying the usage process, improving the accuracy of voice signal transmission, and reducing noise interference.
Smart Images

Figure CN119402791B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of blind source separation and noise-canceling headphone technology, specifically a noise reduction method that preserves speech. Background Technology
[0002] Traditional noise-canceling earmuffs can isolate most everyday noises; however, while isolating noise, they also isolate speech signals, preventing wearers from communicating directly with others while wearing them.
[0003] Currently, one method for communicating while wearing earmuffs is through wireless communication. However, this method requires a microphone from another device to transmit sound signals to the earmuffs via wired or wireless means. This process must be conducted between different devices and lacks independence. Summary of the Invention
[0004] To address the problems of existing technologies, this invention provides a noise reduction method that preserves speech. Based on the physical noise reduction earmuffs that isolate noise, the invention utilizes blind source separation technology and speech signal discrimination methods to separate and preserve speech signals, which are then transmitted into the earmuffs. This allows the earmuffs to retain speech signals while isolating noise, enabling the wearer to have direct voice communication with others.
[0005] This invention provides a noise-canceling earmuff that preserves speech, comprising an earmuff with a physical sound insulation structure, an external microphone sensor for collecting external sound signals, and an internal microprocessor module, an audio driver circuit, a speaker, and a power supply module. The microphone sensor is connected to the microprocessor module via an ADC interface, the microprocessor module is connected to the audio driver circuit via a DAC interface, and the audio driver circuit is connected to the speaker via a filter and an amplifier.
[0006] The present invention also provides a noise reduction method that preserves speech, characterized by comprising the following steps:
[0007] 1) Use earmuffs with physical sound insulation structure and use external microphone sensors on the earmuffs to collect external sound signals, including noise signals and voice signals;
[0008] 2) Use a convolutional blind source separation algorithm to separate noise signals and speech signals;
[0009] 3) Determine whether the noise signal and the speech signal are completely separated:
[0010] 3.1) Determine the similarity between two signals based on the cross-correlation discrimination method. If the cross-correlation is lower than the set threshold, proceed to the next step; if the cross-correlation is higher than the set threshold, return to step 1).
[0011] 3.2) Determine the degree of correlation of a signal between different time points based on the autocorrelation discrimination method. If the autocorrelation is lower than the set threshold, proceed to the next step; if the autocorrelation is higher than the set threshold, return to step 1).
[0012] 4) A short-time energy-based discrimination method is used to determine whether the separated result is a speech signal or a noise signal;
[0013] 4.1) Obtain the discriminant formula, assuming a certain length is L. d The signal is x d (n), whose short-time energy is E nd (n), the short-time energy sum is S d S d The formula is
[0014]
[0015] E nd (n) is normalized to obtain Its formula is
[0016]
[0017] set up The sum of The formula is
[0018]
[0019] From formulas (4), (5) and (6), we can see that
[0020]
[0021] set up The sum of squares is Q d Its formula is
[0022]
[0023] 4.2) Assuming the two signals obtained after separation in step 2) are x(n) and y(n) respectively, extract a segment of length L from each segment. d The signal is used to obtain two truncated signal segments, which are then substituted into the formula. Get Q x and Q y ,like
[0024] Q x Q y (9)
[0025] Then we determine that x(n) is a speech signal and y(n) is a noise signal;
[0026] 5) The speaker inside the earmuff inputs the voice signal into the interior of the earmuff cavity to complete voice retention.
[0027] The discriminant method based on cross-correlation in step 3.1) is specifically as follows:
[0028] Assume that the two separated signals are x(n) and y(n), and the cross-correlation function R xy of the signals x(n) and y(n) is calculated by the formula
[0029]
[0030] In the formula, m represents the time offset;
[0031] Take the absolute mean value of the cross-correlation of x(n) and y(n) as h, and set the cross-correlation threshold as H. If h < H is satisfied, it indicates that the noise and voice signals are successfully separated.
[0032] The discriminant method based on autocorrelation in step 3.2) is specifically as follows:
[0033] Assume that one of the separated signals is x(n), and the autocorrelation function R xx of the signal x(n) is calculated by the formula
[0034]
[0035] In the formula, k represents the time offset;
[0036] Take the absolute mean value of the autocorrelation as z, and set the autocorrelation threshold as Z. If z < Z is satisfied, it indicates that the noise and voice signals are successfully separated.
[0037] The beneficial effects of the present invention are as follows:
[0038] 1. Based on the traditional noise-canceling earmuff, a convolutional blind source separation algorithm and a voice signal discriminant algorithm based on microphone input are introduced to complete the separation and retention of voice signals. While the earmuff reduces noise, it transmits useful voice signals into the interior of the earmuff, enabling the wearer to have direct communication while wearing the earmuff.
[0039] 2. The voice signal passes through the microphone of the earmuff itself and is transmitted into the interior of the earmuff after subsequent processing, without the need to be transmitted through other devices, making it more simple and independent to use.
[0040] 3. The convolutional blind source separation algorithm is used to separate the noise signal and the voice signal, reducing the noise directly introduced into the interior of the earmuff;
[0041] 4. The voice signal discriminant algorithm improves the accuracy of the voice signal transmitted into the interior of the earmuff. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a structural block diagram of a noise-canceling earmuff that preserves speech, as proposed in this invention.
[0044] Figure 2 This is a schematic diagram of a noise-canceling earmuff that preserves speech, as proposed in this invention.
[0045] Figure 3 This is a diagram of the earmuff system proposed in this invention;
[0046] Figure 4 This is a flowchart of the algorithm in an embodiment of the present invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] This invention provides a noise-canceling earmuff that preserves speech, such as Figure 1 As shown, the device includes earmuffs with a physical sound insulation structure. An external microphone sensor for collecting external sound signals is installed on the outside of the earmuffs. Inside the earmuffs are a microprocessor module, an audio driver circuit, a speaker, and a power supply module. The microphone sensor is connected to the microprocessor module via an ADC interface. The microprocessor module is connected to the audio driver circuit via a DAC interface. The audio driver circuit is connected to the speaker via a filter and an amplifier.
[0049] Noise cancellation is achieved by passive noise-isolating earmuffs, which eliminate most noise through sound-absorbing structures and absorbent materials.
[0050] The process of speech preservation is as follows: a microphone collects sound signals from the environment, including noise signals and speech signals. A convolutional blind source separation algorithm is used to process the collected signals to separate the environmental noise signals and speech signals; a speech signal discrimination algorithm is used to identify the speech signals, and the speech signals are transmitted into the earcup cavity through a speaker driven by an audio driver circuit.
[0051] like Figure 2As shown, the passive sound-insulating earmuffs serve as the main structure, including a sound-insulating cavity and sound-absorbing material. Microphones are arranged on the outer side of the earmuffs to collect external sound signals, and the sound signals include two types of signals: noise and speech. Two microphones are respectively connected to the ADC interface, and the ADC interface is connected to the microprocessor module as the input of the microprocessor. The microprocessor module performs algorithm processing on the input signal, and the microprocessor module outputs a signal through the DAC interface. The DAC interface is connected to the audio driving circuit, and the audio driving circuit is connected to the speaker, and the speaker is arranged in the earmuff cavity.
[0052] As Figure 3 shown, the system includes: Physical structure: Physical sound-insulating earmuffs, ADC and DAC interface modules, power supply module, audio driving circuit, microphone group, speaker. Software algorithm composition: Convolutional blind source separation algorithm and speech signal discrimination algorithm.
[0053] The present invention also provides a noise reduction method for retaining speech. The complete process is as Figure 4 shown, including the following steps:
[0054] 1) Use earmuffs with a physical sound-insulating structure, and use the pick-up microphone sensor outside the earmuffs to collect external sound signals, including noise signals and speech signals;
[0055] 2) Use the convolutional blind source separation algorithm to separate the noise signal and the speech signal;
[0056] 3) Determine whether the noise signal and the speech signal are completely separated:
[0057] 3.1) Discrimination method based on cross-correlation
[0058] Cross-correlation can measure the similarity between two signals. The cross-correlation function R xy (m) of signals x(n) and y(n) is calculated as
[0059]
[0060] m represents the time offset.
[0061] Assume that the two separated signals are x(n) and y(n). In the case of normal separation, noise and speech can be completely separated, and at this time, the cross-correlation between x(n) and y(n) is small. When the separation effect is poor, noise and speech are not completely separated, and x(n) and y(n) contain similar signals, and the cross-correlation between the two is large.
[0062] Therefore, a cross-correlation threshold condition h < H can be set. When this condition is met, it indicates that the noise and speech signals have been successfully separated. Among them, h can take the absolute mean value of the cross-correlation between x(n) and y(n), and H is the set threshold.
[0063] Self - correlation - based discrimination method
[0064] Self - correlation represents the degree of correlation between different time points of a signal. The autocorrelation function \(R\) xx (k) of the signal \(x(n)\) is calculated as
[0065]
[0066] where \(k\) represents the time offset.
[0067] The results of blind source separation include three types: speech signals, noise signals, and mixed signals. Speech signals and noise signals are the results of normal separation, while mixed signals are obtained when the separation effect is not good. Speech signals usually do not have obvious periodicity, so their autocorrelation values are small. In contrast, the autocorrelation values of noise signals generated by repetitive mechanical movements are larger than those of speech signals.
[0068] Therefore, a threshold condition of autocorrelation \(z < Z\) can be set. When this condition is met, it indicates that the noise and speech signals are successfully separated, and the signal with the smallest autocorrelation value is likely to be the speech signal. Here, \(z\) can be taken as the absolute mean of the autocorrelation, and \(Z\) is the set threshold.
[0069] 4) Use the short - time energy - based discrimination method to determine whether the separated result is a speech signal or a noise signal;
[0070] The energy of a speech signal changes with time, and its short - time energy is calculated as the energy of each frame of the speech signal under a certain frame length.
[0071] The formula for the short - time energy of the signal \(x(n)\) is as follows
[0072]
[0073] where \(E(n)\) represents the short - time energy at time \(n\), \(\omega(m)\) is the window function, and \(N\) is the window length. Windowing is usually used to emphasize the sample values at the central moment and reduce the influence of sample values at the window boundaries.
[0074] To make a judgment using short - time energy, the short - time energy of the signal needs to be normalized as follows first.
[0075] Suppose there is a signal \(x\) d (n) with a length of \(L\) d , its short - time energy is \(E\) nd (n), and the sum of short - time energies is \(S\) d , and the formula for \(S\) d is
[0076]
[0077] E nd (n) is normalized to obtain Its formula is
[0078]
[0079] set up The sum of The formula is
[0080]
[0081] From formulas (4), (5) and (6), we can see that
[0082]
[0083] This completes the analysis of signal x. d Normalization of short-time energy of (n).
[0084] set up The sum of squares is Q d Its formula is
[0085]
[0086] From formula (7), we can see that and For a constant value, according to mathematical knowledge, if and only if At that time, Q d To obtain the minimum value. The closer Q approaches the average, d The smaller the Q; conversely, Q d The larger.
[0087] Speech signals typically have many interruptions, while noise is usually continuous. As shown in Figure (1), the short-time energy of a speech signal has many peaks and valleys, and the valleys are close to 0; while the short-time energy of a noise signal has fewer peaks and valleys, and the valleys are not 0.
[0088] Assuming the speech signal is x(n) and the noise signal is y(n), and both have the same length, the normalized short-time energy can be obtained from formula (5). and Q is obtained from formula (8) x and Q y Based on the characteristics of the peak and trough values of the two signals, it can be known that the speech signal... The difference is large, and the noise signal Relatively average. Therefore, Q x Closer to the maximum value, Q y Closer to the minimum value, Q x and Qy The following relationship exists
[0089] Q x Q y (9)
[0090] That is, the Q of the speech signal x Q greater than the noise signal y Based on this relationship, it can be determined whether the separated result is a speech signal or a noise signal.
[0091] 5) The speaker inside the earcup inputs the voice signal into the earcup cavity to complete voice retention.
[0092] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, for the device embodiments, the above descriptions are merely preferred embodiments of the present invention. Since they are fundamentally similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The above descriptions are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention, without departing from the principle of the present invention, should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A noise reduction method that preserves speech, characterized in that... It includes the following steps: 1) Adopt earcups with a physical sound insulation structure, and use the pickup microphone sensor outside the earcups to collect external sound signals, including noise signals and voice signals; 2) Use the convolutional blind source separation algorithm to separate the noise signal and the voice signal; 3) Judge whether the noise signal and the voice signal are completely separated: 3.1) Use the discrimination method based on cross-correlation to judge the similarity between the two signals. If the absolute mean value of the cross-correlation function of the two signals is lower than the set cross-correlation threshold, go to the next step. If the absolute mean value of the cross-correlation function of the two signals is higher than the set cross-correlation threshold, return to step 1); The discrimination method based on cross-correlation is as follows: Suppose the two separated signals are x(n) and y(n), and the cross-correlation function R of signals x(n) and y(n) is... xy The formula for calculating (m) is: In the formula, m represents the time offset; Take the absolute mean value of the cross-correlation function of x(n) and y(n) as h, and set the cross-correlation threshold as H. If h < H is satisfied, it indicates that the noise signal and the voice signal are successfully separated; 3.2) Use the discrimination method based on autocorrelation to judge the correlation degree between a certain signal at different time points. If the absolute mean value of the autocorrelation function of the signal is lower than the set autocorrelation threshold, go to the next step. If the absolute mean value of the autocorrelation function of the signal is higher than the set autocorrelation threshold, return to step 1); The discrimination method based on autocorrelation is as follows: Suppose that a certain signal obtained by separation is x(n), and the autocorrelation function of signal x(n) is R. xx The formula for calculating (k) is: In the formula, k represents the time offset; Take the absolute mean value of the autocorrelation function of x(n) as z, and set the autocorrelation threshold as Z. If z < Z is satisfied, it indicates that the noise signal and the voice signal are successfully separated; 4) Use the discrimination method based on short-time energy to judge whether the separated result is a voice signal or a noise signal; 4.1) Obtain the discriminant formula, assuming a certain length is L. d The signal is x d (n), whose short-time energy is E nd (n), the short-time energy sum is S d S d The formula is E nd (n) is normalized to obtain Its formula is set up The sum of The formula is It can be seen from formulas (4), (5) and (6) that set up The sum of squares is Q d Its formula is 4.2) Assuming the two signals obtained after separation in step 2) are x(n) and y(n), respectively, extract a segment of length L from each segment. d The signal is used to obtain two truncated signal segments, which are then substituted into the formula. Get Q x and Q y ,like Q x >Q y (9) Then judge that x(n) is a voice signal and y(n) is a noise signal; 5) The speaker in the earcup inputs the voice signal into the earcup cavity to complete voice retention.
Citation Information
Patent Citations
Active sound insulation earmuff with voice enhancement function
CN110856070A