Natural sound interference stripping algorithm and its application in noise monitoring
By employing a natural sound interference stripping algorithm, utilizing real-time fractional-order differential processing and an acoustic waveguide model, combined with a multi-band canceller, the problem of speech interference stripping in complex natural sound environments is solved, achieving efficient and accurate interference stripping results and improving the quality and robustness of speech signals.
Patent Information
- Application Number
- CN202511247344.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-03
AI Technical Summary
In complex natural sound environments, traditional speech interference removal techniques are unable to effectively distinguish and remove non-stationary, nonlinear, and time-varying natural sound interference, resulting in a large number of interference components remaining in the speech signal and limited improvement in speech quality.
A natural sound interference stripping algorithm is adopted, which uses real-time fractional derivative processing, acoustic waveguide model and phase analysis, combined with multi-band canceller, to reconstruct the time-domain interference waveform based on the acoustic impedance matching principle, so as to achieve efficient and accurate stripping of the target speech signal.
In complex natural acoustic environments, it significantly improves the accuracy and efficiency of interference stripping, generates pure target speech signals of higher quality, reduces signal distortion and interference, and enhances the spatial coherence and robustness of speech signals.
Smart Images

Figure CN120766705B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech analysis, and in particular to a natural sound interference stripping algorithm and its application in noise monitoring. BACKGROUND
[0002] In many fields such as voice communication, voice recognition, hearing aid technology and intelligent acoustic monitoring, accurately obtaining pure target voice signal is the core requirement to ensure system performance and user experience. However, in actual application scenarios, the natural sound environment is complex and changeable, and the target voice is inevitably affected by various natural sound interferences. These interference sounds and target voice are mixed with each other to form a complex acoustic signal, which greatly increases the difficulty of separating the pure target voice from the mixed signal.
[0003] Traditional voice enhancement and interference stripping techniques are mainly based on linear processing methods in frequency domain or time domain, such as spectral subtraction, adaptive filtering, etc. These methods can reduce noise interference to some extent, but have many limitations when dealing with complex natural sound interference. Natural sound interference has the characteristics of non-stationarity, non-linearity and time-varying, and its frequency spectrum has extensive overlap with the target voice. Simple linear processing cannot effectively distinguish and strip the interference signal, resulting in a large amount of interference components remaining in the processed voice signal, and limited voice quality improvement.
[0004] With the continuous development of acoustic theory and signal processing technology, some beamforming techniques based on microphone array are applied to the field of voice interference stripping. This kind of technology uses the spatial distribution characteristics of multiple microphones to enhance the voice signal in the target direction and suppress the interference signal in other directions through beamforming algorithm. However, the actual acoustic environment is complex, and sound waves will reflect, refract and diffract during propagation, making it difficult to accurately model the spatial characteristics of the sound field. In addition, factors such as the calibration accuracy of the microphone array, the array topology structure and the change of environmental acoustic parameters will significantly affect the performance of the beamforming technology, making it difficult to achieve ideal interference stripping effect in actual application.
[0005] Therefore, we propose a natural sound interference stripping algorithm and its application in noise monitoring to solve the above problems. SUMMARY
[0006] The present application provides a natural sound interference stripping algorithm and its application in noise monitoring, which is used to efficiently and accurately strip the interference components in the target voice signal in a complex natural sound environment.
[0007] The first aspect of the present application provides a natural sound interference stripping algorithm, comprising: collecting environmental sound signals to generate original mixed speech signals; performing real-time fractional differential processing on the original mixed speech signals, dynamically adjusting the differential order according to the speech fundamental frequency, and outputting a preprocessed signal; inputting the preprocessed signal into an acoustic waveguide model to generate a separated field containing speech guided wave modes and interference radiation modes through real-time phase analysis; extracting the interference radiation mode components in the separated field, and reconstructing the time-domain interference waveform based on the acoustic impedance matching principle; inputting the original mixed speech signals and the reconstructed interference waveform into a multi-frequency band canceller to output a pure target speech signal.
[0008] Optionally, in the first implementation manner of the first aspect of the present application, the method further comprises: performing time difference of arrival compensation on the original sound signals collected by each channel of the microphone array to generate a multi-channel signal; inputting the multi-channel signal into a phase coherence detector to extract a coherent component dominated by the target speech and output a phase synchronization enhanced signal; calculating spatial filtering parameters in real time based on the environmental sound impedance characteristics to generate a set of adaptive filtering coefficients matched with the current sound field; inputting the phase synchronization enhanced signal into a spatial filter and applying the set of adaptive filtering coefficients for multi-channel fusion to output the original mixed speech signal.
[0009] Optionally, in the second implementation manner of the first aspect of the present application, the method further comprises: performing time-domain zero-crossing analysis on the original mixed speech signal to generate real-time updated speech fundamental frequency parameters; inputting the speech fundamental frequency parameters into an order mapper to output a dynamic differential order control signal through a preset fundamental frequency-differential order relationship curve; inputting the original mixed speech signal into an analog differential circuit, applying the differential order control signal to adjust the circuit impedance characteristics in real time, and generating a primary differential signal; collecting environmental temperature and humidity parameters, generating a circuit compensation signal through a sound speed correction model, and performing amplitude calibration on the primary differential signal; superimposing the calibrated differential signal and the original signal to output a preprocessed signal.
[0010] Optionally, in the third implementation manner of the first aspect of the present application, the method further comprises: calculating the propagation constant of sound waves in the equivalent waveguide based on the environmental air pressure and temperature parameters to generate a set of waveguide physical structure parameters; inputting the preprocessed signal into a phase analysis array to separate the waveguide wave components and the radiation components through orthogonal phase detection and generate an initial modal separation signal; adjusting the waveguide boundary impedance parameters according to the target sound source position information to generate impedance-matched sound field constraint conditions; inputting the initial modal separation signal into an acoustic waveguide model, applying the set of waveguide physical structure parameters and the sound field constraint conditions, and constructing a spatial separation field of speech guided wave modes and interference radiation modes; collecting the energy distribution characteristics of the output spatial separation field, dynamically correcting the waveguide structure parameters, and generating an acoustic separation field.
[0011] Optionally, in a fourth implementation form of the first aspect of the present application, the boundary impedance parameter is set as , then: ; wherein, ; ; real part is the acoustic resistance; imaginary part is the acoustic reactance; is the imaginary unit; is the medium density; is the sound speed; k is the wave number; L is the waveguide length; is the impedance discontinuity position.
[0012] Optionally, in a fifth implementation form of the first aspect of the present application, it comprises: collecting the sound energy distribution of the interference radiation mode from the designated spatial region of the acoustically separated field to generate a radiation mode energy distribution map; calculating the acoustic impedance reference value of the target reconstruction frequency band in real time according to the environmental air density and sound speed to generate an impedance matching parameter set; inputting the radiation mode energy distribution map into the array of tunable acoustic resonators, driving the resonators to vibrate by applying the impedance matching parameter set, and generating a primary reconstructed interference waveform; performing phase correlation comparison on the primary reconstructed interference waveform and the original separated field radiation component to generate a waveform calibration signal; and performing amplitude normalization on the calibrated waveform through a voltage-controlled amplifier to output a time-domain interference waveform.
[0013] Optionally, in a sixth implementation form of the first aspect of the present application, it comprises: inputting the reconstructed interference waveform into an analog filter bank to separate it into 32 reference sub-signals of independent frequency bands through LC resonant circuits; calculating the dynamic gain coefficient set of each frequency band in real time based on the environmental temperature and humidity parameters and the acoustic impedance characteristics; inputting the original mixed voice signal and the reference sub-signals of each frequency band into a parallel voltage subtracter array, applying the corresponding gain coefficients to perform cancellation operations, and generating a primary screening voice sub-signal; performing phase comparison on the primary screening voice sub-signal and the original separated field voice guided wave component to generate phase calibration parameters of each frequency band; and performing impedance matching on the calibrated voice sub-signal to output a pure target voice signal through acoustic-electric conversion.
[0014] Optionally, in a seventh implementation form of the first aspect of the present application, it further comprises converting the residual interference sound energy in the acoustically separated field into electrical energy while suppressing the mechanical vibration of the device, monitoring environmental mutations and triggering system reset, and synchronously outputting the pure voice to a dual-mode interface of bone conduction and air conduction: capturing residual sound energy from the radiation mode region of the acoustically separated field and converting it into electrical energy through a piezoelectric transducer to generate a system power compensation signal; collecting the device shell vibration signal to generate an anti-phase vibration waveform to drive a linear actuator and output a mechanical stability compensation force; monitoring the acoustic impedance mutation gradient in real time, generating a warning signal and triggering the processing parameter reset when the gradient exceeds a threshold; and synchronously converting the pure target voice signal into a bone conduction vibration signal and an air conduction sound signal to generate a dual-mode output.
[0015] The second aspect of the application provides an application of natural sound interference stripping in noise monitoring, comprising a natural sound interference stripping device: an acquisition module for collecting environmental sound signals to generate original mixed voice signals; a processing module for performing real-time fractional order differential processing on the original mixed voice signals, dynamically adjusting the differential order according to the voice fundamental frequency, and outputting a pretreatment signal; a setting module for inputting the pretreatment signal into an acoustic waveguide model to generate a separation field containing voice guide mode and interference radiation mode through real-time phase analysis; a reconstruction module for extracting the interference radiation mode component in the separation field and reconstructing a time-domain interference waveform based on the acoustic impedance matching principle; and a distribution module for inputting the original mixed voice signal and the reconstructed interference waveform into a multi-frequency band canceller to output a pure target voice signal.
[0016] The mechanism of the application is as follows: breaking through the limitations of traditional digital signal processing, the sound wave separation is realized by directly manipulating the physical field, and the significant advantages of ultra-low delay, high robustness and energy efficiency self-optimization are exhibited in complex natural environment.
[0017] Beneficial effects: based on the real-time calculation of spatial filtering parameters according to the environmental acoustic impedance characteristics, the adaptive filter coefficient set is generated, so that the spatial filter can better adapt to the current sound field environment, and the accuracy and effectiveness of multi-channel signal fusion are improved; through this signal fusion mode, the original mixed voice signal generated has higher spatial coherence, reduces signal distortion and interference caused by environmental factors, provides a better basic signal for subsequent processing, and helps to improve the interference stripping performance of the whole system.
[0018] The dynamic fractional order differential processing can adjust the processing mode according to the real-time characteristics of the voice signal, strengthen the resonance characteristics of the voice signal, and make the target voice and interference signal more obviously distinguished in characteristics, and the compensation of environmental temperature and humidity parameters further improves the accuracy of the pretreatment signal, providing more favorable conditions for subsequent modal separation and interference stripping.
[0019] Through the acoustic waveguide model and real-time phase analysis, different modes in the sound field can be accurately separated, providing a clear target signal for subsequent extraction of the interference radiation mode component and reconstruction of the interference waveform, greatly improving the precision and efficiency of interference stripping.
[0020] The time-domain interference waveform highly consistent with the actual interference signal is generated, providing an accurate reference signal for multi-frequency band cancellation operation, so that the interference components in the original mixed voice signal are more effectively removed, and the quality of the pure target voice signal is improved.
[0021] Multi-band processing can more precisely cancel interference signals in different frequency bands, improving the comprehensiveness and accuracy of interference stripping. Phase comparison and calibration, as well as impedance matching synthesis, further optimize the quality of the speech signal, making the output pure target speech signal clearer and more natural. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of one embodiment of the natural sound interference stripping algorithm in this invention;
[0023] Figure 2 This is a schematic diagram of another embodiment of the natural sound interference stripping algorithm in this invention;
[0024] Figure 3 This is a schematic diagram of one embodiment of the natural sound interference stripping device in this invention;
[0025] Figure 4 This is a schematic diagram of one embodiment of the natural sound interference stripping device in this invention. Detailed Implementation
[0026] This invention provides a natural sound interference removal algorithm and its application in noise monitoring, enabling efficient and accurate removal of interference components from target speech signals in complex natural sound environments. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the natural sound interference stripping algorithm in this invention includes:
[0028] 101. Multi-channel speech signal acquisition and fusion: Ambient sound signals are synchronously acquired through a spatially distributed microphone array and fused to generate a spatially coherent original mixed speech signal.
[0029] It can be understood that the subject performing the present application can be a natural sound interference stripping device, but also a terminal or a server, and the specific place is not limited. The server is taken as an example for description in the embodiment of the present application.
[0030] It should be noted that the array structure: a diameter of 20 cm Fibonacci ball array is adopted, containing 7 microphones (odd layout). The microphone position is generated by the Fibonacci spherical algorithm, and the coordinates are as follows:
[0031] Microphone No. Azimuth (°) Elevation (°) Relative coordinate (cm) 1 0.0 0.0 (0,0,10) 2 137.5 47.8 (6.2,4.1,7.3) 3 275.0 25.6 (-5.8,8.9,5.1) ... ... ... ...
[0032] Hardware configuration: each microphone is connected to an independent 24-bit ADC module (sampling rate 16 kHz), and cross-injection synchronization technology is used to ensure that the multi-channel synchronization error is less than 10 μs.
[0033] Synchronization mechanism: 1. Clock distribution: the main controller sends a unified clock signal (frequency 256 x sampling rate) to all ADCs to trigger synchronous sampling. 2. Data buffering: each ADC module has a built-in double buffer (capacity 512 samples), which realizes continuous acquisition without interruption through ping-pong operation. 3. Time stamp alignment: a high-precision time stamp (precision 1 μs) is added to each frame of signal (frame length 20 ms), and the data integrity is checked by CRC.
[0034] Spatial coherence fusion, beamforming algorithm: time delay compensation: taking the microphone as the reference point, the sound wave time delay (microphone 2 time delay ≈ 0.12 ms) of the target sound source (azimuth angle 30°, elevation angle 15°, distance 2 m) to other microphones is calculated. Weighted summation: delay and sum beamforming (DSB) is used, and each channel signal is superimposed according to the weight after time shift (weight coefficient: central microphone 0.25, edge microphone 0.1), generating a single-channel mixed signal.
[0035] Performance indicators: signal-to-noise ratio improvement: in a conference room environment (background noise 45 dB), the target speech signal-to-noise ratio is improved by 5 dB (SNR before fusion = 8 dB, SNR after fusion = 13 dB). Distortion control: signal distortion < 0.8% (THD-N standard).
[0036] Mixed signal characteristics: the output signal contains target human voice (frequency range 300-3400 Hz) and residual environmental noise (main component is air conditioner low-frequency humming sound). Verification method: through cross-correlation function analysis of each channel signal, the spatial coherence coefficient is measured to be 0.92 (ideal value is 1), indicating high spatial consistency.
[0037] 102、Adaptive fractional order resonance enhancement, real-time fractional order differential processing is performed on the original mixed speech signal, the differential order a (0.8 ≤ a ≤ 1.5) is dynamically adjusted according to the speech fundamental frequency, and the preprocessed signal with enhanced resonance characteristics is output;
[0038] Note that the input signal: original mixed speech signal (sampling rate 16 kHz), contains the target voice (male speaker, fundamental frequency range 85-180 Hz) and background air conditioner noise (120 Hz low-frequency hum). The input signal-to-noise ratio is 8 dB. Hardware platform: the server is equipped with a dual-core processor, and the real-time processing frame length is 20 ms (320 sampling points per frame).
[0039] Core algorithm implementation, fractional differential algorithm: discrete implementation using Grünwald-Letnikov (G-L) definition. Perform fractional differential operation on each frame of signal: order dynamic adjustment: the fundamental frequency detection module tracks the speech fundamental frequency in real time (detects that the current frame fundamental frequency is 110 Hz), and maps the differential order according to the preset rule: low frequency band (fundamental frequency < 100 Hz): order α = 1.5 (enhance low frequency formant); medium frequency band (100 Hz ≤ fundamental frequency ≤ 150 Hz): order α = 1.2 (balance enhancement); high frequency band (fundamental frequency > 150 Hz): order α = 0.8 (suppress high frequency noise); real-time calculation: taking the current frame fundamental frequency 110 Hz as an example, mapping α = 1.2. Calculate the fractional differential through a sliding window with a length of 5, and the coefficients are [1, -0.48, 0.28, -0.16, 0.09] (based on G-L coefficient table).
[0040] Dynamic adjustment strategy, fundamental frequency tracking accuracy: use autocorrelation function + peak detection, error less than 2 Hz. When the fundamental frequency jumps from 110 Hz to 170 Hz (explosive sound in speech), the system will switch α from 1.2 to 0.8 in the next frame (within 20 ms). Order smooth transition: to avoid distortion caused by order jump, use linear interpolation transition (α from 1.2 to 0.8 is completed in 3 frames).
[0041] Resonance feature enhancement: in the processed signal, the first resonance peak (F1) of the target voice is enhanced by 6.5 dB, and the air conditioner noise at 120 Hz is attenuated by 4 dB. Signal-to-noise ratio improvement: the output signal-to-noise ratio is improved from 8 dB to 14.5 dB, and the speech intelligibility (STOI index) is improved from 0.68 to 0.93. Real-time performance: single-frame processing time is less than 5 ms, meeting the real-time requirement.
[0042] 103、Acoustic boundary layer separation field construction, input the preprocessed signal into the acoustic waveguide model, and generate a separation field containing speech guided wave mode and interference radiation mode through real-time phase analysis;
[0043] It should be noted that the environmental settings: 10m x 5m x 3m conference room (rectangular waveguide model), background noise is 120Hz air conditioner sound (sound pressure level 65dB), the target sound source is a male speaker 1.5m away from the array (fundamental frequency 110Hz, speech frequency band 300-3400Hz). Input signal: preprocessed multi-channel speech signal (sampling rate 16kHz, frame length 20ms), after step 102 resonance enhancement, the target speech first formant (500Hz) energy is increased by 6.5dB, and the noise suppression is 4dB. Hardware platform: the server is equipped with a spherical array (diameter 20cm, containing 32 measurement points), the array center is 1.2m away from the ground, and the multi-channel phase synchronous data is transmitted in real time.
[0044] Waveguide model construction and modal decomposition, model parameters: the conference room is abstracted as a rigid boundary rectangular waveguide, the height direction (z axis) is 3m, and the sound speed is 340m / s. Calculate the cutoff frequency: the first-order normal wave cutoff frequency is 57Hz ( , ), so the 0-3 order normal wave covers 0-228Hz, the target speech fundamental frequency (110Hz) belongs to the first-order mode, and the air conditioner noise (120Hz) belongs to the second-order mode.
[0045] Real-time phase analysis: short-time Fourier transform (STFT) is performed on 32 channels of signals synchronously, the window function length is 256 points (16ms), and the step length is 10ms. Extract the phase gradient of each frequency point: at 500Hz, the target speech phase difference is ≤15° (inter-channel), and the air conditioner noise phase difference is ≥60° due to the diffusion field.
[0046] Modal separation strategy, guided wave mode identification: low-order modes (0-2 order) correspond to speech propagation, and their phase changes conform to the waveguide axial propagation law (the first-order mode has a phase difference close to 0° along the z axis); through spherical harmonic function expansion (truncated order N=6), the low-order components with energy concentrated in the axial direction are separated out, and 95% of the target speech energy is preserved.
[0047] Interference radiation mode extraction: high-order modes (≥3 order) correspond to noise scattering, and the energy distribution is uniform (the coherence coefficient of 120Hz noise on the spherical array is only 0.3); construct a hemispherical equivalent source plane (radius 25cm), extract high-order spherical harmonic components, and the noise energy is attenuated by 85%.
[0048] Output separation field characteristics, separation field structure:
[0049] Modal type Frequency range Propagation characteristics Energy proportion (target speech) Speech guided wave mode 100-1500 Hz Axial propagation, phase coherence 92% Interference radiation mode 100-400 Hz Uniform in all directions, phase random ≤8%
[0050] Performance indicators: target speech distortion: ≤1.2% (based on pure speech); processing delay: 8ms / frame (satisfies real-time requirements).
[0051] 104. Physical equivalent interference reconstruction: Extract the interference radiation mode components in the separation field and reconstruct the time-domain interference waveform based on the acoustic impedance matching principle;
[0052] It should be noted that the input and scene configurations are as follows: Input signal: Acoustic separation field data from step 103, including voice guided wave mode (energy percentage 92%, frequency 100–1500Hz) and interference radiation mode (energy percentage ≤8%, frequency 100–400Hz). The interference source is air conditioning noise (center frequency 120Hz, sound pressure level 65dB). Environmental parameters: Conference room dimensions 10m×5m×3m, air density ρ=1.2kg / m³. 3 The speed of sound is c = 340 m / s, and the theoretical acoustic impedance is... Rayleigh.
[0053] Interference radiation mode extraction and component separation: Using the spherical harmonic coefficients generated in step 103 (truncation order N=6), the higher-order spherical harmonic components (≥3rd order) of the interference radiation mode are extracted. At 120Hz, the spatial coherence coefficient of the interference component is 0.3 and the energy attenuation is 85%.
[0054] Modal characteristics: The interference components exhibit uniform radiation in all directions with significant phase randomness (phase difference between channels ≥ 60°), which contrasts sharply with the axial propagation characteristics of the target speech.
[0055] Acoustic impedance matching reconstruction, impedance equivalence principle: Treating the interference radiation mode as an equivalent sound source, its reconstruction must satisfy the acoustic impedance matching condition: the equivalent impedance Z1 of the reconstructed interference must be equal to the impedance of the actual noise source. match( ≈ If mismatch ( This will result in 50% energy reflection.
[0056] Reconstruction steps: Impedance modeling: A virtual equivalent source is set at the air conditioner outlet (1.8m from the array), with an initial impedance of 600Rayl (simulating impedance mismatch); Parameter adjustment: The equivalent density of the virtual source is optimized iteratively. ) and speed of sound ( ),make Approaching 408Rayl. Final optimized value: =1.18kg / m 3 , =346m / s =408.3Rayl, mismatch rate <0.1%); Waveform generation: Input the optimized equivalent source parameters into the acoustic transmission equation and output the time-domain interference waveform (sampling rate 16kHz).
[0057] Output reconstruction accuracy:
[0058] Parameter Actual noise Reconstructed waveform Fundamental energy 120 Hz, 65 dB 120 Hz, 64.8 dB Harmonic distortion 2.1% 2.3% Correlation coefficient - 0.93
[0059] 105、Dynamic spectral subtraction output, input the original mixed speech signal and the reconstructed interference waveform into the multi-band canceller, output the pure target speech signal.
[0060] Note that the input signal: the original mixed speech signal: sampling rate 16 kHz, containing the target human voice (male speaker, fundamental frequency 110 Hz, energy concentrated in 300-1500 Hz) and residual air conditioning noise (120 Hz, sound pressure level 60 dB). The reconstructed interference waveform: the time-domain noise signal from step 104 (main peak 120 Hz, harmonic component 240 Hz / 360 Hz, correlation coefficient with actual noise 0.93). Hardware platform: the server is equipped with a multi-core DSP processor, real-time processing frame length 20 ms (320 sampling points), frame overlap 50%.
[0061] Multi-band canceller workflow, frequency band division and parallel processing: divide the signal into 5 sub-bands according to the critical frequency band of the human ear:
[0062] Frequency band index Frequency range (Hz) Processing strategy 1 0–200 Strong suppression (air conditioner noise fundamental frequency) 2 200–500 Moderate suppression (speech F1 transition zone) 3 500–1500 Weak suppression (speech formant) 4 1500–3000 Retained (speech intelligibility key) 5 3000–8000 Retained (consonant component)
[0063] Each sub-band is separated by an elliptical filter bank, with a group time delay difference of less than 0.5 ms to avoid phase distortion.
[0064] Dynamic spectral subtraction operation: perform independently on each sub-band: spectral alignment: synchronize the reconstructed interference waveform and the original signal frame by frame, and calculate the noise power spectrum of each frame (120 Hz noise energy 64.8 dB).
[0065] Adaptive spectral subtraction: low frequency band (0-500 Hz): use the over subtraction factor α=1.8 to completely suppress the air conditioning fundamental frequency and harmonics. High frequency band (>500 Hz): use α=1.2 to retain the energy of the speech formant.
[0066] Spectral floor control: set the spectral floor threshold β=0.05, set the negative value after subtraction to 5% of the noise energy, and suppress the music noise.
[0067] Sub-band recombination and output: superimpose the processed sub-band signals, eliminate the frequency band boundary distortion through phase consistency calibration, and output the full-band pure speech.
[0068] Output noise suppression effect:
[0069] Frequency band Noise energy before processing (dB) Residual energy after processing (dB) 120 Hz 60.0 ≤ 35.2 (attenuation 24.8 dB) 240 Hz 52.5 ≤36.1 300-1500 Hz - Speech energy retention rate ≥ 98%
[0070] Voice quality indicators: SNR improvement: input signal SNR = 14.5 dB → output SNR = 24.1 dB (9.6 dB improvement). Distortion control: voice waveform distortion ≤ 1.8% (based on pure voice), formant frequency shift < 2 Hz. Subjective listening: complete elimination of air conditioner humming, voice fullness (F1 / F2) retention 95%, consonant clarity lossless.
[0071] In the embodiment of the present application, the Fibonacci spherical array and the specific hardware configuration and synchronization mechanism are adopted, combined with the beam forming algorithm, to effectively improve the signal-to-noise ratio and control distortion in a complex environment, to provide high-quality signals for subsequent processing, and to solve the problems of multi-channel synchronization and spatial coherence fusion in traditional methods; real-time fractional differential processing and dynamic adjustment of differential order according to the voice fundamental frequency can accurately strengthen the resonance characteristics according to the voice characteristics, effectively improve the target voice energy and suppress noise, and improve the intelligibility of the voice; the acoustic waveguide model is constructed and modal separation is realized through real-time phase analysis, the conference room environment is abstracted as a rigid boundary rectangular waveguide for modal decomposition, and the phase difference is used to distinguish voice and noise modal, providing a new idea for acoustic signal processing and effectively separating target voice and interference noise; based on the principle of acoustic impedance matching, the time-domain interference waveform is reconstructed, high-precision reconstruction is realized through impedance modeling and parameter adjustment, accurate interference waveform is provided for dynamic spectral subtraction, and the accuracy of noise cancellation is improved; the multi-band canceller divides the sub-band according to the critical frequency band of the human ear and adopts different processing strategies, combines dynamic spectral subtraction operation and spectral bottom control, effectively suppresses noise in different frequency bands, and at the same time preserves the key information of the voice, improves the voice quality, and achieves a good balance in noise suppression and voice preservation.
[0072] Please refer to Figure 2 Another embodiment of the natural sound interference stripping algorithm in the embodiment of the present application includes:
[0073] 201, multi-channel voice signal acquisition and fusion, synchronously acquiring environmental sound signals through a spatially distributed microphone array, and fusing to generate a spatially coherent original mixed voice signal;
[0074] Specifically, multi-channel time difference compensation signal generation, time difference compensation is performed on the original sound signals collected by each channel of the microphone array to generate time-aligned multi-channel signals; phase coherence enhancement signal generation, the time-aligned multi-channel signals are input into a phase coherence detector to extract coherent components dominated by target voice, and a phase synchronization enhancement signal is output; spatial filter coefficient dynamic generation, real-time calculation of spatial filter parameters based on environmental sound impedance characteristics, generation of adaptive filter coefficient set matched with the current sound field; coherent fusion signal output, the phase synchronization enhancement signal is input into a spatial filter, and adaptive filter coefficient set is applied for multi-channel fusion, and a spatially coherent original mixed voice signal is output.
[0075] Note that the microphone array: 4 MEMS microphones (sensitivity -38 dB) are arranged in a linear array with a spacing of 5 cm and are deployed in the center of the conference table. Processor: STM32F429 core controller, supports double buffering mechanism, sampling rate 16 kHz (24-bit quantization precision).
[0076] Signal processing flow, multi-channel time difference compensation, input: 4-channel original speech signal (target sound source azimuth 45°, distance 1.2 m). TDOA calculation: estimate the time delay by the generalized cross-correlation (GCC-PHAT) algorithm, the time difference between microphones 1 and 2 is 42 μs (corresponding to a sound path difference of 1.44 cm). Output: time-aligned multi-channel signal (after time difference compensation, the correlation of each channel signal is improved to 0.93).
[0077] Phase coherence enhancement, phase detector: calculate the mutual power spectrum phase of the frequency sub-band (1 kHz band), generate a phase weight matrix. Component extraction: retain coherent components with a phase difference <15° (target speech proportion increased by 18 dB). Output: phase-synchronous enhanced signal (signal-to-noise ratio SNR improved from initial 5 dB to 12 dB).
[0078] Dynamic generation of spatial filtering coefficients, sound field parameters: calculate the sound speed 346 m / s based on the temperature and humidity sensor data (25℃, RH60%), acoustic impedance 415 Rayl. Filter design: use MVDR beamforming, constrain the main lobe to point to the 45° direction (side lobe suppression -20 dB). Coefficient set: generate a 4×4-dimensional weight matrix (main channel weight 0.82, secondary channel -0.31).
[0079] Coherent fusion output, spatial filtering: convolve the phase-enhanced signal with the weight matrix, the target speech energy proportion of the fused signal reaches 89%. Output: spatially coherent mixed speech (coherence coefficient 0.78, reverberation time reduced to 0.6 s).
[0080] Index:
[0081] Stage Input SNR Output SNR Speech intelligibility (STOI) After time difference compensation 5 dB 8 dB 0.65 After phase coherence enhancement 8 dB 12 dB 0.78 After spatial filtering fusion 12 dB 18 dB 0.92
[0082] Note: The measured data is based on the NOIZEUS database (vehicle noise background), and the PESQ score is improved from 2.1 to 3.4.
[0083] 202, adaptive fractional order resonance enhancement, real-time fractional order differential processing is performed on the original mixed speech signal, and the differential order α (0.8≤α≤1.5) is dynamically adjusted according to the speech fundamental frequency, and a preprocessed signal with enhanced resonance characteristics is output.
[0084] Specifically, the voice fundamental frequency feature extraction, the original mixed voice signal is analyzed by time domain zero-crossing, and the real-time updated voice fundamental frequency parameter is generated; the differential order control signal generation, the voice fundamental frequency parameter is input into the order mapper, and the dynamic differential order control signal is output through the preset fundamental frequency-differential order relationship curve; the analog fractional order differential processing, the original mixed voice signal is input into the analog differential circuit, the differential order control signal is applied to adjust the circuit impedance characteristics in real time, and the primary differential signal is generated; the environmental temperature and humidity compensation, the environmental temperature and humidity parameters are collected, the circuit compensation signal is generated through the sound speed correction model, and the amplitude of the primary differential signal is calibrated; the resonance enhancement signal output, the calibrated differential signal and the original signal are superimposed, and the preprocessed signal with enhanced resonance characteristics is output.
[0085] It should be noted that the signal acquisition end: high sensitivity MEMS microphone (frequency response 100 Hz-8 kHz), sampling rate 16 kHz; sensor: temperature and humidity sensor (accuracy ±0.5℃, ±3%RH); core processor: STM32F4 series (supports real-time floating-point operation); differential circuit: adjustable fractional order capacitor circuit (order adjustment range 0.8-1.5);
[0086] Signal processing flow, voice fundamental frequency feature extraction, input: original mixed voice (target human voice fundamental frequency range: male voice 85-180 Hz); time domain zero-crossing analysis: detect the number of zero-crossing points in 20 ms frame, and calculate the fundamental frequency. Example: 32 zero-crossing points are detected in the frame→fundamental frequency=160 Hz (sampling rate 16 kHz / frame length 320 points)
[0087] Differential order control signal generation, order mapping curve: preset relationship between fundamental frequency and differential order a (fundamental frequency↑→a↓):
[0088] Fundamental frequency (Hz) 50–100 100–150 150–200 α 1.5 1.2 0.9
[0089] Current adjustment: fundamental frequency 160 Hz→output a=0.95;
[0090] Analog fractional order differential processing, circuit adjustment: a=0.95→control voltage 1.8V drives phase shift circuit, adjusts RC parameters (R=15kΩ, C=10nF); output: primary differential signal (high frequency component enhancement 3dB, phase shift +42°);
[0091] Environmental temperature and humidity compensation, parameter acquisition: temperature 25℃, humidity 60%→sound speed 346m / s; amplitude calibration: look up table to get compensation coefficient K=1.05→differential signal amplitude x 1.05;
[0092] Resonance enhancement signal output, signal superposition: calibrated differential signal+original signal (mixed ratio 1:0.7);
[0093] Verification:
[0094] Index Before processing After processing Target speech signal-to-noise ratio 12 dB 20 dB Fundamental harmonic distortion rate 9% 4% Processing delay — 8 ms
[0095] Note: Data based on real conference room noise environment (air conditioner noise + keyboard click sound) test, PESQ score from 2.8 to 3.6.
[0096] 203、Acoustic boundary layer separation field construction, input the pretreated signal into the acoustic waveguide model, and generate a separation field containing the speech guided wave mode and the interference radiation mode through real-time phase analysis;
[0097] Specifically, waveguide structure parameter generation, based on environmental air pressure and temperature parameters, calculate the propagation constant of sound waves in the equivalent waveguide, generate the waveguide physical structure parameter set; multi-modal phase decomposition, input the pretreated signal into the phase analysis array, separate the waveguide component and the radiation component through orthogonal phase detection, generate the initial modal separation signal; boundary impedance matching, according to the target sound source position information, adjust the waveguide boundary impedance parameter, generate the impedance matched sound field constraint condition; physical separation field construction, input the initial modal separation signal into the acoustic waveguide model, apply the waveguide physical structure parameter set and the sound field constraint condition, construct the spatial separation field of the speech guided wave mode and the interference radiation mode; real-time modal feedback calibration, collect the energy distribution characteristics of the output separation field, dynamically correct the waveguide structure parameter, generate the optimized acoustic separation field.
[0098] It should be noted that the waveguide structure parameter generation, the environmental parameters: temperature 25℃, air pressure 101.3kPa, humidity 60%, calculate the sound speed m / s (T is Celsius).
[0099] Propagation constant calculation: the equivalent waveguide is a cylindrical air cavity with a diameter D=0.4m, the target speech main frequency Hz, wave number rad / m.
[0100] Output parameter set: waveguide length L=2.5m, cutoff frequency Hz (only supports fundamental mode propagation).
[0101] Multi-modal phase decomposition, input signal: pretreated signal (after resonance enhancement, SNR=20dB). Orthogonal phase detection: use a 4-channel phase analysis array to calculate the signal mutual power spectrum phase difference in a 1kHz subband. Guided wave component: phase difference (coherence>0.9). Radiation component: phase difference (coherence<0.2). Output: the energy proportion of the separated guided wave mode is 85%, and the radiation mode is 15%.
[0102] Boundary impedance matching, target sound source localization: azimuth 30°, distance 1.2m (based on microphone array TDOA estimation). Impedance parameters: boundary acoustic impedance , is the imaginary unit, where:
[0103] ; ;
[0104] = 1.18 kg / m3 (air density), = 1.5 m (impedance discontinuity location). It is calculated that .
[0105] Physical separation field construction, sound field constraint: guided wave mode region [0, ] is a traveling wave field (phase linearly varying), and the radiation mode region [ , L] is a standing wave field (sound pressure amplitude fluctuation > 8dB). Separation field output: at = 1.5 m, the guided wave component transmittance > 90%, and the radiation component reflectance > 85%.
[0106] Real-time modal feedback calibration, energy distribution monitoring: guided wave region sound pressure amplitude fluctuation , radiation region Pa. Parameter correction: when the temperature rises to 28℃, the sound speed increases to 349 m / s, dynamically adjust = 1.52 m and update .
[0107] Verification:
[0108] Index Before calibration After calibration Guided wave mode energy purity 78% 92% Radiation mode suppression ratio 12 dB 21 dB Processing delay 15 ms 8 ms
[0109] 204、Physical equivalent interference reconstruction, extract the interference radiation mode component in the separation field, and reconstruct the time-domain interference waveform based on the acoustic impedance matching principle;
[0110] Specifically, radiation mode energy extraction, collect the sound energy distribution of the interference radiation mode from the specified spatial region of the acoustic separation field, and generate a radiation mode energy distribution map; acoustic impedance matching parameter generation, real-time calculation of the acoustic impedance reference value of the target reconstruction frequency band according to the environmental air density and sound speed, and generation of an impedance matching parameter set; physical waveform synthesis, input the radiation mode energy distribution map into the tunable acoustic resonator array, drive the resonator to vibrate by applying the impedance matching parameter set, and generate a primary reconstructed interference waveform; time-domain characteristic calibration, compare the primary reconstructed interference waveform with the original separation field radiation component in terms of phase correlation, and generate a waveform calibration signal; equivalent interference waveform output, amplitude normalization of the calibrated waveform through a voltage-controlled amplifier, and output of a physically equivalent time-domain interference waveform.
[0111] Note that the acoustic separation field acquisition array: 8 MEMS pressure sensors (frequency response 50Hz-5kHz), deployed on the ceiling of the conference room, 20cm apart, covering the radiated mode area. Resonator array: 32 piezoelectric ceramic resonant units (resonant frequency width 80Hz-1.2kHz), integrated adjustable LC circuit. Calibration module: high-speed phase detector (resolution 0.5°), voltage-controlled amplifier (gain range 0.1-5 times).
[0112] Signal processing flow, radiated mode energy extraction, input: interference radiated mode components in the acoustic separation field (main frequency band 500-800Hz). Energy distribution map generation: collect sound pressure amplitude distribution at 500Hz frequency band, sensor 3 detects peak sound pressure 0.8Pa, sensor 5 detects 0.4Pa, generate spatial energy gradient map.
[0113] Acoustic impedance matching parameter generation, environmental parameters: temperature 25°C, humidity 60%→ sound speed 346m / s, air density 1.18kg / m 3 Impedance calculation: the acoustic impedance reference value of the target reconstruction frequency band 600Hz =412 Rayl (Rayl). Parameter set output: generate LC tuning parameters for 32 resonant units (600Hz unit: inductance L=12mH, capacitance C=5.8nF).
[0114] Physical waveform synthesis, resonator driving: input the energy distribution map into the resonator array, apply 8V driving voltage to the 600Hz unit, generate a sinusoidal waveform. Primary reconstructed waveform: the synthesized waveform has an amplitude of 1.2V at 600Hz, with a correlation of 0.85 with the original interference spectrum.
[0115] Time-domain characteristic calibration, phase comparison: detect the phase difference between the reconstructed waveform and the original radiated component (about 15°), generate an inverted calibration signal; amplitude adjustment: increase the gain of the 600Hz component to 1.3 times through the voltage-controlled amplifier, and the phase difference is reduced to 3°.
[0116] Equivalent interference waveform output, normalization processing: amplitude normalization to the range of ±1.5V for the full frequency band (limiting clipping distortion <2%); output: time-domain interference waveform signal-to-noise ratio 22dB, correlation with original interference waveform 0.93.
[0117] Verification:
[0118] Index Before calibration After calibration Interference waveform correlation 0.85 0.93 Fundamental band amplitude error 18% 5% Processing delay 12 ms 3 ms
[0119] 205、Dynamic spectral subtraction output, input the original mixed speech signal and the reconstructed interference waveform into the multi-frequency band canceller, output the pure target speech signal.
[0120] Specifically, the frequency band reference signal is generated, the reconstructed interference waveform is input into the analog filter bank, and the LC resonance circuit is separated into 32 independent frequency band reference sub-signals; the environment adaptive gain coefficient is generated, the dynamic gain coefficient set of each frequency band is calculated in real time based on the environmental temperature and humidity parameters and the acoustic impedance characteristics; the frequency band cancellation is performed, the original mixed speech signal and the frequency band reference sub-signals are input into the parallel voltage subtractor array, and the cancellation operation is performed by applying the corresponding gain coefficient to generate the preliminary screening speech sub-signal; the phase coherence calibration is performed, the preliminary screening speech sub-signal is compared with the original separated field speech guided wave component in phase to generate the phase calibration parameter of each frequency band; and the pure speech synthesis output is performed, the calibrated speech sub-signal is impedance matched, and the final pure target speech signal is output through the acoustoelectric conversion.
[0121] It should be noted that the filter bank: 32-channel LC resonance circuit (center frequency 80Hz-4kHz, Q value 0.8), inductance tolerance ±2%. The subtractor array: parallel operational amplifier (gain accuracy 0.1dB), supports ±5V dynamic range. Sensor: temperature and humidity sensor (temperature ±0.3℃, humidity ±2%RH). Phase calibrator: high-speed comparator (delay <0.5ms), reference speech guided wave component phase reference.
[0122] Signal processing flow, frequency band reference signal generation, input: reconstructed interference waveform (main noise frequency band 500-800Hz, amplitude 1.2V). LC filter bank: 500Hz frequency band: L=15mH, C=6.8nF, output reference sub-signal amplitude 0.8V. 3kHz frequency band: L=2mH, C=1.3nF, output reference sub-signal amplitude 0.3V. Output: 32 independent frequency band reference sub-signals (bandwidth 125Hz).
[0123] Environment adaptive gain coefficient generation, parameter acquisition: temperature 26℃, humidity 55%→sound speed 348m / s, acoustic impedance 418Rayl. Gain calculation: low frequency band (<1kHz): gain coefficient 0.92 (air absorption compensation). High frequency band (>2kHz): gain coefficient 1.05 (humidity-induced sound attenuation compensation).
[0124] Frequency band cancellation execution, input: original mixed speech signal (target speech energy ratio 45%, SNR 18dB). Voltage subtractor: 500Hz frequency band: mixed signal 1.5V-reference signal 0.8V×0.92→output 0.76V. 3kHz frequency band: mixed signal 0.9V-reference signal 0.3V×1.05→output 0.59V. Output: preliminary screening speech sub-signal (residual noise energy reduced by 12dB).
[0125] Phase coherent calibration, phase alignment: take the voice guided wave component in the acoustic isolation field as the reference (phase fluctuation <10°). Calibration operation: detect the phase deviation of 500Hz sub-signal 8° → generate time delay compensation 0.04ms. After calibration, the phase difference of each frequency band is ≤5°.
[0126] Pure speech synthesis output, impedance matching: 32 sub-signals are weighted and superimposed (low frequency weight 0.85, high frequency 0.75). Acoustic-electric conversion: output signal-to-noise ratio 28dB, speech intelligibility STOI 0.94.
[0127] Performance verification:
[0128] Index Before processing After processing Full-band residual noise -18 dB -42 dB Speech harmonic distortion rate 9% 2% Environmental mutation response delay — 20 ms
[0129] 206、Recycle the residual interference sound energy in the acoustic isolation field into electrical energy, while suppressing the mechanical vibration of the device, monitoring environmental mutations and triggering system reset, and synchronously outputting pure speech to the bone conduction and air conduction dual-mode interface;
[0130] Specifically, the residual sound energy is captured from the radiation mode area of the acoustic isolation field, converted into electrical energy by a piezoelectric transducer, and a system power compensation signal is generated; the device shell vibration signal is collected, a linear actuator is driven by a reverse vibration waveform, and a mechanical stability compensation force is output; the sound impedance mutation gradient is monitored in real time, and when the gradient exceeds the threshold, a warning signal is generated and the processing parameter reset is triggered; the pure target speech signal is synchronously converted into bone conduction vibration signal and air conduction sound signal, and a dual-mode output is generated.
[0131] It should be noted that the hardware configuration, sound energy recycling module: 8 piezoelectric ceramic transducer pieces (sensitivity 1.2V / Pa, frequency response 80-500Hz) are arranged in the radiation mode area of the acoustic isolation field. Vibration suppression module: 3-axis MEMS accelerometer (range ±8g), linear actuator (response frequency 0-200Hz, thrust ±5N). Environmental monitoring module: acoustic impedance sensor (accuracy ±0.5 Rayl), temperature and humidity sensor (±0.3℃, ±2%RH). Output interface: bone conduction vibrator (resonant frequency 250Hz), air conduction loudspeaker (frequency response 100Hz-10kHz).
[0132] Processing flow, residual sound energy recycling, input: residual sound energy in the radiation mode area (main frequency 300Hz, sound pressure 0.6Pa). Piezoelectric conversion: 8 transducer pieces are connected in parallel, peak voltage 4.2V, conversion efficiency 18%. Output: system power compensation signal (stable power supply +5V / 200mA).
[0133] Mechanical vibration suppression, vibration collection: device shell vibration (main frequency 120 Hz, amplitude 0.4g). Anti-phase waveform generation: generate anti-phase vibration waveform with phase difference 180°, amplitude 0.4g. Active cancellation: linear actuator output compensation force 3.2N, vibration energy reduction 15dB.
[0134] Environmental mutation monitoring and resetting, monitoring parameters: sound impedance gradient threshold is set to 10 Rayl / s (room temperature 25℃ reference). Mutation response: air conditioner opening causes sound impedance gradient to rise to 15 Rayl / s→trigger parameter reset (waveguide structure, filter coefficient full refresh). Delay: response time 8ms is monitored.
[0135] Dual-mode voice output, bone conduction output: pure voice signal drives bone conduction vibrator (gain 0.8, delay <2ms), temporal bone conduction sound pressure level 60dB. Air conduction output: air conduction speaker output (frequency band compression to 300-3500Hz, signal-to-noise ratio 28dB).
[0136] Output:
[0137] Function Before processing After processing Residual sound energy utilization rate — 18% Device vibration amplitude 0.4g 0.05g Environmental mutation response delay — 8 ms Bone conduction output intelligibility (STOI) — 0.89
[0138] In the embodiment of the application, the spatially distributed microphone array synchronously collects environmental sound signals, and through multi-channel time difference compensation, phase coherence enhancement, and dynamic generation of spatial filter coefficients, an original mixed voice signal with spatial coherence is fused and generated, spatial information is fully utilized, the correlation of the signal and the extraction accuracy of the target voice are effectively improved, compared with the traditional single-channel processing method, the voice collection demand in a complex environment can be better adapted to; the original mixed voice signal is subjected to real-time fractional order differential processing, the differential order α is dynamically adjusted according to the voice fundamental frequency, adaptive processing can be performed according to the real-time characteristics of the voice signal, the resonance characteristics of the voice are effectively enhanced, and the target voice information is highlighted; the preprocessed signal is input into an acoustic waveguide model, a separation field containing a voice guided wave mode and an interference radiation mode is generated through real-time phase analysis, effective separation of the voice signal and the interference signal in space is realized, and a more accurate signal basis is provided for subsequent interference suppression;
[0139] Extracting the interference radiation modal component in the separation field, reconstructing the time-domain interference waveform based on the acoustic impedance matching principle can accurately simulate the interference signal in the environment, provide a more accurate reference signal for dynamic spectrum subtraction, and more effectively suppress interference; inputting the original mixed speech signal and the reconstructed interference waveform into the multi-frequency band canceller, through the steps of generating a frequency band reference signal, generating an environment adaptive gain coefficient, frequency band cancellation, and phase coherence calibration, precise cancellation of interference signals in different frequency bands is realized, and a pure target speech signal is output; capturing residual sound energy from the radiation modal area of the acoustic separation field, converting it into electrical energy through a piezoelectric transducer to generate a system power compensation signal, realizing energy recycling and reuse, improving the energy utilization efficiency of the system, and reducing the dependence on external power supply;
[0140] Collecting device shell vibration signals, generating reverse vibration waveforms to drive linear actuators, and outputting mechanical stability compensation forces can effectively suppress device mechanical vibration, improve device stability and reliability, and reduce the impact of vibration on speech signal processing; real-time monitoring of acoustic impedance mutation gradient, generating a warning signal and triggering processing parameter reset when the gradient exceeds the threshold, enabling the system to quickly adapt to environmental changes and ensuring stable operation in different environments; synchronously converting the pure target speech signal into bone conduction vibration signals and air conduction sound signals to generate dual-mode output, meeting the use requirements of different users and improving the flexibility and applicability of speech output. At the same time, the bone conduction output clarity reaches 0.89, ensuring the quality of the output speech.
[0141] The natural sound interference stripping algorithm in the embodiment of the application is described above, and the natural sound interference stripping device in the embodiment of the application is described below. Please refer to Figure 3 The natural sound interference stripping device in the embodiment of the application includes: an acquisition module 301 for collecting environmental sound signals and fusing to generate an original mixed speech signal; a processing module 302 for performing real-time fractional order differentiation processing on the original mixed speech signal, dynamically adjusting the differentiation order a according to the speech fundamental frequency, and outputting a preprocessed signal; a setting module 303 for inputting the preprocessed signal into an acoustic waveguide model to generate a separation field containing speech guided wave modes and interference radiation modes through real-time phase analysis; a reconstruction module 304 for extracting the interference radiation modal component in the separation field and reconstructing a time-domain interference waveform based on the acoustic impedance matching principle; and a distribution module 305 for inputting the original mixed speech signal and the reconstructed interference waveform into a multi-frequency band canceller and outputting a pure target speech signal.
[0142] In the embodiment of the present application, the processing module adopts real-time fractional differential to process the original mixed speech signal, and can dynamically adjust the differential order according to the speech fundamental frequency. Compared with the traditional fixed parameter processing, it can more accurately adapt to different speech characteristics, effectively improve the preprocessing effect of the original signal, and lay a good foundation for subsequent separation of interference and extraction of target speech; the setting module inputs the preprocessed signal into the acoustic waveguide model, generates a separation field containing speech guided wave mode and interference radiation mode through real-time phase analysis, and can clearly distinguish the mode of speech and interference, providing a key basis for accurate stripping of interference; the reconstruction module extracts the interference radiation mode component in the separation field, and reconstructs the time-domain interference waveform based on the acoustic impedance matching principle. The application of the acoustic impedance matching principle makes the reconstructed interference waveform closer to the actual interference situation, improving the accuracy and scientificity of the interference waveform reconstruction; the distribution module inputs the original mixed speech signal and the reconstructed interference waveform into the multi-band canceller, and outputs the pure target speech signal. The use of the multi-band canceller can efficiently cancel the interference of different frequency bands, effectively improve the purity of the output target speech signal, and obtain high-quality target speech in complex acoustic environment.
[0143] The above Figure 3 The natural sound interference stripping device in the embodiment of the present application is described in detail from the perspective of modular functional entities. The natural sound interference stripping device in the embodiment of the present application is described in detail from the perspective of hardware processing.
[0144] Figure 4 Fig. 1 is a structural schematic diagram of a natural sound interference stripping device provided by the embodiment of the present application. The natural sound interference stripping device 400 can have great differences due to different configurations or performances. The natural sound interference stripping device 400 includes a transmitter 401, a receiver 402, and a processor 403. The processor 403 can also be a controller, Figure 4 indicated as "controller / processor 403" in the figure. Optionally, the natural sound interference stripping device 400 can also include a modem processor 405. The modem processor 405 can include an encoder 406, a modulator 407, a decoder 408, and a demodulator 409.
[0145] In one example, the transmitter 401 conditions (e.g., analog converts, filters, amplifies, and upconverts, etc.) the output samples and generates an uplink signal, which is transmitted via an antenna to an access network device. On the downlink, the antenna receives the downlink signal transmitted by the access network device. The receiver 402 conditions (e.g., filters, amplifies, downconverts, and digitizes, etc.) the signal received from the antenna and provides input samples. In the modem processor 405, the encoder 406 receives traffic data and signaling messages to be sent on the uplink and processes (e.g., formats, encodes, and interleaves, etc.) the traffic data and signaling messages. The modulator 407 further processes (e.g., symbol maps and modulates) the encoded traffic data and signaling messages and provides output samples. The demodulator 409 processes (e.g., demodulates) the input samples and provides symbol estimates. The decoder 408 processes (e.g., deinterleaves and decodes) the symbol estimates and provides decoded data and signaling messages to the natural sound interference cancellation device 400. The encoder 406, the modulator 407, the demodulator 409, and the decoder 408 can be implemented by a composite modem processor 405. These units process in accordance with the radio access technology employed by the wireless access network (e.g., the access technology of LTE and other evolved systems). Note that when the natural sound interference cancellation device 400 does not include the modem processor 405, the above functions of the modem processor 405 can also be completed by the processor 403.
[0146] The processor 403 controls and manages the actions of the natural sound interference cancellation device 400, for performing the processing procedures performed by the natural sound interference cancellation device 400 in the above embodiments of the present disclosure. For example, the processor 403 is further configured to perform each step of the transmitting device or the receiving device in the above algorithm embodiments, and / or other steps of the technical solutions described in the embodiments of the present disclosure.
[0147] Further, the natural sound interference cancellation device 400 can further include a memory 404, which is configured to store program codes and data for the natural sound interference cancellation device 400.
[0148] It can be understood that, Figure 4 Only a simplified design of the natural sound interference cancellation device 400 is shown. In actual applications, the natural sound interference cancellation device 400 can include any number of transmitters, receivers, processors, modem processors, memories, etc., and all devices that can implement the embodiments of the present disclosure are within the protection scope of the embodiments of the present disclosure.
[0149] The application further provides a natural sound interference stripping device, comprising a memory and a processor, the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to make the processor execute the steps of the natural sound interference stripping algorithm in each of the embodiments.
[0150] The application further provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores instructions, and the instructions make a computer execute the steps of the natural sound interference stripping algorithm when the instructions are run on the computer.
[0151] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing algorithm embodiments, which will not be described here.
[0152] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the algorithm described in each embodiment of the application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0153] The above-described embodiments are only used to illustrate the technical solutions of the application, rather than limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the application.
Claims
1. A natural sound interference peeling algorithm, characterized by, The natural sound interference stripping algorithm comprises: Collecting ambient sound signals, and fusing to generate an original mixed voice signal; Performing real-time fractional order differentiation processing on the original mixed voice signal, dynamically adjusting a differentiation order according to a voice fundamental frequency, and outputting a preprocessed signal; Inputting the preprocessed signal into an acoustic waveguide model, and generating a separation field containing a voice guided wave mode and an interference radiation mode through real-time phase analysis; Extracting an interference radiation mode component in the separation field, and reconstructing a time-domain interference waveform based on an acoustic impedance matching principle; Inputting the original mixed voice signal and the reconstructed interference waveform into a multi-frequency band canceller, and outputting a pure target voice signal.
2. The natural sound interference peeling algorithm according to claim 1, characterized in that, Comprise: Performing time difference of arrival compensation on original sound signals collected by each channel of a microphone array to generate a multi-channel signal; Inputting the multi-channel signal into a phase coherence detector to extract a coherent component dominated by target voice, and outputting a phase synchronization enhancement signal; Real-time calculation of spatial filtering parameters based on ambient sound impedance characteristics to generate an adaptive filtering coefficient set matched to the current sound field; Inputting the phase synchronization enhancement signal into a spatial filter to perform multi-channel fusion using the adaptive filtering coefficient set, and outputting an original mixed voice signal.
3. The natural sound interference peeling algorithm according to claim 1, characterized in that, Comprise: Performing time-domain zero-crossing analysis on an original mixed voice signal to generate real-time updated voice fundamental frequency parameters; Inputting the voice fundamental frequency parameters into an order mapper to output a dynamic differentiation order control signal through a preset fundamental frequency-differentiation order relationship curve; Inputting the original mixed voice signal into an analog differentiation circuit to real-time adjust circuit impedance characteristics using the differentiation order control signal to generate a primary differentiated signal; Collecting ambient temperature and humidity parameters, generating a circuit compensation signal through a sound speed correction model to calibrate the amplitude of the primary differentiated signal; Superimposing the calibrated differentiated signal and the original signal to output a preprocessed signal.
4. The natural sound interference peeling algorithm according to claim 1, characterized in that, Comprise: Based on ambient air pressure and temperature parameters, calculate the propagation constant of sound waves in the equivalent waveguide to generate a waveguide physical structure parameter set; Inputting the preprocessed signal into a phase analysis array to separate acoustic waveguide wave components and radiation components through orthogonal phase detection to generate an initial modal separation signal; Adjusting the waveguide boundary impedance parameters according to the target sound source position information to generate impedance-matched sound field constraint conditions; Inputting the initial modal separation signal into an acoustic waveguide model to construct a spatial separation field of voice guided wave modes and interference radiation modes using the waveguide physical structure parameter set and the sound field constraint conditions; Collecting the energy distribution characteristics of the output spatial separation field to dynamically correct the waveguide structure parameters to generate an acoustic separation field.
5. The natural sound interference peeling algorithm according to claim 4, characterized in that, Setting the boundary impedance parameter to then: ; wherein ; ; Real part is the acoustic resistance; imaginary part is the acoustic reactance; is the imaginary unit; is the medium density; is the sound velocity; k is the wave number; L is the waveguide length; is the impedance discontinuity location.
6. The natural sound interference peeling algorithm according to claim 4, characterized in that, Comprise: Collecting the sound energy distribution of the interference radiation mode from a specified spatial region of the acoustic separation field to generate a radiation mode energy distribution map; Real-time calculation of the acoustic impedance reference value of the target reconstruction frequency band based on ambient air density and sound speed to generate an impedance matching parameter set; Inputting the radiation mode energy distribution map into a tunable acoustic resonator array to drive the resonator to vibrate using the impedance matching parameter set to generate a primary reconstructed interference waveform; Comparing the primary reconstructed interference waveform with the original separation field radiation component in terms of phase correlation to generate a waveform calibration signal; The calibrated waveform is amplitude normalized by a voltage control amplifier, and a time domain interference waveform is output.
7. The natural sound interference peeling algorithm according to claim 6, characterized in that, The method comprises the following steps: The reconstructed interference waveform is input into an analog filter bank, and is separated into 32 reference sub-signals of independent frequency bands through an LC resonant circuit; Based on the environmental temperature and humidity parameters and the acoustic impedance characteristics, the dynamic gain coefficient set of each frequency band is calculated in real time; The original mixed speech signal and the reference sub-signals of each frequency band are input into a parallel voltage subtractor array, and the corresponding gain coefficients are applied to perform offset operation to generate a preliminary screening speech sub-signal; The preliminary screening speech sub-signal and the original separated field speech guided wave component are phase-compared to generate phase calibration parameters of each frequency band; The calibrated speech sub-signal is impedance-matched, and a pure target speech signal is output through acoustic-electric conversion.
8. The natural sound interference peeling algorithm according to claim 7, characterized in that, The method also comprises the following steps: Residual interference sound energy in the acoustic separation field is converted into electric energy, and mechanical vibration of the equipment is suppressed, environmental mutations are monitored, and the system is reset, and the pure speech is output to a bone conduction and air conduction dual-mode interface: Residual sound energy in the radiation mode area of the acoustic separation field is captured, converted into electric energy through a piezoelectric transducer, and a system power compensation signal is generated; A device shell vibration signal is collected, a reverse vibration waveform is generated to drive a linear actuator, and a mechanical stability compensation force is output; The acoustic impedance mutation gradient is monitored in real time, and a warning signal is generated and the processing parameters are reset when the gradient exceeds a threshold value; 9. Use of natural sound interference peeling in noise monitoring, characterized by The pure target speech signal is converted into a bone conduction vibration signal and an air conduction sound signal synchronously, and a dual-mode output is generated. The natural sound interference stripping device comprises: An acquisition module is configured to collect environmental sound signals and fuse to generate an original mixed speech signal; A processing module is configured to perform real-time fractional order differential processing on the original mixed speech signal, dynamically adjust the differential order α according to the speech fundamental frequency, and output a preprocessed signal; A setting module is configured to input the preprocessed signal into an acoustic waveguide model, generate a separation field containing a speech guided wave mode and an interference radiation mode through real-time phase analysis, and output a reconstructed interference waveform; A reconstruction module is configured to extract the interference radiation mode component in the separation field, and reconstruct a time domain interference waveform based on the acoustic impedance matching principle; A distribution module is configured to input the original mixed speech signal and the reconstructed interference waveform into a multi-frequency band canceller, and output a pure target speech signal.
Citation Information
Patent Citations
Signal separation method based on fractional wavelet transform
CN101655834A
Speech separation method and device, mobile terminal and computer readable storage medium
CN110808061A