Natural sound interference stripping algorithm and application thereof in noise monitoring
Through the natural sound interference stripping algorithm, using real-time fractional-order differential processing and acoustic waveguide model, combined with multi-band canceller and acoustic impedance matching principles, the problem of interference stripping in voice signals in complex natural sound environments is solved, achieving efficient and accurate interference stripping and improving voice signal quality.
Patent Information
- Application Number
- CN202511247344.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-03
AI Technical Summary
In complex natural sound environments, traditional speech interference stripping technology has difficulty in effectively distinguishing and stripping non-stationary and nonlinear natural sound interference, resulting in a large amount of interference components remaining in the speech signal and limited improvement in speech quality.
It adopts the natural acoustic interference stripping algorithm, realizes efficient separation of speech signals through real-time fractional-order differential processing, acoustic waveguide model and multi-band canceller, combined with the principle of acoustic impedance matching.
In complex natural sound environments, the accuracy and efficiency of interference stripping are significantly improved, generating high-quality pure target speech signals, reducing signal distortion and interference, and improving the clarity and intelligibility of speech signals.
Smart Images

Figure CN120766705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech analysis, and in particular to a natural sound interference stripping algorithm and its application in noise monitoring. Background Art
[0002] In numerous fields, including voice communication, speech recognition, hearing aid technology, and intelligent acoustic monitoring, accurately acquiring pure target speech signals is a core requirement for ensuring system performance and user experience. However, in real-world applications, the natural acoustic environment is complex and ever-changing, and the target speech is often inevitably affected by various natural sound interferences. These interferences mix with the target speech, forming a complex acoustic signal, greatly increasing the difficulty of separating the pure target speech from the mixed signal.
[0003] Traditional speech enhancement and interference removal technologies are primarily based on linear processing methods in the frequency or time domain, such as spectral subtraction and adaptive filtering. While these methods can reduce noise interference to a certain extent, they have numerous limitations when dealing with complex natural sound interference. Natural sound interference is non-stationary, nonlinear, and time-varying, and its spectral characteristics overlap extensively with the target speech. Simple linear processing cannot effectively distinguish and remove the interference signal, resulting in a significant amount of interference remaining in the processed speech signal, limiting speech quality improvement.
[0004] With the continuous development of acoustic theory and signal processing technology, some beamforming technologies based on microphone arrays have been applied to the field of speech interference removal. These technologies utilize the spatial distribution characteristics of multiple microphones and use beamforming algorithms to enhance the speech signal in the target direction and suppress interference signals from other directions. However, the actual acoustic environment is complex, and sound waves are subject to reflection, refraction, and diffraction during propagation, making it difficult to accurately model the spatial characteristics of the sound field. In addition, factors such as the calibration accuracy of the microphone array, the array topology, and variations in environmental acoustic parameters can significantly affect the performance of beamforming technology, making it difficult to achieve ideal interference removal results in practical applications.
[0005] Therefore, we propose a natural acoustic interference stripping algorithm and its application in noise monitoring to solve the above problems. Summary of the Invention
[0006] The present invention provides a natural sound interference stripping algorithm and its application in noise monitoring, which is used to efficiently and accurately strip interference components from a target speech signal in a complex natural sound environment.
[0007] The first aspect of the present invention provides a natural sound interference stripping algorithm, which includes: collecting ambient sound signals and fusing them to generate an original mixed voice signal; performing real-time fractional-order differential processing on the original mixed voice signal, dynamically adjusting the differential order α according to the voice fundamental frequency, and outputting a preprocessed signal; inputting the preprocessed signal into an acoustic waveguide model, and generating a separation field containing voice guided wave modes and interference radiation modes through real-time phase analysis; extracting the interference radiation mode component in the separation field, and reconstructing the time domain interference waveform based on the acoustic impedance matching principle; inputting the original mixed voice signal and the reconstructed interference waveform into a multi-band canceller, and outputting a pure target voice signal.
[0008] Optionally, in a first implementation method of the first aspect of the present invention, it includes: compensating for the arrival time difference of the original sound signal collected by each channel of the microphone array to generate a multi-channel signal; inputting the multi-channel signal into a phase coherent detector, extracting the coherent component dominated by the target voice, and outputting a phase-synchronized enhanced signal; calculating the spatial filtering parameters in real time based on the environmental acoustic impedance characteristics to generate an adaptive filter coefficient set that matches the current sound field; inputting the phase-synchronized enhanced signal into the spatial filter, applying the adaptive filter coefficient set to perform multi-channel fusion, and outputting the original mixed voice signal.
[0009] Optionally, in a second implementation method of the first aspect of the present invention, it includes: performing time domain zero crossing analysis on the original mixed voice signal to generate real-time updated voice fundamental frequency parameters; inputting the voice fundamental frequency parameters into the order mapper, and outputting a dynamic differential order control signal through a preset fundamental frequency-differential order relationship curve; inputting the original mixed voice signal into an analog differential circuit, applying the differential order control signal to adjust the circuit impedance characteristics in real time to generate a primary differential signal; collecting ambient temperature and humidity parameters, generating a circuit compensation signal through a sound speed correction model, and performing amplitude calibration on the primary differential signal; superimposing the calibrated differential signal with the original signal, and outputting a preprocessed signal.
[0010] Optionally, in a third implementation method of the first aspect of the present invention, it includes: calculating the propagation constant of the sound wave in the equivalent waveguide based on the ambient air pressure and temperature parameters, and generating a waveguide physical structure parameter set; inputting the preprocessed signal into the phase analysis array, separating the acoustic waveguide wave component and the radiation component through orthogonal phase detection, and generating an initial modal separation signal; adjusting the waveguide boundary impedance parameters according to the target sound source position information, and generating an impedance matching sound field constraint condition; inputting the initial modal separation signal into the acoustic waveguide model, applying the waveguide physical structure parameter set and the sound field constraint condition to construct a spatial separation field of the voice guided wave mode and the interference radiation mode; collecting the energy distribution characteristics of the output spatial separation field, dynamically correcting the waveguide structure parameters, and generating an acoustic separation field.
[0011] Optionally, in a fourth implementation of the first aspect of the present invention, the boundary impedance parameter is set to ,but: ;in, ; ; Real part is the acoustic impedance; imaginary part For sound resistance; is an imaginary unit; is the medium density; is the speed of sound; k is the wave number; L is the length of the waveguide; is the impedance discontinuity location.
[0012] Optionally, in a fifth implementation method of the first aspect of the present invention, it includes: collecting the acoustic energy distribution of the interference radiation mode from a specified spatial area of the acoustic separation field to generate a radiation mode energy distribution diagram; calculating the acoustic impedance reference value of the target reconstruction frequency band in real time according to the ambient air density and sound speed to generate an impedance matching parameter set; inputting the radiation mode energy distribution diagram into a tunable acoustic resonator array, applying the impedance matching parameter set to drive the resonator vibration to generate a primary reconstructed interference waveform; performing phase correlation comparison on the primary reconstructed interference waveform and the original separation field radiation component to generate a waveform calibration signal; performing amplitude normalization on the calibrated waveform through a voltage-controlled amplifier to output a time domain interference waveform.
[0013] Optionally, in a sixth implementation method of the first aspect of the present invention, it includes: inputting the reconstructed interference waveform into an analog filter group, and separating it into reference sub-signals of 32 independent frequency bands through an LC resonant circuit; calculating the dynamic gain coefficient set of each frequency band in real time based on the ambient temperature and humidity parameters and acoustic impedance characteristics; inputting the original mixed voice signal and the reference sub-signal of each frequency band into a parallel voltage subtractor array, applying the corresponding gain coefficient to perform an offset operation to generate a preliminary screening voice sub-signal; performing a phase comparison between the preliminary screening voice sub-signal and the original separation field voice guide wave component to generate a phase calibration parameter for each frequency band; performing impedance matching synthesis on the calibrated voice sub-signal, and outputting a pure target voice signal through acoustic-to-electrical conversion.
[0014] Optionally, in the seventh implementation method of the first aspect of the present invention, it also includes recovering the residual interference sound energy in the acoustic separation field and converting it into electrical energy, while suppressing the mechanical vibration of the equipment, monitoring sudden changes in the environment and triggering a system reset, and synchronously outputting the pure voice to the bone conduction and air conduction dual-modal interface: capturing the residual sound energy from the radiation mode area of the acoustic separation field, converting it into electrical energy through a piezoelectric transducer, and generating a system power supply compensation signal; collecting the vibration signal of the equipment shell, generating an inverse vibration waveform to drive the linear actuator, and outputting a mechanically stable compensation force; monitoring the acoustic impedance mutation gradient in real time, and generating a warning signal and triggering a processing parameter reset when the gradient exceeds a threshold; synchronously converting the pure target voice signal into a bone conduction vibration signal and an air conduction sound signal to generate a dual-modal output.
[0015] The second aspect of the present invention provides an application of natural sound interference stripping in noise monitoring, including a natural sound interference stripping device: an acquisition module for collecting environmental sound signals and fusing them to generate an original mixed voice signal; a processing module for performing real-time fractional-order differential processing on the original mixed voice signal, dynamically adjusting the differential order α according to the voice fundamental frequency, and outputting a preprocessed signal; a setting module for inputting the preprocessed signal into an acoustic waveguide model, and generating a separation field containing voice guided wave modes and interference radiation modes through real-time phase analysis; a reconstruction module for extracting the interference radiation mode components in the separation field, and reconstructing the time domain interference waveform based on the acoustic impedance matching principle; and a distribution module for inputting the original mixed voice signal and the reconstructed interference waveform into a multi-band canceller to output a pure target voice signal.
[0016] The mechanism of this invention is as follows: it breaks through the limitations of traditional digital signal processing and realizes sound wave separation through direct manipulation of physical fields, showing the significant advantages of ultra-low latency, high robustness and energy efficiency self-optimization in complex natural environments; Beneficial Effects: Based on the environmental acoustic impedance characteristics, spatial filter parameters are calculated in real time to generate an adaptive filter coefficient set, enabling the spatial filter to better adapt to the current sound field environment, improving the accuracy and effectiveness of multi-channel signal fusion. Through this signal fusion method, the generated original mixed speech signal has higher spatial coherence, reducing signal distortion and interference caused by environmental factors, providing a higher-quality basic signal for subsequent processing, and helping to improve the interference stripping performance of the entire system. Dynamic fractional-order differential processing can adjust the processing method based on the real-time characteristics of the speech signal, enhance the resonance characteristics of the speech signal, and more clearly distinguish the target speech from the interference signal. Compensation for ambient temperature and humidity parameters further improves the accuracy of the preprocessed signal, providing more favorable conditions for subsequent modal separation and interference stripping. Through the acoustic waveguide model and real-time phase analysis, different modes in the sound field can be accurately separated, providing a clear target signal for the subsequent extraction of interference radiation modal components and reconstruction of interference waveforms, greatly improving the accuracy and efficiency of interference stripping; Generates a time-domain interference waveform that is highly consistent with the actual interference signal, providing an accurate reference signal for multi-band cancellation operations, thereby more effectively removing interference components from the original mixed speech signal and improving the quality of the pure target speech signal; Multi-band processing can more meticulously offset interference signals in different frequency bands, improving the comprehensiveness and accuracy of interference stripping. Phase comparison and calibration as well as impedance matching synthesis further optimize the quality of voice signals, making the output pure target voice signals clearer and more natural. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic diagram of an embodiment of a natural sound interference stripping algorithm according to an embodiment of the present invention; Figure 2 Schematic diagram of another embodiment of the natural sound interference stripping algorithm in an embodiment of the present invention; Figure 3 A schematic diagram of an embodiment of a natural sound interference stripping device according to an embodiment of the present invention; Figure 4 Schematic diagram of an embodiment of a natural sound interference stripping device in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] An embodiment of the present invention provides a natural sound interference stripping algorithm and its application in noise monitoring, which is used to efficiently and accurately strip interference components from a target speech signal in a complex natural sound environment. The terms "first," "second," "third," "fourth," and so on (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products, or devices.
[0019] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 An embodiment of the natural sound interference stripping algorithm in the embodiment of the present invention includes: 101. Multi-channel speech signal acquisition and fusion: A spatially distributed microphone array synchronously acquires ambient sound signals and fuses them to generate an original mixed speech signal with spatial coherence. It is understandable that the subject of the present invention may be a natural sound interference stripping device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0020] It should be noted that the array structure uses a Fibonacci sphere array with a diameter of 20 cm, including 7 microphones (odd number layout). The microphone positions are generated using the Fibonacci sphere algorithm, and the coordinates are as follows: Microphone serial number Azimuth (°) Elevation angle (°) Relative coordinates (cm) 1 0.0 0.0 (0,0,10) 2 137.5 47.8 (6.2,4.1,7.3) 3 275.0 25.6 (-5.8,8.9,5.1) ... ... ... ... Hardware configuration: Each microphone is connected to an independent 24-bit ADC module (sampling rate 16kHz), and cross-injection synchronization technology is used to ensure multi-channel synchronization error <10μs.
[0021] Synchronization Mechanism: 1. Clock Distribution: The master controller sends a unified clock signal (frequency 256 × sampling rate) to all ADCs, triggering synchronized sampling. 2. Data Buffering: Each ADC module has a built-in double buffer (capacity 512 samples), which uses a ping-pong operation to ensure continuous, uninterrupted acquisition. 3. Timestamp Alignment: A high-precision timestamp (accuracy 1μs) is added to each signal frame (frame length 20ms), and data integrity is verified using a CRC.
[0022] Spatial coherence fusion and beamforming algorithms: Delay compensation: Using the microphone as the reference point, the delay from the target sound source (30° azimuth, 15° elevation, 2m distance) to the other microphones is calculated (microphone 2 delay ≈ 0.12ms). Weighted summation: Using delay-sum beamforming (DSB), each channel signal is time-shifted and weighted (weight coefficient: 0.25 for the center microphone, 0.1 for the edge microphones) to generate a single-channel mixed signal.
[0023] Performance Indicators: Signal-to-Noise Ratio Improvement: In a conference room environment (background noise 45dB), the target speech signal-to-noise ratio improved by 5dB (SNR before fusion = 8dB, SNR after fusion = 13dB). Distortion Control: Signal distortion <0.8% (THD-N standard).
[0024] Mixed Signal Characteristics: The output signal consists of the target human voice (frequency range 300–3400 Hz) and residual ambient noise (primarily the low-frequency hum of an air conditioner). Verification: Cross-correlation analysis of each channel signal measured a spatial coherence coefficient of 0.92 (the ideal value is 1), indicating high spatial coherence.
[0025] 102. Adaptive fractional-order resonance enhancement: performs real-time fractional-order differential processing on the original mixed speech signal, dynamically adjusts the differential order α (0.8≤α≤1.5) according to the speech fundamental frequency, and outputs a preprocessed signal with enhanced resonance characteristics; The input signal is a raw mixed speech signal (sampling rate 16kHz), consisting of the target human voice (male speaker, fundamental frequency range 85–180Hz) and background air conditioning noise (120Hz low-frequency hum). The input signal-to-noise ratio is 8dB. The hardware platform is a server equipped with a dual-core processor, with a real-time processing frame length of 20ms (320 samples per frame).
[0026] Core algorithm implementation: Fractional-order differentiation algorithm: This algorithm uses the discretization defined by Grünwald-Letnikov (GL) to implement fractional-order differentiation for each frame. Dynamic order adjustment: The fundamental frequency detection module tracks the speech fundamental frequency in real time (detecting that the current frame's fundamental frequency is 110 Hz) and maps the differentiation order according to preset rules: Low-frequency band (fundamental frequency < 100 Hz): order α = 1.5 (enhancing low-frequency formants); mid-frequency band (100 Hz ≤ fundamental frequency ≤ 150 Hz): order α = 1.2 (balancing and enhancing); high-frequency band (fundamental frequency > 150 Hz): order α = 0.8 (suppressing high-frequency noise). Real-time calculation: For the current frame with a fundamental frequency of 110 Hz, α = 1.2 is used. Fractional-order differentiation is calculated using a sliding window of length 5, with coefficients in the range [1, -0.48, 0.28, -0.16, 0.09] (based on the GL coefficient table).
[0027] Dynamic adjustment strategy, fundamental frequency tracking accuracy: using autocorrelation function + peak detection, with an error of less than 2Hz. When the fundamental frequency jumps from 110Hz to 170Hz (plosive sounds in speech), the system switches α from 1.2 to 0.8 in the next frame (within 20ms). Order transition smoothing: To avoid distortion caused by order jumps, linear interpolation is used (α gradually changes from 1.2 to 0.8 in three frames).
[0028] Resonance feature enhancement: In the processed signal, the first resonance peak of the target speech ( ) energy increased by 6.5dB, and air conditioning noise energy at 120Hz was attenuated by 4dB. Signal-to-noise ratio improvement: The output signal signal-to-noise ratio increased from 8dB to 14.5dB, and speech intelligibility (STOI index) improved from 0.68 to 0.93. Real-time performance: Single-frame processing takes less than 5ms, meeting real-time requirements.
[0029] 103. Acoustic boundary layer separation field construction: input the pre-processed signal into the acoustic waveguide model, and generate a separation field containing the voice guided wave mode and the interference radiation mode through real-time phase analysis; The environment used was a 10m×5m×3m conference room (rectangular waveguide model). The background noise consisted of 120Hz air conditioning noise (65dB sound pressure level). The target sound source was a male speaker (fundamental frequency 110Hz, voice band 300–3400Hz) located 1.5m from the array. The input signal consisted of a preprocessed multi-channel speech signal (sampling rate 16kHz, frame length 20ms). After resonance enhancement in step 102, the energy of the first formant (500Hz) of the target speech was increased by 6.5dB, and noise was suppressed by 4dB. The hardware platform consisted of a server equipped with a spherical array (20cm diameter, with 32 measurement points), with the center of the array 1.2m above the ground, transmitting multi-channel phase-synchronized data in real time.
[0030] Waveguide model construction and modal decomposition, model parameters: the conference room is abstracted as a rigid boundary rectangular waveguide, the height direction (z axis) is 3m, the sound speed is 340m / s. Calculation cutoff frequency: the first-order simple normal wave cutoff frequency is 57Hz ( , ), so the 0–3 order simple normal waves cover 0–228 Hz, the target speech fundamental frequency (110 Hz) belongs to the first order mode, and the air conditioning noise (120 Hz) belongs to the second order mode.
[0031] Real-time phase analysis: Short-time Fourier transform (STFT) is performed synchronously on 32 channels, using a 256-point (16ms) window and a 10ms step size. Phase gradients are extracted at each frequency point. At 500 Hz, the phase difference of the target speech is ≤15° (between channels), and the phase difference of the air-conditioning noise due to the diffuse field is ≥60°.
[0032] Modal separation strategy, waveguide mode identification: Low-order modes (0–2) correspond to speech propagation, and their phase variations conform to the law of axial waveguide propagation (the first-order mode has a phase difference of nearly 0° along the z-axis). Spherical harmonic expansion (truncation order N=6) is used to separate low-order components with energy concentrated in the axial direction, retaining 95% of the target speech energy.
[0033] Interference radiation mode extraction: High-order modes (≥3rd order) correspond to noise scattering, and the energy distribution is uniform (the coherence coefficient of 120Hz noise on the spherical array is only 0.3); a hemispherical equivalent source surface (radius 25cm) is constructed to extract high-order spherical harmonic components, and the noise energy is attenuated by 85%.
[0034] Output separation field characteristics, separation field structure: Modal Type Frequency range Propagation characteristics Energy ratio (target speech) Speech Guided Wave Mode 100–1500Hz Axial propagation, phase coherent 92% Interference radiation mode 100–400Hz Uniform in all directions, random in phase ≤8% Performance indicators: Target speech distortion: ≤1.2% (based on pure speech); processing delay: 8ms / frame (meets real-time requirements).
[0035] 104. Physical equivalent interference reconstruction: extract the interference radiation modal components in the separation field and reconstruct the time domain interference waveform based on the acoustic impedance matching principle; It should be noted that the input and scenario configurations are as follows: the input signal: the acoustic separation field data from step 103, including the voice guided wave mode (92% energy, frequency 100–1500 Hz) and the interference radiation mode (≤8% energy, frequency 100–400 Hz). The interference source is air conditioning noise (center frequency 120 Hz, sound pressure level 65 dB). Environmental parameters: conference room dimensions 10 m × 5 m × 3 m, air density ρ = 1.2 kg / m 3 , sound velocity c=340m / s, theoretical acoustic impedance Rayleigh (Rayl).
[0036] Interference radiation mode extraction and component separation: Use the spherical harmonic coefficients (truncation order N = 6) generated in step 103 to extract the high-order spherical harmonic components (≥ 3rd order) of the interference radiation mode. At 120 Hz, the spatial coherence coefficient of the interference component is 0.3, and the energy is attenuated by 85%.
[0037] Modal characteristics: The interference component exhibits uniform radiation in all directions and significant phase randomness (phase difference between channels ≥ 60°), which is in sharp contrast to the axial propagation characteristics of the target speech.
[0038] Acoustic impedance matching reconstruction, impedance equivalence principle: the interference radiation mode is regarded as an equivalent sound source, and its reconstruction must meet the acoustic impedance matching condition: the equivalent impedance Z1 of the reconstructed interference must be the same as the actual noise source impedance match( ≈ If there is a mismatch ( ), will result in 50% energy reflection.
[0039] Reconstruction steps: Impedance modeling: Set up a virtual equivalent source at the air-conditioning outlet (1.8m away from the array), and set the initial impedance to 600Rayl (to simulate impedance mismatch); Parameter adjustment: Optimize the equivalent density of the virtual source through iteration ( ) and the speed of sound ( ),make Approaching 408Rayl. Final optimized value: =1.18kg / m 3 , =346m / s( =408.3Rayl, mismatch rate <0.1%); Waveform generation: Input the optimized equivalent source parameters into the acoustic transmission equation and output the time domain interference waveform (sampling rate 16kHz).
[0040] Output reconstruction accuracy: parameter Actual noise Reconstructed waveform Main frequency energy 120Hz, 65dB 120Hz, 64.8dB Harmonic distortion 2.1% 2.3% Correlation coefficient - 0.93 105. Dynamic spectrum subtraction output, the original mixed speech signal and the reconstructed interference waveform are input into the multi-band canceller, and the pure target speech signal is output.
[0041] The input signal is: the original mixed speech signal (sampled at 16kHz), containing the target human voice (male speaker, fundamental frequency 110Hz, energy concentrated between 300–1500Hz) and residual air conditioning noise (120Hz, sound pressure level 60dB). The reconstructed interference waveform is: the time-domain noise signal from step 104 (main peak 120Hz, harmonic components 240Hz / 360Hz, correlation coefficient with actual noise 0.93). The hardware platform is: a server equipped with a multi-core DSP processor, real-time processing frame length 20ms (320 samples), with 50% frame overlap.
[0042] Multi-band canceller workflow, frequency band division and parallel processing: Divide the signal into 5 sub-bands according to the critical frequency band of the human ear: Frequency band index Frequency range (Hz) Treatment strategy 1 0–200 Strong suppression (main frequency of air conditioning noise) 2 200–500 Moderate suppression (speech F1 transition zone) 3 500–1500 Weak suppression (speech formants) 4 1500–3000 Retention (critical for speech clarity) 5 3000–8000 Retention (consonant component) Each sub-band is separated by an elliptic filter bank with a group delay difference of <0.5ms to avoid phase distortion.
[0043] Dynamic spectrum subtraction operation: performed independently for each subband: Spectrum alignment: synchronize the reconstructed interference waveform with the original signal and frame it, and calculate the noise power spectrum of each frame (noise energy 64.8dB at 120Hz).
[0044] Adaptive spectral subtraction: Low frequency band (0–500Hz): Use a subtraction factor of α = 1.8 to completely suppress the air conditioning fundamental frequency and harmonics. Mid- and high-frequency band (>500Hz): Use α = 1.2 to preserve the speech formant energy.
[0045] Spectral bottom control: Set the spectral bottom threshold β=0.05 and set the negative value after subtraction to 5% of the noise energy to suppress musical noise.
[0046] Sub-band reorganization and output: The processed sub-band signals are superimposed, and the frequency band boundary distortion is eliminated through phase consistency calibration to output full-band pure speech.
[0047] Output noise suppression effect: frequency band Noise energy before processing (dB) Residual energy after treatment (dB) 120Hz 60.0 ≤35.2 (attenuation 24.8dB) 240Hz 52.5 ≤36.1 300–1500Hz - Voice energy retention rate ≥98% Voice Quality: Signal-to-Noise Ratio Improvement: Input signal SNR = 14.5dB → Output SNR = 24.1dB (9.6dB improvement). Distortion Control: Voice waveform distortion ≤ 1.8% (based on pure speech), formant frequency shift < 2Hz. Subjective Listening Experience: Air conditioning hum is completely eliminated, voice richness (F1 / F2) is retained at 95%, and consonant clarity is intact.
[0048] In the embodiment of the present invention, a Fibonacci sphere array and a specific hardware configuration and synchronization mechanism are used, combined with a beamforming algorithm, to effectively improve the signal-to-noise ratio and control distortion in a complex environment, provide high-quality signals for subsequent processing, and solve the difficulties of traditional methods in multi-channel synchronization and spatial coherence fusion; real-time fractional-order differential processing and dynamic adjustment of the differential order according to the voice fundamental frequency can accurately enhance the resonance characteristics according to the voice characteristics, effectively enhance the target voice energy and suppress noise, and improve voice intelligibility; an acoustic waveguide model is constructed and modal separation is achieved through real-time phase analysis, abstracting the conference room environment into a rigid-boundary rectangular waveguide. Row mode decomposition uses phase differences to distinguish speech and noise modes, providing new ideas for acoustic signal processing and effectively separating target speech from interfering noise; reconstructing the time domain interference waveform based on the principle of acoustic impedance matching, achieving high-precision reconstruction through impedance modeling and parameter adjustment, providing an accurate interference waveform for dynamic spectrum subtraction and improving the accuracy of noise cancellation; the multi-band canceller divides the subbands according to the critical frequency band of the human ear and adopts different processing strategies, combined with dynamic spectrum subtraction operations and spectral floor control, effectively suppressing noise in different frequency bands while retaining key speech information, improving speech quality, and achieving a good balance between noise suppression and speech preservation.
[0049] See also Figure 2 Another embodiment of the natural sound interference stripping algorithm in the embodiment of the present invention includes: 201. Multi-channel speech signal acquisition and fusion: A spatially distributed microphone array synchronously acquires ambient sound signals and fuses them to generate an original mixed speech signal with spatial coherence. Specifically, multi-channel time difference compensation signal generation, the original sound signal collected by each channel of the microphone array is compensated for the arrival time difference, and a time-aligned multi-channel signal is generated; phase coherence enhancement signal generation, the time-aligned multi-channel signal is input into the phase coherence detector, the coherent component dominated by the target speech is extracted, and a phase-synchronized enhancement signal is output; spatial filter coefficients are dynamically generated, the spatial filter parameters are calculated in real time based on the environmental acoustic impedance characteristics, and an adaptive filter coefficient set that matches the current sound field is generated; coherent fusion signal output, the phase-synchronized enhancement signal is input into the spatial filter, the adaptive filter coefficient set is applied for multi-channel fusion, and the original mixed speech signal with spatial coherence is output.
[0050] The microphone array consists of four MEMS microphones (−38dB sensitivity) arranged in a linear array, spaced 5cm apart and positioned at the center of the conference table. The processor is an STM32F429 core controller with double buffering and a 16kHz sampling rate (24-bit quantization precision).
[0051] Signal processing flow, multi-channel time difference compensation. Input: 4-channel original speech signal (target sound source azimuth angle 45°, distance 1.2m). TDOA calculation: Time delay estimation using the generalized cross-correlation (GCC-PHAT) algorithm. The time difference between microphones 1 and 2 is 42μs (corresponding to a sound path difference of 1.44cm). Output: Time-aligned multi-channel signals (after time difference compensation, the correlation between each channel signal increases to 0.93).
[0052] Phase coherence enhancement: Phase detector: Calculates the cross-power spectrum phase of frequency subbands (1kHz band) and generates a phase weight matrix. Component extraction: Retains coherent components with a phase difference of less than 15° (increasing the target speech ratio by 18dB). Output: Phase-synchronized enhanced signal (SNR increased from 5dB to 12dB).
[0053] Spatial filter coefficients are dynamically generated. Sound field parameters: a sound velocity of 346 m / s and an acoustic impedance of 415 Rayleigh are calculated based on temperature and humidity sensor data (25°C, RH 60%). Filter design: MVDR beamforming is used, with the main lobe constrained to point at a 45° angle (sidelobe suppression −20 dB). Coefficient set: A 4×4 weight matrix is generated (weights 0.82 for the primary channel and −0.31 for the secondary channel).
[0054] Coherent fusion output, spatial filtering: The phase-enhanced signal is convolved with the weight matrix. The fused signal accounts for 89% of the target speech energy. Output: Spatially coherent mixed speech (coherence coefficient 0.78, reverberation time reduced to 0.6s).
[0055] index: stage Input SNR Output SNR Speech Intelligibility (STOI) After time difference compensation 5dB 8dB 0.65 After phase coherence enhancement 8dB 12dB 0.78 After spatial filtering fusion 12dB 18dB 0.92 Note: The measured data is based on the NOIZEUS database (vehicle noise background), and the PESQ score has increased from 2.1 to 3.4.
[0056] 202. Adaptive fractional-order resonance enhancement: performs real-time fractional-order differential processing on the original mixed speech signal, dynamically adjusts the differential order α (0.8≤α≤1.5) according to the speech fundamental frequency, and outputs a preprocessed signal with enhanced resonance characteristics; Specifically, the speech fundamental frequency feature is extracted, and the time domain zero crossing analysis is performed on the original mixed speech signal to generate real-time updated speech fundamental frequency parameters; the differential order control signal is generated, and the speech fundamental frequency parameters are input into the order mapper, and the dynamic differential order control signal is output through the preset fundamental frequency-differential order relationship curve; the analog fractional-order differential processing is performed, and the original mixed speech signal is input into the analog differential circuit, and the differential order control signal is applied to adjust the circuit impedance characteristics in real time to generate a primary differential signal; the ambient temperature and humidity compensation is performed, and the ambient temperature and humidity parameters are collected, and the circuit compensation signal is generated through the sound speed correction model, and the amplitude of the primary differential signal is calibrated; the resonance enhancement signal is output, and the calibrated differential signal is superimposed on the original signal to output a preprocessed signal with enhanced resonance characteristics.
[0057] It should be noted that the signal acquisition end is a high-sensitivity MEMS microphone (frequency response 100Hz–8kHz), with a sampling rate of 16kHz; the sensor is a temperature and humidity sensor (accuracy ±0.5°C, ±3%RH); the core processor is an STM32F4 series (supporting real-time floating-point operations); the differential circuit is an adjustable fractional-order capacitor circuit (order adjustment range 0.8–1.5); Signal processing flow, speech fundamental frequency feature extraction, input: original mixed speech (target human voice fundamental frequency range: male voice 85–180Hz); time domain zero-crossing analysis: detect the number of zero crossings in the signal within a 20ms frame and calculate the fundamental frequency. Example: 32 zero crossings detected within a frame → fundamental frequency = 160Hz (sampling rate 16kHz / frame length 320 points) Differential order control signal generation, order mapping curve: preset relationship between fundamental frequency and differential order α (fundamental frequency ↑→α↓): Fundamental frequency (Hz) 50–100 100–150 150–200 α 1.5 1.2 0.9 Current adjustment: base frequency 160Hz → output α=0.95; Analog fractional-order differential processing, circuit adjustment: α = 0.95 → control voltage 1.8V to drive the phase shift circuit, adjust RC parameters (R = 15kΩ, C = 10nF); output: primary differential signal (high-frequency component boosted by 3dB, phase shifted by +42°); Ambient temperature and humidity compensation, parameter collection: temperature 25°C, humidity 60% → sound speed 346m / s; amplitude calibration: look up the table to obtain the compensation coefficient K = 1.05 → differential signal amplitude × 1.05; Resonance enhanced signal output, signal superposition: calibrated differential signal + original signal (mixing ratio 1:0.7); verify: index Before treatment After processing Target speech signal-to-noise ratio 12dB 20dB Fundamental frequency harmonic distortion rate 9% 4% Processing delays — 8ms Note: The data is based on actual conference room noise environment (air conditioning noise + keyboard typing) testing, and the PESQ score increased from 2.8 to 3.6.
[0058] 203. Acoustic boundary layer separation field construction: input the pre-processed signal into the acoustic waveguide model, and generate a separation field containing the voice guided wave mode and the interference radiation mode through real-time phase analysis; Specifically, the waveguide structure parameter generation is based on the ambient air pressure and temperature parameters to calculate the propagation constant of the sound wave in the equivalent waveguide and generate the waveguide physical structure parameter set; multimodal phase decomposition is performed to input the preprocessed signal into the phase analysis array, and the acoustic waveguide wave component and the radiation component are separated through orthogonal phase detection to generate the initial modal separation signal; boundary impedance matching is performed to adjust the waveguide boundary impedance parameters according to the target sound source position information to generate the impedance matching sound field constraint conditions; physical separation field construction is performed to input the initial modal separation signal into the acoustic waveguide model, and the waveguide physical structure parameter set and the sound field constraint conditions are applied to construct the spatial separation field of the voice waveguide mode and the interference radiation mode; real-time modal feedback calibration is performed to collect the energy distribution characteristics of the output separation field, dynamically correct the waveguide structure parameters, and generate the optimized acoustic separation field.
[0059] It should be noted that the waveguide structure parameters are generated, the environmental parameters are: temperature 25℃, air pressure 101.3kPa, humidity 60%, and the sound speed is calculated. m / s (T is degrees Celsius).
[0060] Propagation constant calculation: The equivalent waveguide is a cylindrical air cavity with a diameter of D = 0.4m, and the target speech main frequency Hz, wave number rad / m.
[0061] Output parameter set: waveguide length L = 2.5m, cutoff frequency Hz (supports only fundamental mode propagation).
[0062] Multimodal phase decomposition, input signal: pre-processed signal (after resonance enhancement, SNR = 20dB). Orthogonal phase detection: using a 4-channel phase analysis array, calculate the signal cross-power spectrum phase difference in the 1kHz sub-band. Guided wave component: phase difference (Coherence>0.9). Radiation component: Phase difference (Coherence < 0.2) Output: The separated guided wave mode energy accounts for 85%, and the radiating mode accounts for 15%.
[0063] Boundary impedance matching, target sound source positioning: azimuth angle 30°, distance 1.2m (based on microphone array TDOA estimation). Impedance parameter: boundary acoustic impedance , is the imaginary unit, where: ; ; =1.18kg / m3 (air density), =1.5m (impedance discontinuity location). Calculated .
[0064] Physical separation field construction, acoustic field constraints: guided wave mode area [0, ] is the traveling wave field (phase linearly changes), the radiation mode area [ ,L] is the standing wave field (sound pressure amplitude fluctuation>8dB). Separation field output: = An acoustic barrier is set at 1.5m, with the transmittance of the guided wave component greater than 90% and the reflectivity of the radiation component greater than 85%.
[0065] Real-time modal feedback calibration, energy distribution monitoring: sound pressure amplitude fluctuation in the waveguide area , radiation zone Pa. Parameter correction: When the temperature rises to 28℃, the speed of sound increases to 349m / s, and the dynamic adjustment =1.52m and update .
[0066] verify: index Before calibration After calibration Guided wave mode energy purity 78% 92% Radiation Mode Suppression Ratio 12dB 21dB Processing delays 15ms 8ms 204. Physical equivalent interference reconstruction, extracting interference radiation modal components in the separation field and reconstructing the time domain interference waveform based on the acoustic impedance matching principle; Specifically, radiation mode energy extraction collects the acoustic energy distribution of the interference radiation mode from the specified spatial area of the acoustic separation field to generate a radiation mode energy distribution diagram; acoustic impedance matching parameter generation calculates the acoustic impedance reference value of the target reconstruction frequency band in real time according to the ambient air density and sound speed to generate an impedance matching parameter set; physical waveform synthesis inputs the radiation mode energy distribution diagram into the tunable acoustic resonator array, applies the impedance matching parameter set to drive the resonator vibration, and generates a primary reconstructed interference waveform; time domain characteristic calibration compares the phase correlation of the primary reconstructed interference waveform with the original separation field radiation component to generate a waveform calibration signal; equivalent interference waveform output normalizes the amplitude of the calibrated waveform through a voltage-controlled amplifier to output a physically equivalent time domain interference waveform.
[0067] The acoustic separation field acquisition array consists of eight MEMS sound pressure sensors (frequency response 50 Hz–5 kHz) deployed on the conference room ceiling, spaced 20 cm apart, covering the radiation mode area. The resonator array consists of 32 piezoelectric ceramic resonator units (resonant bandwidth 80 Hz–1.2 kHz) with an integrated tunable LC circuit. The calibration module consists of a high-speed phase detector (resolution 0.5°) and a voltage-controlled amplifier (gain range 0.1–5x).
[0068] Signal processing flow, radiation mode energy extraction, input: interference radiation mode components in the acoustic separation field (main frequency band 500–800 Hz). Energy distribution map generation: Acquire the sound pressure amplitude distribution in the 500 Hz frequency band. Sensor 3 detects a peak sound pressure of 0.8 Pa, and sensor 5 detects 0.4 Pa. Generate a spatial energy gradient map.
[0069] Acoustic impedance matching parameter generation, environmental parameters: temperature 25°C, humidity 60% → sound speed 346m / s, air density 1.18kg / m 3 Impedance calculation: Acoustic impedance reference value of the target reconstruction frequency band 600Hz = 412Rayl. Parameter set output: Generates LC tuning parameters for 32 resonant units (600Hz unit: inductor L = 12mH, capacitor C = 5.8nF).
[0070] Physical waveform synthesis, resonator drive: The energy profile is input into the resonator array, and an 8V drive voltage is applied to the 600Hz unit to generate a sinusoidal waveform. Primary reconstructed waveform: The synthesized waveform has an amplitude of 1.2V at 600Hz and a correlation of 0.85 with the original interference spectrum.
[0071] Time domain characteristic calibration, phase comparison: Detect the phase difference (approximately 15°) between the reconstructed waveform and the original radiation component to generate an inverted calibration signal; amplitude adjustment: Use a voltage-controlled amplifier to increase the gain of the 600Hz component to 1.3 times and reduce the phase difference to 3°.
[0072] Equivalent interference waveform output, normalization processing: full-band amplitude normalized to ±1.5V range (limited clipping distortion <2%); output: time domain interference waveform signal-to-noise ratio 22dB, correlation with the original interference waveform 0.93.
[0073] verify: index Before calibration After calibration Interference waveform correlation 0.85 0.93 Main frequency band amplitude error 18% 5% Processing Delays 12ms 3ms 205. Dynamic spectrum subtraction output, the original mixed speech signal and the reconstructed interference waveform are input into the multi-band canceller, and the pure target speech signal is output.
[0074] Specifically, the frequency band reference signal is generated, the reconstructed interference waveform is input into the analog filter bank, and the LC resonance circuit is separated into 32 independent frequency band reference sub-signals; the environment adaptive gain coefficient is generated, the dynamic gain coefficient set of each frequency band is calculated in real time based on the environmental temperature and humidity parameters and the acoustic impedance characteristics; the frequency band cancellation is performed, the original mixed speech signal and the frequency band reference sub-signals are input into the parallel voltage subtractor array, and the cancellation operation is performed by applying the corresponding gain coefficient to generate the preliminary screening speech sub-signal; the phase coherence calibration is performed, the preliminary screening speech sub-signal is compared with the original separated field speech guided wave component in phase to generate the phase calibration parameter of each frequency band; and the pure speech synthesis output is performed, the calibrated speech sub-signal is impedance matched, and the final pure target speech signal is output through the acoustoelectric conversion.
[0075] It should be noted that the filter bank: 32-channel LC resonance circuit (center frequency 80Hz-4kHz, Q value 0.8), inductance tolerance ±2%. The subtractor array: parallel operational amplifier (gain accuracy 0.1dB), supports ±5V dynamic range. Sensor: temperature and humidity sensor (temperature ±0.3℃, humidity ±2%RH). Phase calibrator: high-speed comparator (delay <0.5ms), reference speech guided wave component phase reference.
[0076] Signal processing flow, frequency band reference signal generation, input: reconstructed interference waveform (main noise frequency band 500-800Hz, amplitude 1.2V). LC filter bank: 500Hz frequency band: L=15mH, C=6.8nF, output reference sub-signal amplitude 0.8V. 3kHz frequency band: L=2mH, C=1.3nF, output reference sub-signal amplitude 0.3V. Output: 32 independent frequency band reference sub-signals (bandwidth 125Hz).
[0077] Environment adaptive gain coefficient generation, parameter acquisition: temperature 26℃, humidity 55%→sound speed 348m / s, acoustic impedance 418Rayl. Gain calculation: low frequency band (<1kHz): gain coefficient 0.92 (air absorption compensation). High frequency band (>2kHz): gain coefficient 1.05 (humidity-induced sound attenuation compensation).
[0078] Frequency band cancellation execution, input: original mixed speech signal (target speech energy ratio 45%, SNR 18dB). Voltage subtractor: 500Hz frequency band: mixed signal 1.5V-reference signal 0.8V×0.92→output 0.76V. 3kHz frequency band: mixed signal 0.9V-reference signal 0.3V×1.05→output 0.59V. Output: preliminary screening speech sub-signal (residual noise energy reduced by 12dB).
[0079] Phase coherence calibration and phase comparison: Using the voice-guided wave component of the acoustic separation field as the benchmark (phase fluctuation <10°). Calibration operation: Detecting an 8° phase deviation of the 500Hz sub-signal → generating a 0.04ms delay compensation. After calibration, the phase difference in each frequency band is ≤5°.
[0080] Pure speech synthesis output, impedance-matched synthesis: 32-channel signal is weighted and superimposed (low frequency weight 0.85, high frequency weight 0.75). Acoustic-to-electrical conversion: output signal-to-noise ratio 28dB, speech clarity STOI 0.94.
[0081] Performance Verification: index Before treatment After processing Full-band residual noise -18dB -42dB Speech harmonic distortion rate 9% 2% Delayed response to environmental changes — 20ms 206. Recover the residual interference sound energy in the acoustic separation field and convert it into electrical energy, while suppressing the mechanical vibration of the equipment, monitoring sudden changes in the environment and triggering system reset, and synchronously outputting pure voice to the bone conduction and air conduction dual-modal interfaces; Specifically, the residual acoustic energy is captured from the radiation mode area of the acoustic separation field, converted into electrical energy through a piezoelectric transducer, and a system power supply compensation signal is generated; the vibration signal of the equipment shell is collected, and an anti-phase vibration waveform is generated to drive the linear actuator to output a mechanically stable compensation force; the acoustic impedance mutation gradient is monitored in real time, and when the gradient exceeds the threshold, an early warning signal is generated and the processing parameters are reset; the pure target voice signal is synchronously converted into a bone-conducted vibration signal and an air-conducted sound signal to generate a dual-modal output.
[0082] The hardware configuration includes the acoustic energy recovery module, which features eight piezoelectric ceramic transducers (sensitivity 1.2 V / Pa, frequency response 80–500 Hz) deployed in the acoustic separation field radiation mode area. The vibration suppression module includes a 3-axis MEMS accelerometer (range ±8 g) and a linear actuator (response frequency 0–200 Hz, thrust ±5 N). The environmental monitoring module includes an acoustic impedance sensor (accuracy ±0.5 Rayleigh), a temperature and humidity sensor (±0.3°C, ±2% RH). Output interfaces include a bone conduction vibrator (resonance frequency 250 Hz) and an air conduction speaker (frequency response 100 Hz–10 kHz).
[0083] Processing flow for residual acoustic energy recovery: Input: Residual acoustic energy in the radiation mode (main frequency 300Hz, sound pressure 0.6Pa). Piezoelectric conversion: 8 transducers connected in parallel, output peak voltage 4.2V, conversion efficiency 18%. Output: System power compensation signal (stable +5V / 200mA).
[0084] Mechanical vibration suppression and vibration acquisition: Equipment housing vibration (main frequency 120Hz, amplitude 0.4g). Anti-phase waveform generation: Generates anti-phase vibration waveforms with a phase difference of 180° and an amplitude of 0.4g. Active cancellation: The linear actuator outputs a compensation force of 3.2N, reducing vibration energy by 15dB.
[0085] Sudden environmental change monitoring and reset. Monitoring parameters: Acoustic impedance gradient threshold set to 10 Rayl / s (based on room temperature at 25°C). Sudden change response: Turning on the air conditioner causes the acoustic impedance gradient to rise to 15 Rayl / s, triggering a parameter reset (waveguide structure and filter coefficients fully refreshed). Latency: 8ms to monitor response time.
[0086] Dual-mode speech output: Bone conduction output: Pure speech signal drives the bone conduction vibrator (gain 0.8, latency <2ms), with a temporal bone conduction sound pressure level of 60dB. Air conduction output: Air conduction speaker output (bandwidth compression to 300–3500Hz, signal-to-noise ratio 28dB).
[0087] Output: Function Before treatment After processing Residual sound energy utilization — 18% Equipment vibration amplitude 0.4g 0.05g Delayed response to sudden environmental changes — 8ms Bone conduction output clarity (STOI) — 0.89 In an embodiment of the present invention, a spatially distributed microphone array is used to synchronously collect ambient sound signals. Through multi-channel time difference compensation, phase coherence enhancement, dynamic generation of spatial filter coefficients and other technologies, an original mixed voice signal with spatial coherence is generated by fusion, which makes full use of spatial information, effectively improves the correlation of signals and the extraction accuracy of target voice. Compared with traditional single-channel processing methods, it can better adapt to the voice collection needs in complex environments; real-time fractional-order differential processing is performed on the original mixed voice signal, and the differential order α is dynamically adjusted according to the voice fundamental frequency. It can perform adaptive processing according to the real-time characteristics of the voice signal, effectively enhance the resonance characteristics of the voice, and highlight the target voice information; the preprocessed signal is input into the acoustic waveguide model, and a separation field containing the voice guided wave mode and the interference radiation mode is generated through real-time phase analysis, thereby achieving effective separation of the voice signal and the interference signal in space, providing a more accurate signal basis for subsequent interference suppression; Extracting the interference radiation modal components in the separation field and reconstructing the time domain interference waveform based on the acoustic impedance matching principle can accurately simulate the interference signal in the environment, provide a more accurate reference signal for dynamic spectrum subtraction, and thus more effectively suppress interference; inputting the original mixed voice signal and the reconstructed interference waveform into the multi-band canceller, through the steps of frequency band reference signal generation, environment adaptive gain coefficient generation, frequency band cancellation, phase coherence calibration, etc., it realizes the precise cancellation of interference signals in different frequency bands and outputs a pure target voice signal; capturing the residual sound energy from the radiation mode area of the acoustic separation field, converting it into electrical energy through a piezoelectric transducer, generating a system power supply compensation signal, realizing energy recovery and reuse, improving the energy utilization efficiency of the system, and reducing dependence on external power supply; The system collects vibration signals from the device housing, generates an anti-phase vibration waveform to drive the linear actuator, and outputs a mechanically stable compensating force. This effectively suppresses mechanical vibration, improves device stability and reliability, and reduces the impact of vibration on voice signal processing. It also monitors sudden changes in acoustic impedance gradients in real time. When the gradient exceeds a threshold, it generates a warning signal and triggers a reset of processing parameters, enabling the system to quickly adapt to environmental changes and ensuring stable operation in various environments. It also simultaneously converts pure target voice signals into bone-conducted vibration signals and air-conducted acoustic signals, generating a dual-modal output that meets the needs of diverse users and improves the flexibility and applicability of voice output. Furthermore, the bone-conducted output achieves a clarity of 0.89, ensuring the quality of the output voice.
[0088] The above describes the natural sound interference stripping algorithm in the embodiment of the present invention. The following describes the natural sound interference stripping device in the embodiment of the present invention. Figure 3 In one embodiment of the present invention, an embodiment of the natural sound interference stripping device includes: an acquisition module 301, which is used to collect environmental sound signals and fuse them to generate an original mixed speech signal; a processing module 302, which is used to perform real-time fractional-order differential processing on the original mixed speech signal, dynamically adjust the differential order α according to the speech fundamental frequency, and output a preprocessed signal; a setting module 303, which is used to input the preprocessed signal into an acoustic waveguide model, and generate a separation field containing a speech guided wave mode and an interference radiation mode through real-time phase analysis; a reconstruction module 304, which is used to extract the interference radiation mode component in the separation field and reconstruct the time domain interference waveform based on the acoustic impedance matching principle; and an allocation module 305, which is used to input the original mixed speech signal and the reconstructed interference waveform into a multi-band canceller to output a pure target speech signal.
[0089] In an embodiment of the present invention, the processing module uses real-time fractional-order differentiation to process the original mixed speech signal and can dynamically adjust the differential order α according to the speech fundamental frequency. Compared with traditional fixed parameter processing, it can more accurately adapt to different speech characteristics, effectively improve the preprocessing effect of the original signal, and lay a good foundation for subsequent separation of interference and extraction of target speech. The setting module inputs the preprocessed signal into the acoustic waveguide model and generates a separation field containing the speech guided wave mode and the interference radiation mode through real-time phase analysis. It can clearly distinguish the modes of speech and interference, providing a key basis for accurate interference removal. The reconstruction module extracts the interference radiation mode component in the separation field and reconstructs the time domain interference waveform based on the acoustic impedance matching principle. The application of the acoustic impedance matching principle makes the reconstructed interference waveform closer to the actual interference situation, improving the accuracy and scientificity of the interference waveform reconstruction. The distribution module inputs the original mixed speech signal and the reconstructed interference waveform into the multi-band canceller and outputs a pure target speech signal. The use of the multi-band canceller can efficiently cancel the interference in different frequency bands, effectively improve the purity of the output target speech signal, and obtain high-quality target speech even in complex acoustic environments.
[0090] above Figure 3 The natural sound interference stripping device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The natural sound interference stripping device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0091] Figure 4 4 is a schematic diagram of the structure of a natural sound interference stripping device provided by an embodiment of the present invention. The natural sound interference stripping device 400 may have relatively large differences due to different configurations or performances. The natural sound interference stripping device 400 includes a transmitter 401, a receiver 402 and a processor 403. The processor 403 may also be a controller. Figure 4 denoted as “controller / processor 403 ”. Optionally, the natural sound interference stripping device 400 may further include a modem processor 405 , wherein the modem processor 405 may include an encoder 406 , a modulator 407 , a decoder 408 and a demodulator 409 .
[0092] In one example, transmitter 401 conditions (e.g., performs analog-to-analog conversion, filtering, amplification, and frequency upconversion) the output samples and generates an uplink signal, which is transmitted via an antenna to an access network device. On the downlink, the antenna receives the downlink signal transmitted by the access network device. Receiver 402 conditions (e.g., performs filtering, amplification, frequency downconversion, and digitization) the signal received from the antenna and provides input samples. Within modem processor 405, encoder 406 receives service data and signaling messages to be transmitted on the uplink and processes them (e.g., formats, encodes, and interleaves them). Modulator 407 further processes (e.g., performs symbol mapping and modulation) the encoded service data and signaling messages and provides output samples. Demodulator 409 processes (e.g., demodulates) the input samples and provides symbol estimates. Decoder 408 processes (e.g., deinterleaves and decodes) the symbol estimates and provides decoded data and signaling messages, which are transmitted to natural acoustic interference stripping device 400. The encoder 406, modulator 407, demodulator 409, and decoder 408 can be implemented by the combined modem processor 405. These units perform processing based on the radio access technology (e.g., LTE and other evolved system access technologies) used by the radio access network. It should be noted that when the natural sound interference removal device 400 does not include the modem processor 405, the above functions of the modem processor 405 can also be performed by the processor 403.
[0093] Processor 403 controls and manages the operations of natural sound interference removal device 400, and is configured to execute the processing performed by natural sound interference removal device 400 in the above-described embodiments of the present disclosure. For example, processor 403 is also configured to execute the various steps of the transmitting device or receiving device in the above-described algorithm embodiments, and / or other steps of the technical solutions described in the embodiments of the present disclosure.
[0094] Furthermore, the natural sound interference removal device 400 may further include a memory 404 , and the memory 404 is used to store program codes and data used for the natural sound interference removal device 400 .
[0095] It is understandable that Figure 4 Only a simplified design of the natural sound interference stripping device 400 is shown. In actual applications, the natural sound interference stripping device 400 may include any number of transmitters, receivers, processors, modem processors, memories, etc., and all devices that can implement the embodiments of the present disclosure are within the scope of protection of the embodiments of the present disclosure.
[0096] The present invention also provides a natural sound interference stripping device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the natural sound interference stripping algorithm in the above-mentioned embodiments.
[0097] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the natural sound interference stripping algorithm.
[0098] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned algorithm embodiments and will not be repeated here.
[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the algorithm described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0100] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A natural sound interference stripping algorithm, characterized in that: The natural sound interference stripping algorithm includes: Collect environmental sound signals and fuse them to generate original mixed speech signals; Performing real-time fractional-order differential processing on the original mixed speech signal, dynamically adjusting the differential order α according to the speech fundamental frequency, and outputting a preprocessed signal; Inputting the preprocessed signal into an acoustic waveguide model, and generating a separation field including a speech guided wave mode and an interference radiation mode through real-time phase analysis; Extracting interference radiation modal components in the separation field and reconstructing the time domain interference waveform based on the acoustic impedance matching principle; The original mixed speech signal and the reconstructed interference waveform are input into a multi-band canceller to output a pure target speech signal.
2. The natural sound interference stripping algorithm according to claim 1, characterized in that: include: Perform time difference of arrival compensation on the original sound signals collected by each channel of the microphone array to generate multi-channel signals; Inputting the multi-channel signal into a phase coherent detector, extracting the coherent component dominated by the target speech, and outputting a phase-synchronized enhanced signal; Calculate spatial filtering parameters in real time based on the environmental acoustic impedance characteristics and generate an adaptive filter coefficient set that matches the current sound field; The phase synchronization enhanced signal is input into a spatial filter, and the adaptive filter coefficient set is applied to perform multi-channel fusion to output an original mixed speech signal.
3. The natural sound interference stripping algorithm according to claim 1, characterized in that: include: Perform time domain zero-crossing analysis on the original mixed speech signal to generate real-time updated speech fundamental frequency parameters; Input the voice fundamental frequency parameter into the order mapper, and output a dynamic differential order control signal through a preset fundamental frequency-differential order relationship curve; Inputting the original mixed voice signal into an analog differential circuit, and applying the differential order control signal to adjust the circuit impedance characteristics in real time to generate a primary differential signal; Collecting environmental temperature and humidity parameters, generating a circuit compensation signal through a sound velocity correction model, and performing amplitude calibration on the primary differential signal; The calibrated differential signal is superimposed on the original signal to output the preprocessed signal.
4. The natural sound interference stripping algorithm according to claim 1, characterized in that: include: Based on the ambient air pressure and temperature parameters, the propagation constant of the acoustic wave in the equivalent waveguide is calculated to generate the waveguide physical structure parameter set; Inputting the preprocessed signal into a phase analysis array, separating the acoustic waveguide component and the radiation component by orthogonal phase detection, and generating an initial modal separation signal; According to the target sound source location information, the waveguide boundary impedance parameters are adjusted to generate the acoustic field constraint conditions for impedance matching; Inputting the initial modal separation signal into an acoustic waveguide model, applying the waveguide physical structure parameter set and acoustic field constraints, and constructing a spatial separation field of the voice guided wave mode and the interference radiation mode; The energy distribution characteristics of the output spatial separation field are collected, the waveguide structure parameters are dynamically modified, and the acoustic separation field is generated.
5. The natural sound interference stripping algorithm according to claim 4, characterized in that: Set the boundary impedance parameter to ,but: ; in, ; ; Real part is the acoustic impedance; imaginary part For sound resistance; is an imaginary unit; is the medium density; is the speed of sound; k is the wave number; L is the length of the waveguide; is the impedance discontinuity location.
6. The natural sound interference stripping algorithm according to claim 4, characterized in that: include: Acquiring the acoustic energy distribution of the interference radiation mode from a specified spatial region of the acoustic separation field to generate a radiation mode energy distribution diagram; Calculate the acoustic impedance reference value of the target reconstruction frequency band in real time based on the ambient air density and sound velocity, and generate an impedance matching parameter set; Inputting the radiation mode energy distribution diagram into a tunable acoustic resonator array, applying the impedance matching parameter set to drive the resonators to vibrate, and generating a primary reconstructed interference waveform; Performing phase correlation comparison between the primary reconstructed interference waveform and the original separation field radiation component to generate a waveform calibration signal; The calibrated waveform is amplitude normalized through a voltage-controlled amplifier to output a time-domain interference waveform.
7. The natural sound interference stripping algorithm according to claim 6, characterized in that: include: The reconstructed interference waveform is input into the analog filter bank and separated into 32 reference sub-signals of independent frequency bands through the LC resonant circuit; Based on the ambient temperature and humidity parameters and acoustic impedance characteristics, the dynamic gain coefficient set of each frequency band is calculated in real time; The original mixed speech signal and the reference sub-signals of each frequency band are input into the parallel voltage subtractor array, and the corresponding gain coefficients are applied to perform an offset operation to generate the primary screening speech sub-signals; Perform phase comparison between the initial screening speech sub-signal and the original separation field speech guided wave component to generate phase calibration parameters for each frequency band; The calibrated speech sub-signals are subjected to impedance matching synthesis, and the pure target speech signal is output through acoustic-to-electrical conversion.
8. The natural sound interference stripping algorithm according to claim 7, characterized in that: It also includes converting residual interfering acoustic energy in the acoustic separation field into electrical energy, suppressing mechanical vibration of the device, monitoring sudden changes in the environment and triggering system resets, and synchronously outputting pure speech to the bone conduction and air conduction dual-modal interfaces: Residual acoustic energy is captured from the radiation mode region of the acoustic separation field and converted into electrical energy through a piezoelectric transducer to generate a system power compensation signal; Collect the vibration signal of the equipment shell, generate the anti-phase vibration waveform to drive the linear actuator, and output the mechanical stabilization compensation force; Real-time monitoring of acoustic impedance gradients. When the gradient exceeds a threshold, a warning signal is generated and processing parameters are reset. The pure target speech signal is synchronously converted into a bone-conducted vibration signal and an air-conducted acoustic signal to generate a dual-modal output.
9. An application of natural sound interference stripping in noise monitoring, characterized in that: Includes natural sound interference stripping device: The acquisition module is used to collect environmental sound signals and fuse them to generate the original mixed speech signal; A processing module, configured to perform real-time fractional-order differential processing on the original mixed speech signal, dynamically adjust the differential order α according to the speech fundamental frequency, and output a preprocessed signal; A setting module is used to input the preprocessed signal into the acoustic waveguide model, and generate a separation field including a voice guided wave mode and an interference radiation mode through real-time phase analysis; A reconstruction module, configured to extract interference radiation modal components in the separation field and reconstruct a time domain interference waveform based on the acoustic impedance matching principle; The distribution module is used to input the original mixed speech signal and the reconstructed interference waveform into the multi-band canceller and output a pure target speech signal.
Citation Information
Patent Citations
Signal separation method based on fractional wavelet transform
CN101655834A
Method for separating monaural overlapping speeches based on fractional Fourier transform (FrFT)
CN102054480A
Speech separation method and device, mobile terminal and computer readable storage medium
CN110808061A
De-noising system based on adaptive beam forming and ICA (independent component analysis)
CN120375849A
Systems, methods, apparatus, and computer-readable media for far-field multi-source tracking and separation
US20120099732A1
Cited By
Indirect-coherent measurement method for forward sound scattering target intensity of underwater object
CN121364467A