Partial Speech Reconstruction for Vehicle Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech processing systems in vehicles struggle to effectively improve speech intelligibility and quality due to environmental noise and interference, often suppressing noise across large frequency bands, which can lead to residual noise in lower frequencies and unintelligible voiced segments.
Innovation Solution
A speech reconstruction system that includes a low-frequency reconstruction controller, a harmonic generator, and a gain controller to selectively generate and adjust low-frequency harmonics within a predetermined frequency range, matching signal strength to the original input, while blocking noise above and below the selected band.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If systems suppress a fixed amount of noise across large frequency bands, then noise suppression is achieved, but residual noise remains in lower frequencies and speech quality degrades
Solution Approach 1:
The system divides the frequency spectrum into multiple bands (low-frequency band below first threshold, mid-frequency band between thresholds, high-frequency band above second threshold) and applies different processing strategies to each band. This segmentation allows targeted noise suppression in specific frequency ranges while preserving speech quality in others, resolving the contradiction between overall noise suppression and speech quality maintenance.
Solution Approach 2:
Different frequency bands receive different processing treatments: low-frequency band gets harmonic reconstruction, mid-frequency band gets selective noise suppression, and high-frequency band gets different handling. This local quality approach ensures that noise suppression is applied where appropriate without degrading speech quality in critical frequency ranges, particularly preserving voiced segments.
2Object-affected harmful factors
If systems attenuate or eliminate large portions of speech while suppressing noise, then noise reduction is improved, but voiced segments become unintelligible
Solution Approach 1:
The system dynamically adjusts processing parameters based on real-time analysis of the input signal. The low-frequency reconstruction controller continuously monitors the signal and adapts the harmonic generation and gain adjustment parameters accordingly. This dynamic approach allows the system to suppress noise when present while preserving speech intelligibility when speech is present, resolving the contradiction between noise reduction and information loss.
Solution Approach 2:
The gain controller uses feedback from the original input signal to adjust the amplitude of reconstructed harmonics. By comparing the reconstructed low-frequency signal with the original input and adjusting gains accordingly, the system ensures that noise is reduced while speech intelligibility is maintained, preventing the loss of important speech information.
3Manufacturing precision
If systems reconstruct speech across the full frequency band, then speech quality is improved, but latency increases and accuracy decreases
Solution Approach 1:
The system reconstructs speech selectively in specific frequency bands (primarily low-frequency band below the first threshold) rather than attempting full-band reconstruction. This segmentation approach reduces computational complexity and processing time, thereby reducing latency while still improving speech quality in the most critical frequency ranges where harmonics are most important for intelligibility.
Solution Approach 2:
The system applies partial reconstruction by focusing computational resources on reconstructing only the low-frequency harmonics that are most critical for speech intelligibility, rather than attempting to reconstruct the entire frequency spectrum. This partial action approach achieves acceptable speech quality improvement with reduced computational load and lower latency.
Data Source
AI summary
A system improves speech intelligibility by reconstructing speech segments. The system includes a low-frequency reconstruction controller programmed to select a predetermined portion of a time domain signal. The low-frequency reconstruction controller substantially blocks signals above and below the selected predetermined portion. A harmonic generator generates low-frequency harmonics in the time domain that lie within a frequency range controlled by a background noise modeler. A gain controller adjusts the low-frequency harmonics to substantially match the signal strength to the time domain original input signal.


