Speech Processing Apparatus Harmonic Structure Noise Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech processing systems face challenges in accurately distinguishing between noise components and speech components, especially in environments with high levels of periodic noise, leading to erroneous detection and increased power consumption due to heavy processing loads.
Innovation Solution
A speech processing apparatus and method that extracts frames from input signals, converts them into the frequency domain, detects peak spectra, and determines harmonic structures to differentiate between speech and noise components, using a harmonic-overtone determination unit to accurately identify speech segments and attenuate periodic noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cepstrum analysis is used to detect speech segments, then speech detection accuracy is improved, but power consumption increases and processing load becomes heavy
Solution Approach 1:
The invention extracts only the essential spectral features (peak spectra and harmonic structure) needed for speech detection, rather than performing complete cepstrum analysis. This selective extraction maintains speech detection accuracy while significantly reducing computational load and power consumption in battery-powered systems
Solution Approach 2:
The speech detection process is segmented into distinct stages: frame extraction, spectrum generation, peak detection, and harmonic-overtone determination. This segmentation allows the system to process only relevant portions of the signal at each stage, reducing overall processing requirements while maintaining detection precision
2Device complexity
If known speech detection processes are used in noisy environments, then processing is simple, but speech detection accuracy deteriorates due to erroneous detection of noise as speech
Solution Approach 1:
The invention applies partial action by implementing only the necessary steps (peak detection and harmonic structure analysis) rather than complete spectral analysis. This partial approach maintains simplicity while improving accuracy in noisy environments by focusing computational resources on the most discriminative features
Solution Approach 2:
The invention changes the detection parameters from general spectral analysis to specific harmonic structure parameters. By analyzing the relationships between fundamental pitch and harmonic overtones, the system can distinguish speech from noise even in complex acoustic environments, improving detection accuracy without proportionally increasing complexity
3Speed
If periodicity-based voice detection is used, then detection speed is improved, but reliability deteriorates because periodic noises are erroneously identified as voices
Solution Approach 1:
The invention adds a new dimension to voice detection by analyzing the harmonic structure relationships between different frequency components. Instead of relying solely on temporal periodicity, the system examines the spectral relationships between fundamental pitch and harmonic overtones, creating a multi-dimensional detection approach that distinguishes voice from periodic noise
Solution Approach 2:
The harmonic structure analysis serves as an intermediary between simple periodicity detection and complete spectral analysis. This intermediate approach maintains the speed benefits of periodicity-based methods while adding the reliability of structural analysis to differentiate voice from noise
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables accurate speech segmentation with reduced power consumption and improved sound quality in noisy environments, enhancing speech recognition and communication systems by effectively distinguishing between speech and noise components.
Implementation Method 1
a spectrum generation unit configured to convert the per-frame input signal in a time domain into a per-frame input signal in a frequency domain, thereby generating a spectral pattern of spectra
Data Source
AI summary
A signal portion is extracted per frame having a specific duration from an input signal, thus generating a per-frame input signal. The per-frame input signal in the time domain is converted into a per-frame input signal in the frequency domain, thereby generating a spectral pattern of spectra. Peak spectra having peaks are detected in the spectral pattern. A harmonic spectrum is determined, in the peak spectra, having a harmonic structure showing a relationship between a fundamental pitch and a harmonic overtone.


