Speech Processing Apparatus Harmonic Structure Noise Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing systems face challenges in accurately distinguishing between noise components and speech components, especially in environments with high levels of periodic noise, leading to erroneous detection and increased power consumption due to heavy processing loads.

Innovation Solution

A speech processing apparatus and method that extracts frames from input signals, converts them into the frequency domain, detects peak spectra, and determines harmonic structures to differentiate between speech and noise components, using a harmonic-overtone determination unit to accurately identify speech segments and attenuate periodic noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If cepstrum analysis is used to detect speech segments, then speech detection accuracy is improved, but power consumption increases and processing load becomes heavy

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The invention extracts only the essential spectral features (peak spectra and harmonic structure) needed for speech detection, rather than performing complete cepstrum analysis. This selective extraction maintains speech detection accuracy while significantly reducing computational load and power consumption in battery-powered systems

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The speech detection process is segmented into distinct stages: frame extraction, spectrum generation, peak detection, and harmonic-overtone determination. This segmentation allows the system to process only relevant portions of the signal at each stage, reducing overall processing requirements while maintaining detection precision

Inventive Principle:
Principle #1Segmentation

2Device complexity

If known speech detection processes are used in noisy environments, then processing is simple, but speech detection accuracy deteriorates due to erroneous detection of noise as speech

Engineering Contradiction:
Improveprocessing complexityVSAvoidspeech detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The invention applies partial action by implementing only the necessary steps (peak detection and harmonic structure analysis) rather than complete spectral analysis. This partial approach maintains simplicity while improving accuracy in noisy environments by focusing computational resources on the most discriminative features

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The invention changes the detection parameters from general spectral analysis to specific harmonic structure parameters. By analyzing the relationships between fundamental pitch and harmonic overtones, the system can distinguish speech from noise even in complex acoustic environments, improving detection accuracy without proportionally increasing complexity

Inventive Principle:
Principle #35Parameter changes

3Speed

If periodicity-based voice detection is used, then detection speed is improved, but reliability deteriorates because periodic noises are erroneously identified as voices

Engineering Contradiction:
Improvedetection speedVSAvoiddetection reliability
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The invention adds a new dimension to voice detection by analyzing the harmonic structure relationships between different frequency components. Instead of relying solely on temporal periodicity, the system examines the spectral relationships between fundamental pitch and harmonic overtones, creating a multi-dimensional detection approach that distinguishes voice from periodic noise

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The harmonic structure analysis serves as an intermediary between simple periodicity detection and complete spectral analysis. This intermediate approach maintains the speed benefits of periodicity-based methods while adding the reliability of structural analysis to differentiate voice from noise

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables accurate speech segmentation with reduced power consumption and improved sound quality in noisy environments, enhancing speech recognition and communication systems by effectively distinguishing between speech and noise components.

Implementation Method 1

a spectrum generation unit configured to convert the per-frame input signal in a time domain into a per-frame input signal in a frequency domain, thereby generating a spectral pattern of spectra

Methodology Applied
Scientific EffectFourier transform:

Data Source

PatentUS8818806B2Speech processing apparatus and speech processing method
Publication Date: 2014.08.26 JVC KENWOOD CORP
  • US8818806B2 patent drawing
  • US8818806B2 patent drawing
  • US8818806B2 patent drawing

AI summary

A signal portion is extracted per frame having a specific duration from an input signal, thus generating a per-frame input signal. The per-frame input signal in the time domain is converted into a per-frame input signal in the frequency domain, thereby generating a spectral pattern of spectra. Peak spectra having peaks are detected in the spectral pattern. A harmonic spectrum is determined, in the peak spectra, having a harmonic structure showing a relationship between a fundamental pitch and a harmonic overtone.