Voice Activity Detection for Continuous Listening Noise Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In Continuous Listening environments, voice recognition systems face malfunctions due to activation by various noise signals instead of voice signals, as they cannot anticipate start and end points of voice utterance, leading to incorrect signal processing.

Innovation Solution

A voice activity detection method and apparatus that extracts feature parameters from frame signals, compares them with model parameters of noise and voice signals, and outputs only voice signals, using techniques like energy calculation, likelihood ratio analysis, and index value correction to differentiate between voice and noise signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition systems operate in Continuous Listening environment without physical interface activation, then ease of operation is improved, but reliability deteriorates due to activation by noise signals instead of voice signals

Engineering Contradiction:
Improveease of operationVSAvoidreliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces a voice activity detection module as an intermediary between the continuous listening environment and the voice recognition system. This module extracts feature parameters from input signals, compares them with model parameters of noise and voice signals, and determines whether the signal is actual voice or noise. By adding this intermediary classification layer, the system maintains ease of continuous operation while improving reliability through accurate voice-noise discrimination before processing

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by performing voice activity detection and signal classification before the main voice recognition processing. The system pre-processes incoming signals by extracting features, comparing them against predefined voice and noise models, and filtering out noise signals before they reach the voice recognition engine. This preliminary filtering prevents noise-induced malfunctions while maintaining continuous listening capability

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the system processes all incoming signals in Continuous Listening environment, then productivity is improved, but loss of information increases due to processing noise signals incorrectly

Engineering Contradiction:
ImproveproductivityVSAvoidloss of information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent extracts and removes noise signals from the processing stream by comparing feature parameters against model parameters of various noise types. The voice activity detection module identifies and extracts only genuine voice signals while discarding noise signals, preventing information loss that would occur from processing incorrect signals. This selective extraction maintains productivity by ensuring continuous processing while avoiding information loss from noise interference

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8762144B2Method and apparatus for voice activity detection
Publication Date: 2014.06.24 SAMSUNG ELECTRONICS CO LTD
  • US8762144B2 patent drawing
  • US8762144B2 patent drawing
  • US8762144B2 patent drawing

AI summary

A method and apparatus for detecting voice activity are disclosed. The method of detecting voice activity includes: extracting a feature parameter from a frame signal; determining whether the frame signal is a voice signal or a noise signal by comparing the feature parameter with model parameters of a plurality of comparison signals, respectively; and outputting the frame signal when the frame signal is determined to be a voice signal. The apparatus includes a classifier module which extracts a feature parameter from a frame signal, and generating labeling information with respect to the frame signal by comparing the feature parameter with model parameters of a plurality of comparison signals; and a voice detection unit which determines whether the frame signal is a noise signal or a voice signal with reference to the labeling information, and outputting the frame signal when the frame signal is determined to be a voice signal.