Robot Voice Direction Detection via Multi-Stage Sound Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound source separation techniques struggle to effectively isolate a signal voice from noise voices in robotic systems, hindering the ability of robots to accurately detect the direction or position of a target person's voice for communication and movement purposes.

Innovation Solution

A sound source separation information detecting device and method that utilizes a voice acquisition unit with predetermined directivity, combined with first and second direction detection units, to identify the arrival directions of signal and noise voices, and a detection unit to determine the sound source separation direction or position, employing beam forming and multiple signal classification techniques for effective separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional beam forming techniques are used for sound source separation, then the signal-to-noise ratio can be improved to some extent, but the robot cannot accurately detect the direction or position of the target person's voice in complex noise environments

Engineering Contradiction:
Improvedirection detection precisionVSAvoidvoice recognition reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The invention segments the sound source separation process into multiple stages: first separating signal voices from noise voices using beam forming, then further separating multiple signal voices from each other using spectral information and direction detection. This multi-stage segmentation enables accurate direction detection even in complex noise environments where conventional single-stage beam forming fails.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention introduces spectral information as an intermediary element between the acoustic signal and the direction detection process. By analyzing spectral characteristics of sound sources and using this information as a mediator for direction estimation, the system achieves more accurate voice direction detection than conventional methods that rely solely on spatial filtering.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the robot uses multiple microphones for sound source separation, then the ability to separate signal voice from noise voice improves, but the device complexity increases

Engineering Contradiction:
Improvesound source separation precisionVSAvoidmicrophone array complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The invention dynamically adjusts the beam forming parameters and spectral analysis thresholds based on the detected sound source characteristics and environmental conditions. This dynamic adaptation allows the system to maintain high separation precision while using a moderate number of microphones, as the system optimizes its performance through software algorithms rather than requiring a fixed complex hardware architecture.

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution enhances the signal-to-noise ratio, allowing robots to accurately recognize and respond to a target person's voice by optimizing the separation of signal voices from noise, enabling successful voice recognition and communication.

Implementation Method 1

a voice acquisition unit having predetermined directivity to acquire a voice

Methodology Applied
Scientific EffectDirectivity:

Implementation Method 2

detect a first direction, which is an arrival direction of a signal voice, from the voice acquired by the voice acquisition unit

Methodology Applied
Scientific EffectSound wave propagation: Sound

Data Source

PatentUS10665249B2Sound source separation for robot from target voice direction and noise voice direction
Publication Date: 2020.05.26 CASIO COMPUTER CO LTD
  • US10665249B2 patent drawing
  • US10665249B2 patent drawing
  • US10665249B2 patent drawing

AI summary

A voice input unit has predetermined directivity for acquiring a voice. A sound source arrival direction estimation unit operating as a first direction detection unit detects a first direction, which is an arrival direction of a signal voice of a predetermined target, from the acquired voice. Moreover, a sound source arrival direction estimation unit operating as a second direction detection unit detects a second direction, which is an arrival direction of a noise voice, from the acquired voice. A sound source separation unit, a sound volume calculation unit, and a detection unit having an S/N ratio calculation unit detect a sound source separation direction or a sound source separation position, based on the first direction and the second direction.