Robot Voice Detection Using SLAM and Acoustic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice detection systems in humanoid robots fail to accurately determine whether a detected sound source is a human voice, leading to unnecessary rotations and interactions with non-human sound sources, such as TVs or radios.

Innovation Solution

A voice detection apparatus equipped with multiple microphones and a processor that creates a SLAM map of the environment, uses sound source localization techniques like MUSIC, and incorporates voice occurrence probability databases to differentiate between human voices and other sound sources, allowing the robot to selectively rotate its head towards human voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the robot rotates its head towards every detected sound source, then it can respond to all potential voice inputs, but it causes unnecessary rotations when the sound source is not a human voice

Engineering Contradiction:
Improvevoice detection accuracyVSAvoidoperational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary analysis of the sound source before triggering the head rotation action. By examining acoustic characteristics and comparing against human voice profiles in advance, the system determines whether the sound source is likely to be a human voice, thereby avoiding unnecessary rotations while maintaining reliable detection of actual human callers.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the robot uses simple amplitude threshold detection, then the system remains simple and fast, but it cannot accurately distinguish human voices from other sound sources

Engineering Contradiction:
Improvesound source identification accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies different analysis methods to different aspects of the sound signal. It uses amplitude thresholding for initial detection while applying more sophisticated acoustic characteristic analysis specifically to the frequency and temporal patterns of the detected sound. This localized application of complex analysis only where needed improves identification accuracy without proportionally increasing overall system complexity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system transitions from detecting only a single parameter (amplitude) to analyzing multiple parameters including frequency spectrum, temporal patterns, and acoustic characteristics. By changing from monolithic single-parameter detection to multi-parameter analysis, the system achieves better sound source identification while managing complexity through selective application of analysis methods.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10424320B2Voice detection, apparatus, voice detection method, and non-transitory computer-readable storage medium
Publication Date: 2019.09.24 CASIO COMPUTER CO LTD
  • US10424320B2 patent drawing
  • US10424320B2 patent drawing
  • US10424320B2 patent drawing

AI summary

A robot determines whether a voice is a voice emanated directly from an actual person or is a voice output from a speaker of an electronic device.A controller of a robot detects a voice by means of microphones, determines whether or not a voice generating source of the detected voice is a specific voice generating source, and controls, based on a result of the determination, the robot by means of a neck joint and a chassis.