Movable Robot Speaker Tracking via Diagonal Microphone Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Movable robots, especially larger ones, face difficulties in accurately tracking the position of a speaker due to the robot's body obstructing the voice signal, leading to poor voice reception by microphones and degradation in sound quality, which hampers the estimation of the speaker's direction using conventional GCC algorithms.
Innovation Solution
The method involves a movable robot with microphones installed at the vertices of a quadrangle cross-section, where the robot receives wake-up voices through diagonally opposed microphones, calculates reference values based on gain or confidence scores, and uses sound source localization (SSL) values to track the speaker's position, selecting appropriate microphones to ensure reliable voice reception and accurate direction determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Volume of moving object
If the robot is made larger to provide better services, then the robot can provide more comprehensive services, but the robot body obstructs the voice signal from reaching the microphones
Solution Approach 1:
The patent divides the robot's microphone system into multiple segments positioned at different locations (front, rear, left, right) of the robot body. This segmentation allows the system to receive voice signals from multiple directions simultaneously, overcoming the obstruction problem caused by the robot's large size.
Solution Approach 2:
The patent transitions from a single-point microphone reception (one-dimensional) to a distributed multi-point microphone array (three-dimensional spatial distribution). By positioning microphones at multiple vertices of a quadrangle cross-section, the system creates a spatial dimension for voice signal reception, enabling accurate direction tracking despite the robot's size.
2Area of stationary object
If microphones are installed at the rear of the robot to expand coverage, then the coverage area increases, but the voice receiving quality degrades due to robot body obstruction
Solution Approach 1:
The patent segments the voice reception function across multiple microphones positioned at different spatial locations. Each microphone captures voice signals from its specific direction, and the system integrates these segmented signals to achieve both wide coverage and high reception quality.
Solution Approach 2:
The patent applies local quality by optimizing the position and function of each individual microphone based on its location. Front microphones are optimized for forward-facing voice reception, while rear microphones handle backward-facing signals, with each microphone's characteristics tailored to its specific spatial position and reception conditions.
3Measurement precision
If the GCC algorithm is used to estimate speaker direction, then the direction estimation can be performed, but the system requires two microphones to receive voice properly and fails when one microphone does not receive the voice
Solution Approach 1:
The patent divides the direction estimation function into multiple independent calculation paths, each using a different pair of microphones. By segmenting the estimation process across multiple microphone pairs, the system ensures that if one path fails due to poor voice reception, other paths can still provide reliable direction estimation.
Solution Approach 2:
The patent changes the parameters of the direction estimation system by using multiple microphone pairs with different spatial configurations. This allows the system to adaptively select the most suitable microphone pair based on the actual voice reception conditions, thereby maintaining high reliability in direction estimation.
4Area of stationary object
If more microphones are installed to improve voice reception from all directions, then the directional coverage improves, but the device complexity increases
Solution Approach 1:
The patent segments the voice reception coverage into four distinct directional zones (front, rear, left, right), with one microphone positioned at each vertex of a quadrangle cross-section. This segmentation achieves comprehensive directional coverage while maintaining a simple and symmetrical microphone configuration.
Solution Approach 2:
The patent merges the functions of multiple microphones into a unified direction tracking system. By combining the signals from all four microphones and processing them through a single integrated algorithm, the system achieves comprehensive coverage without proportionally increasing processing complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows the robot to accurately track the speaker's direction by selecting the most reliable microphones, enhancing the confidence score and overcoming the limitations of conventional methods, particularly for larger robots where voice obstruction is a significant issue.
Implementation Method 1
receiving a wake-up voice through first and third microphones disposed respectively at first and third vertices
Data Source
AI summary
Proposed is a method for determining, by a movable robot, a position of a speaker, wherein the movable robot includes first to fourth microphones installed at four vertexes of a quadrangle of a horizontal cross section of the robot respectively, wherein the method includes: receiving a wake-up voice through first and third microphones disposed respectively at first and third vertices in a diagonal direction; obtaining a first reference value of the first microphone and a second reference value of the third microphone based on the received wake-up voice; comparing the obtained first and second reference values to select the first microphone; selecting a second microphone disposed at a second vertex, wherein the first and second microphones are on a front side of the quadrangle; calculating a sound source localization (SSL) value based on the selected first and second microphones; and tracking a position of the speaker based on the SSL value.


