Robot Attention Shifting Using Audio-Visual Speaker Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Human-Robot Interaction systems face challenges in effectively shifting attention to multiple speakers in group conversations, particularly in telepresence applications, due to limited audio-visual perception and rigid camera adjustments, leading to a lack of sense of presence and belongingness for remote participants.

Innovation Solution

A processor-implemented method and system that uses sound source localization and audio-visual perception to estimate direction of arrivals of voice activity, generate clusters of qualified DOAs, and dynamically rotate the robot to focus on the current speaker, enhancing visual perception and interaction by continuously updating a knowledge repository and applying Human-Robot Interaction rules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the robot uses audio-visual perception to track multiple speakers in real-time, then the sense of presence and belongingness for remote participants is improved, but the device complexity increases due to integration of multiple sensors and processing modules

Engineering Contradiction:
Improvesense of presenceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple independent modules (audio processing with microphone array, visual processing with camera, speaker tracking algorithms) into an integrated attention shifting system. The state representation model merges audio-visual data streams to jointly represent speaker positions and robot attention states, achieving coordinated multi-modal perception while managing system complexity through modular integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The robot system performs multiple functions simultaneously: it captures audio from multiple directions using a microphone array, processes visual information from cameras, estimates direction of arrivals for multiple speakers, tracks speaker positions over time, and dynamically adjusts its attention orientation. This multi-functional capability allows a single system to handle complex group conversation scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If the robot dynamically adjusts its orientation to follow active speakers in group conversations, then the interaction quality is improved, but the response time may be delayed due to continuous processing and clustering computations

Engineering Contradiction:
Improveinteraction qualityVSAvoidresponse time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system continuously estimates direction of arrivals and maintains updated speaker position information in a knowledge repository even during periods of no active speech. This preliminary processing ensures that when a speaker becomes active, the robot can quickly retrieve pre-computed position data and respond without delay, rather than waiting for full processing cycles.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The robot's attention orientation is dynamically adjusted based on real-time speaker activity detection and position changes. The system adapts its response behavior by detecting voice activity and transitioning to tracking mode, allowing flexible response timing that balances processing accuracy with interaction responsiveness.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If the robot uses a fixed camera adjustment mechanism, then the device complexity is reduced, but the ability to track multiple speakers and maintain visual contact is compromised

Engineering Contradiction:
Improvecamera system complexityVSAvoidspeaker tracking capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The camera system transitions from a fixed adjustment mechanism to a dynamic orientation system that actively tracks speakers. The camera orientation is continuously updated based on estimated speaker positions and robot attention direction, enabling the system to maintain visual contact with active speakers while keeping the overall device architecture relatively simple through software-controlled adjustment.

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system improves user experience by accurately shifting attention to active speakers, increasing the sense of presence and belongingness for remote attendees, and enabling effective interaction in group conversations with multiple participants, as validated by user feedback and experimental results.

Implementation Method 1

estimating, via one or more hardware processors comprised in the state representation model, direction of arrivals (DOAs) of voice activity captured continuously via a microphone array, through sound source localization (SSL)

Methodology Applied
Scientific EffectSound source localization: Sound

Implementation Method 2

the DOAs are triggered by Voice Activity Detection (VAD) and are estimated based on Time Difference of Arrival (TDOA) associated with the captured voice activity

Methodology Applied
Scientific EffectTime difference of arrival: Time of Flight

Data Source

PatentEP3797938B1Attention shifting of a robot in a group conversation using audio-visual perception based speaker localization
Publication Date: 2024.01.03 TATA CONSULTANCY SERVICES LTD
  • EP3797938B1 patent drawingFigure 1
  • EP3797938B1 patent drawingFigure 2A
  • EP3797938B1 patent drawingFigure 2B

AI summary

This disclosure relates to attention shifting of a robot in a group conversation with two or more attendees, wherein at least one of them is a speaker. State of the art has dealt with several aspects of Human-Robot Interaction (HRI) including responding to a source of sound at a time, addressing a fixed viewing area or determining who is the speaker based on eye gaze direction. However, attention shifting to make the conversation human-like is a challenge. The present disclosure uses audio-visual perception for speaker localization. Only qualified direction of arrivals (DOAs) are used for the audio perception. Further the audio perception is complimented by visual perception employing real time face detection and lip movement detection. Use of HRI rules, clustering of the DOAs, dynamic adjustment of rotation of the robot and a dynamically updated knowledge repository enriches the robot with intelligence to shift attention with minimum human intervention.