Robot Attention Shifting Using Audio-Visual Speaker Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Human-Robot Interaction systems face challenges in effectively shifting attention to multiple speakers in group conversations, particularly in telepresence applications, due to limited audio-visual perception and rigid camera adjustments, leading to a lack of sense of presence and belongingness for remote participants.
Innovation Solution
A processor-implemented method and system that uses sound source localization and audio-visual perception to estimate direction of arrivals of voice activity, generate clusters of qualified DOAs, and dynamically rotate the robot to focus on the current speaker, enhancing visual perception and interaction by continuously updating a knowledge repository and applying Human-Robot Interaction rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the robot uses audio-visual perception to track multiple speakers in real-time, then the sense of presence and belongingness for remote participants is improved, but the device complexity increases due to integration of multiple sensors and processing modules
Solution Approach 1:
The patent combines multiple independent modules (audio processing with microphone array, visual processing with camera, speaker tracking algorithms) into an integrated attention shifting system. The state representation model merges audio-visual data streams to jointly represent speaker positions and robot attention states, achieving coordinated multi-modal perception while managing system complexity through modular integration.
Solution Approach 2:
The robot system performs multiple functions simultaneously: it captures audio from multiple directions using a microphone array, processes visual information from cameras, estimates direction of arrivals for multiple speakers, tracks speaker positions over time, and dynamically adjusts its attention orientation. This multi-functional capability allows a single system to handle complex group conversation scenarios.
2Ease of operation
If the robot dynamically adjusts its orientation to follow active speakers in group conversations, then the interaction quality is improved, but the response time may be delayed due to continuous processing and clustering computations
Solution Approach 1:
The system continuously estimates direction of arrivals and maintains updated speaker position information in a knowledge repository even during periods of no active speech. This preliminary processing ensures that when a speaker becomes active, the robot can quickly retrieve pre-computed position data and respond without delay, rather than waiting for full processing cycles.
Solution Approach 2:
The robot's attention orientation is dynamically adjusted based on real-time speaker activity detection and position changes. The system adapts its response behavior by detecting voice activity and transitioning to tracking mode, allowing flexible response timing that balances processing accuracy with interaction responsiveness.
3Device complexity
If the robot uses a fixed camera adjustment mechanism, then the device complexity is reduced, but the ability to track multiple speakers and maintain visual contact is compromised
Solution Approach 1:
The camera system transitions from a fixed adjustment mechanism to a dynamic orientation system that actively tracks speakers. The camera orientation is continuously updated based on estimated speaker positions and robot attention direction, enabling the system to maintain visual contact with active speakers while keeping the overall device architecture relatively simple through software-controlled adjustment.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system improves user experience by accurately shifting attention to active speakers, increasing the sense of presence and belongingness for remote attendees, and enabling effective interaction in group conversations with multiple participants, as validated by user feedback and experimental results.
Implementation Method 1
estimating, via one or more hardware processors comprised in the state representation model, direction of arrivals (DOAs) of voice activity captured continuously via a microphone array, through sound source localization (SSL)
Implementation Method 2
the DOAs are triggered by Voice Activity Detection (VAD) and are estimated based on Time Difference of Arrival (TDOA) associated with the captured voice activity
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
This disclosure relates to attention shifting of a robot in a group conversation with two or more attendees, wherein at least one of them is a speaker. State of the art has dealt with several aspects of Human-Robot Interaction (HRI) including responding to a source of sound at a time, addressing a fixed viewing area or determining who is the speaker based on eye gaze direction. However, attention shifting to make the conversation human-like is a challenge. The present disclosure uses audio-visual perception for speaker localization. Only qualified direction of arrivals (DOAs) are used for the audio perception. Further the audio perception is complimented by visual perception employing real time face detection and lip movement detection. Use of HRI rules, clustering of the DOAs, dynamic adjustment of rotation of the robot and a dynamically updated knowledge repository enriches the robot with intelligence to shift attention with minimum human intervention.