Robot Attention Shifting with Audio-Visual DOA Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Human-Robot Interaction systems face challenges in effectively shifting attention to multiple speakers in group conversations, particularly in telepresence applications, due to limited audio-visual perception and inflexible camera adjustments, leading to a lack of sense of presence and belongingness for remote participants.
Innovation Solution
A processor-implemented method and system that uses sound source localization and audio-visual perception to estimate the direction of arrivals of voice activity, generate clusters of qualified DOAs, associate them with attendees, and dynamically rotate the robot to focus attention on the speaker, enhancing visual perception and interaction through a state representation model and HRI rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the robot uses audio-visual perception to track multiple speakers in group conversations, then the sense of presence and belongingness for remote participants is enhanced, but the device complexity increases due to multiple microphones and cameras required
Solution Approach 1:
The system segments the audio spectrum into multiple frequency bands using filter banks, allowing independent processing of different speaker frequencies. This segmentation enables the robot to distinguish and track multiple speakers simultaneously using a single microphone array, enhancing presence without proportionally increasing hardware complexity
Solution Approach 2:
The microphone array serves multiple functions: it performs both sound source localization and spectral analysis for multiple speakers. The same hardware infrastructure supports both audio tracking and visual perception through the camera, reducing the need for separate dedicated components and thereby managing device complexity
2Productivity
If the robot dynamically adjusts camera rotation to focus on active speakers in real-time, then the interaction effectiveness is improved, but the response time and processing delay increase
Solution Approach 1:
The system pre-processes audio signals through continuous Voice Activity Detection and maintains a ready-state camera control system. When a speaker becomes active, the camera rotation and focus adjustments are triggered immediately from a pre-prepared state, minimizing processing delay and maintaining real-time interaction effectiveness
Solution Approach 2:
The system implements continuous feedback loops where microphone arrays constantly monitor for voice activity, and camera positioning is adjusted based on real-time detection of active speakers. This closed-loop feedback ensures the robot responds promptly to speaker changes, maintaining interaction effectiveness while optimizing response time through continuous monitoring and immediate action
3Device complexity
If the robot uses a single microphone array for both sound source localization and spectral analysis, then the device complexity is reduced, but the measurement precision of speaker localization deteriorates
Solution Approach 1:
The system applies different processing qualities to different frequency bands of the audio signal. By analyzing specific frequency ranges where speakers operate and applying targeted spectral analysis, the system achieves accurate speaker localization and identification using a single microphone array, maintaining measurement precision without requiring multiple specialized microphones
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Improves user experience by accurately shifting attention to the active speaker in real-time, enhancing the sense of presence and belongingness for remote attendees, and facilitating more effective interaction in group conversations, with a User Experience/Satisfaction Index of 7 as measured in experimental scenarios.
Implementation Method 1
estimating, via one or more hardware processors comprised in the state representation model, direction of arrivals (DOAs) of voice activity captured continuously via a microphone array, through sound source localization (SSL) for the robot to respond in real-time, wherein the DOAs are triggered by Voice Activity Detection (VAD) and are estimated based on Time Difference of Arrival (TDOA) associated with the captured voice activity
Data Source
AI summary
This disclosure relates to attention shifting of a robot in a group conversation with two or more attendees, wherein at least one of them is a speaker. State of the art has dealt with several aspects of Human-Robot Interaction (HRI) including responding to a source of sound at a time, addressing a fixed viewing area or determining who is the speaker based on eye gaze direction. However, attention shifting to make the conversation human-like is a challenge. The present disclosure uses audio-visual perception for speaker localization. Only qualified direction of arrivals (DOAs) are used for the audio perception. Further the audio perception is complimented by visual perception employing real time face detection and lip movement detection. Use of HRI rules, clustering of the DOAs, dynamic adjustment of rotation of the robot and a dynamically updated knowledge repository enriches the robot with intelligence to shift attention with minimum human intervention.


