Telepresence Robot Internal State Expression Through Multimodal Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Telepresence robots lack non-verbal cues for expressing internal states, leading to confusion and frustration for teleoperators during interactions, particularly in turn-taking and understanding the robot's comprehension of instructions.

Innovation Solution

A method and system that utilize a combination of multiple modalities, including voice activity detection, wake word recognition, automatic speech recognition, natural language understanding, and a robot internal state predictor model, to express internal states through emotional expressions and text messages to the teleoperator.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If telepresence robots use spoken interaction for natural control, then ease of operation is improved, but the robot's ability to express internal states deteriorates due to lack of non-verbal cues

Engineering Contradiction:
Improveease of controlVSAvoidloss of non-verbal cues
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces a new dimension of communication by displaying the robot's internal states (attention, comprehension, processing) as visual information on a screen interface. This transforms the robot's internal cognitive states into observable visual cues, compensating for the absence of physical non-verbal cues in telepresence robots.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent uses an intermediary display interface that translates the robot's internal states into visual representations. This intermediary layer bridges the gap between the robot's internal processing and the teleoperator's understanding, making invisible cognitive states visible without requiring physical non-verbal capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If telepresence robots process multiple audio signals for wake word detection and instruction recognition, then functionality is improved, but device complexity increases

Engineering Contradiction:
ImprovefunctionalityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the audio processing functionality into distinct modules: wake word detection module, voice activity detection module, and instruction recognition module. Each module handles a specific aspect of audio processing independently, making the complex system more manageable and maintainable while providing versatile functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a multi-functional audio processing system where a single speech interface handles multiple tasks: wake word detection, voice activity detection, and instruction recognition. This universal approach consolidates multiple functions into one integrated system rather than requiring separate systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4607508A1Method and system for expressing telepresence robot internal states using combination of multiple modalities
Publication Date: 2025.08.27 TATA CONSULTANCY SERVICES LTD
  • EP4607508A1 patent drawingFigure 1
  • EP4607508A1 patent drawingFigure 2
  • EP4607508A1 patent drawingFigure 3

AI summary

State of art techniques relate to telepresence robots expressing internal states using emotional expressions and corresponding text messages are based on the perceived input from the co-located user or the participant, and not related to the teleoperator. Embodiments of the present disclosure provide a method and system for expressing telepresence robot internal states using combination of multiple modalities. The telepresence robot expresses its own internal states in the form of a plurality of emotional expressions and corresponding text messages to the teleoperator, using a robot internal state predictor model. Unlike the state of art techniques which express an emotion as its internal state, the disclosed method expresses the internal state as the emotional expression with respect to a task processing of the telepresence robot.