Context-Aware Voice Interaction via External User Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition in public mobile robots is often hindered by noisy environments and user preferences for avoiding direct interaction, especially in crowded spaces, and there is a risk of disease transmission through voice interaction.

Innovation Solution

An electronic device equipped with a camera module, short-range communication, and a processor that identifies users and, if direct voice interaction is not feasible, uses external devices for voice recognition via a server, enabling interaction through external electronic devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If direct voice interaction is used in public spaces, then the robot can provide voice-based services, but voice recognition accuracy deteriorates due to noisy environments

Engineering Contradiction:
Improvevoice-based service capabilityVSAvoidvoice recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an external electronic device as an intermediary to receive and process user voice inputs. Instead of the robot directly receiving voice commands in noisy environments, the external device captures the voice and transmits it to the robot via short-range wireless communication, thereby isolating the robot from environmental noise and improving recognition accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The interaction system is segmented into multiple components: the robot, an external electronic device, and a voice recognition server. The voice capture function is separated from the robot and assigned to the external device, allowing the robot to focus on processing and execution while the external device handles the challenging task of voice capture in noisy environments.

Inventive Principle:
Principle #1Segmentation

2Extent of automation

If direct voice interaction is implemented, then user commands can be recognized, but user privacy and comfort deteriorate due to forced interaction in crowded spaces

Engineering Contradiction:
Improveautomatic voice command recognitionVSAvoiduser discomfort and privacy intrusion
Core Design Contradiction:
Extent of automationVSObject-affected harmful factors

Solution Approach 1:

The system dynamically adapts the interaction mode based on the environment. When the robot detects a noisy or crowded environment, it automatically switches from direct voice interaction to using an external electronic device for voice capture. This dynamic adjustment respects user comfort preferences while maintaining automated command recognition capability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The external electronic device serves as a mediator that users can voluntarily use for interaction. This intermediary approach gives users control over their interaction method, allowing them to use their own device's microphone (which they control) rather than being forced to use the robot's microphone in uncomfortable situations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If voice interaction is used in public places, then services can be provided, but disease transmission risk increases through direct voice contact

Engineering Contradiction:
Improveservice accessibilityVSAvoiddisease transmission risk
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The external electronic device acts as a buffer or intermediary layer between the user and the robot. Voice commands are captured by the user's own device and transmitted digitally, eliminating the need for close physical proximity and direct voice projection toward the robot, thereby reducing droplet transmission risk while maintaining service accessibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of operation

If the robot uses its own microphone for voice recognition, then direct interaction is possible, but interaction reliability deteriorates in noisy environments

Engineering Contradiction:
Improvedirect interaction capabilityVSAvoidinteraction reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system uses the external electronic device's microphone as an intermediary voice capture component. This intermediary microphone is positioned close to the user's mouth (in their own device), providing a cleaner voice signal that is then transmitted to the robot. This maintains the convenience of direct interaction while significantly improving reliability by avoiding the robot's distant and noise-exposed microphone.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12521887B2Electronic device for providing interaction on basis of user voice, and method therefor
Publication Date: 2026.01.13 SAMSUNG ELECTRONICS CO LTD
  • US12521887B2 patent drawing
  • US12521887B2 patent drawing
  • US12521887B2 patent drawing

AI summary

An electronic device can include a microphone; a camera module; a short-range communication module supporting short-range wireless communication; a communication module configured to communicate with a voice recognition server; a memory; and a processor. The processor may be configured to: identify whether an object accessing the electronic device is a user; determine whether a voice interaction condition is satisfied on the basis of context information; when the user's access is identified, if the voice interaction condition is satisfied, receive user voice from the microphone, and if the voice interaction condition is not satisfied, output external interaction information that enables an external electronic device to interact with the voice recognition server by using the short-range communication module; receive user voice analysis information from the voice recognition server by using the communication module; and perform at least one operation on the basis of the received user voice analysis information.