Multimodal Sound Device Interaction Using Voice and Location Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional interaction methods, such as voice-based interfaces, lack the ability to provide rich user experiences due to limitations in visual information usage, restricting user movement and orientation, whereas display-based interactions require a fixed orientation and location, limiting flexibility.

Innovation Solution

A multimodal interaction system that combines voice inputs with location and orientation information to enable more flexible user experiences, using a voice-based interface to receive commands and integrate additional data from peripheral devices like smartphones or smartwatches, allowing for various operations and content interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If display-based interaction is used, then visual information is provided to users, but user movement and orientation are restricted

Engineering Contradiction:
Improvevisual informationVSAvoiduser movement flexibility
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent combines voice-based interaction with display-based interaction into a multimodal system. The voice interface allows users to interact without being constrained by display orientation, while the display provides visual feedback. This merging of interaction modes resolves the contradiction by allowing visual information delivery while maintaining user movement freedom.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements multiple interaction modes (voice and display) within a single interface framework. The voice-based interface serves as a universal input method that works regardless of user position, while the display provides supplemental visual information. This multi-functionality allows the system to adapt to different user scenarios without restriction.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If voice-based interaction is used, then user movement freedom is improved, but visual information capability is lost

Engineering Contradiction:
Improveuser movement freedomVSAvoidvisual information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges voice-based interaction (which provides movement freedom) with display-based interaction (which provides visual information). The system processes voice inputs for commands while simultaneously utilizing the display for visual feedback, content display, and contextual information, thus resolving the information loss issue while maintaining movement freedom.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The display acts as an intermediary that compensates for the lack of visual information in pure voice-based interaction. It provides visual feedback, confirms voice commands, displays content, and offers contextual information, thereby mediating between the voice interface and the user's visual information needs.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If fixed orientation interaction is used, then interaction precision is improved, but interaction flexibility deteriorates

Engineering Contradiction:
Improveinteraction precisionVSAvoidinteraction flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts its interaction mode based on the situation. Voice-based interaction provides flexibility for any user position, while display-based interaction provides precision when the user is properly oriented. The system can switch between or combine these modes dynamically, resolving the contradiction between precision and flexibility.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11004452B2Method and system for multimodal interaction with sound device connected to network
Publication Date: 2021.05.11 NAVER CORP
  • US11004452B2 patent drawing
  • US11004452B2 patent drawing
  • US11004452B2 patent drawing

AI summary

A method and a system for multimodal interaction with a sound device connected to a network are provided. The method for multimodal interaction comprises the steps of: outputting audio information for playing content through a voice-based interface included in an electronic device; receiving a speaker's voice input associated with the outputted audio information through the voice-based interface; generating location information associated with the speaker's voice input; and determining an operation associated with the playing of the content by using the voice input and the location information associated with the voice input.