Multimodal Sound Device Interaction Using Voice and Location Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional interaction methods, such as voice-based interfaces, lack the ability to provide rich user experiences due to limitations in visual information usage, restricting user movement and orientation, whereas display-based interactions require a fixed orientation and location, limiting flexibility.
Innovation Solution
A multimodal interaction system that combines voice inputs with location and orientation information to enable more flexible user experiences, using a voice-based interface to receive commands and integrate additional data from peripheral devices like smartphones or smartwatches, allowing for various operations and content interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If display-based interaction is used, then visual information is provided to users, but user movement and orientation are restricted
Solution Approach 1:
The patent combines voice-based interaction with display-based interaction into a multimodal system. The voice interface allows users to interact without being constrained by display orientation, while the display provides visual feedback. This merging of interaction modes resolves the contradiction by allowing visual information delivery while maintaining user movement freedom.
Solution Approach 2:
The system implements multiple interaction modes (voice and display) within a single interface framework. The voice-based interface serves as a universal input method that works regardless of user position, while the display provides supplemental visual information. This multi-functionality allows the system to adapt to different user scenarios without restriction.
2Ease of operation
If voice-based interaction is used, then user movement freedom is improved, but visual information capability is lost
Solution Approach 1:
The patent merges voice-based interaction (which provides movement freedom) with display-based interaction (which provides visual information). The system processes voice inputs for commands while simultaneously utilizing the display for visual feedback, content display, and contextual information, thus resolving the information loss issue while maintaining movement freedom.
Solution Approach 2:
The display acts as an intermediary that compensates for the lack of visual information in pure voice-based interaction. It provides visual feedback, confirms voice commands, displays content, and offers contextual information, thereby mediating between the voice interface and the user's visual information needs.
3Measurement precision
If fixed orientation interaction is used, then interaction precision is improved, but interaction flexibility deteriorates
Solution Approach 1:
The system dynamically adapts its interaction mode based on the situation. Voice-based interaction provides flexibility for any user position, while display-based interaction provides precision when the user is properly oriented. The system can switch between or combine these modes dynamically, resolving the contradiction between precision and flexibility.
Data Source
AI summary
A method and a system for multimodal interaction with a sound device connected to a network are provided. The method for multimodal interaction comprises the steps of: outputting audio information for playing content through a voice-based interface included in an electronic device; receiving a speaker's voice input associated with the outputted audio information through the voice-based interface; generating location information associated with the speaker's voice input; and determining an operation associated with the playing of the content by using the voice input and the location information associated with the voice input.


