Mixed Reality Avatar Voice Gesture Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer interfaces are cumbersome and lack natural interaction, requiring users to rely on mechanical inputs like typing and clicking, and often use passwords that are difficult to remember and insecure, failing to provide a human-like interaction experience.
Innovation Solution
A system and method for mixed reality interactions using a virtual avatar that responds to voice and gestural inputs, allowing users to interact naturally through speech and gestures, with facial recognition for authentication and personalized responses, enabling simultaneous participation of multiple users in an organized manner.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional mechanical interfaces (keyboard, mouse, touchscreen) are used for interaction, then information can be input and output, but user effort and time consumption increase significantly
Solution Approach 1:
The patent replaces mechanical input devices (keyboard, mouse, touchscreen) with voice-based acoustic input and gesture-based visual input. The system substitutes mechanical interaction with natural human behaviors - speaking and gesturing - thereby reducing physical effort and time required for information input and processing
Solution Approach 2:
The system changes the interaction paradigm from mechanical parameters (key presses, mouse movements) to biological parameters (voice frequency, gesture position). By transforming the input modality from mechanical to natural human expressions, the system improves ease of operation while reducing time loss
2Reliability
If passwords are used for authentication, then security can be implemented, but usability deteriorates due to difficulty in remembering and security risks
Solution Approach 1:
The patent replaces the mechanical/password-based authentication system with biometric authentication using voice recognition and facial recognition. This substitution eliminates the need for users to remember complex passwords while providing more secure and reliable authentication through unique biological characteristics
Solution Approach 2:
The system uses the user's own biological characteristics (voice, facial features) as the authentication key. The user authenticates themselves naturally without external assistance or memorization, improving both security and usability simultaneously
3Loss of information
If static text or image output is used, then information can be presented, but user effort to assimilate information increases
Solution Approach 1:
The patent transforms the output parameter from static visual text/images to dynamic audio-visual content. By presenting information through multiple sensory channels (heard and seen simultaneously), the system reduces the effort required for information assimilation while maintaining complete information delivery
Solution Approach 2:
The system adds the audio dimension to the traditional visual output. Information is delivered not only through visual text and images but also through spoken audio, creating a multi-dimensional presentation that reduces cognitive load and assimilation effort
4Productivity
If conventional computer interfaces are used, then data processing can be performed, but natural interaction is lost
Solution Approach 1:
The patent substitutes mechanical interface interactions with natural human communication methods. Users can process data and interact with the system through voice commands and gestures, maintaining full data processing capability while restoring natural interaction patterns
Solution Approach 2:
The system integrates multiple interaction modalities (voice, gesture, touch) into a single universal interface. This multi-functional approach allows users to choose the most natural method for each task while maintaining full access to data processing capabilities
Data Source
AI summary
A method (200) for mixed reality interactions with avatar, comprises steps of receiving (210) one or more of an audio input through a microphone (104) and a visual input through a camera (106), displaying (220) one or more avatars (110) that interact with a user through one or more of an audio outputted from one or more speakers (112) and a video outputted from a display device (108) and receiving (230) one or more of a further audio input through the microphone (104) and a further visual input through the camera (106). Further, a system (600) for mixed reality interactions with avatar has also been provided.


