Audible-Visual Session Transfer for Sensitive Voice Interactions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-based virtual assistants face challenges in handling complex or sensitive interactions, as audible interfaces can be non-intuitive and may not suitably manage sensitive information, leading to user abandonment of interactions.
Innovation Solution
A system and method that determine sensitive or non-intuitive interactions during an audible session and map them to a visual interface, enabling the interaction to be conducted on a device with a visual output, such as a smartphone or tablet, using a server to assess interactions based on predefined criteria and push the visual interface to the appropriate device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice-based virtual assistants are used for all interactions, then the interface simplicity and hands-free operation are maintained, but the ability to handle complex or sensitive interactions deteriorates
Solution Approach 1:
The system enables a single voice assistant platform to perform multiple interaction modes by dynamically routing conversations to different interface types. The server determines whether to maintain audio-only mode or transition to visual interface based on conversation context, allowing the system to universally handle both simple and complex interactions through one unified platform.
Solution Approach 2:
The interface type dynamically changes during the conversation based on the determined complexity or sensitivity. The system transitions from static audio-only interface to dynamic multi-modal interface (audio+visual) when needed, allowing the interaction mode to adapt to the conversation requirements rather than being fixed.
2Ease of operation
If audible interface is used for sensitive interactions, then the hands-free operation is maintained, but the security and privacy of sensitive information deteriorates
Solution Approach 1:
The server acts as an intermediary that mediates between the user's voice input and the appropriate interface selection. It determines when sensitive information is being discussed and automatically routes the interaction to a visual interface, which provides a more secure environment for handling such information while maintaining the convenience of voice initiation.
3Device complexity
If voice assistant is used for complex interactions, then the sequential processing is maintained, but the user comprehension and correction ability deteriorates
Solution Approach 1:
The system transitions from one-dimensional sequential audio processing to two-dimensional multi-modal interaction by introducing visual display. This adds a spatial dimension where users can see the conversation context, options, and results simultaneously, enabling better comprehension and correction of complex interactions while maintaining the sequential processing logic.
Data Source
AI summary
Methods and systems for transferring a user session between at least two electronic devices are described. The user session is conducted as an audible session via an audible interface provided by a primarily audible first electronic device. Input data is received from the audible interface, wherein the input data causes the audible interface to progress through audible interface states. An interaction may be determined to be sensitive or non-intuitive based on a logic rule or based on tracking interactions in the user session. A current audible interface state is mapped to a visual interface state defined for a visual interface. The mapped visual interface state is pushed to a second electronic device having a visual output device for displaying the visual interface, to enable the user session to be continued as a visual session on the second electronic device.


