Intent-Based Voice Response Using Image Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing IVR systems rely on users to explicitly state their intent, limiting their functionality and user experience, as they can only recognize specific terms and fail to determine intent accurately without user input.
Innovation Solution
An electronic device method that uses a combination of voice input and image recognition to identify user intent, context, and usage characteristics, generating relevant voice responses by associating physical objects with intents, super-intents, and sub-intents, and selecting the most relevant voice prompts based on these determinations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If IVR systems use speech recognition to evaluate user queries, then the system can automatically respond to user input, but the system is limited to recognizing only specific pre-defined terms and cannot accurately determine user intent beyond those terms
Solution Approach 1:
The patent combines speech recognition with image recognition and contextual analysis to create a hybrid intent determination system. The IVR system integrates multiple data sources (audio, visual, contextual) to comprehensively understand user intent, moving beyond relying on a single recognition method.
Solution Approach 2:
The patent introduces an intermediary processing layer that analyzes both speech and image data to infer user intent. This intermediary layer processes the raw inputs (voice query and captured image) through contextual characteristics and usage history to determine the underlying user intent, which may not be explicitly stated in the voice input alone.
2Measurement precision
If IVR systems rely on users to explicitly state their intent in the query, then the system can provide accurate responses, but the user experience is compromised and the interaction becomes less natural
Solution Approach 1:
The system performs self-service by automatically determining user intent through analysis of speech, images, and contextual data without requiring the user to explicitly state their intent. The system serves itself by inferring intent from the interaction context, reducing the burden on the user while maintaining accurate intent determination.
Solution Approach 2:
The system uses feedback from multiple sources (voice input, image capture, contextual characteristics, usage history) to continuously refine and improve intent determination. By analyzing patterns across these feedback loops, the system becomes more accurate in understanding user intent over time without requiring explicit user confirmation.
3Ease of manufacture
If IVR systems use only pre-recorded responses, then the system is simple to implement, but the system cannot adapt to diverse user queries and contexts
Solution Approach 1:
The patent transforms the static pre-recorded response system into a dynamic response generation system. Instead of using fixed pre-recorded responses, the system dynamically generates appropriate responses by analyzing real-time inputs (voice, image, context) and selecting from multiple possible responses based on the determined intent and contextual characteristics.
Data Source
AI summary
A method for providing intent-based interactive voice response by an electronic device. The method includes receiving, by the electronic device, a voice input while obtaining an image of an object by using an image sensor, and generating an interactive voice response associated with the object based on the voice input. The method may further include determining a first intent and a second intent from the voice input, and generating an interactive voice response to the voice input, based on the first intent and the second intent.


