User Intent Detection via Image Pose Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual assistant systems require user-initiated conversations to determine intent, leading to missed opportunities for assistance, resource wastage, and poor insights in both physical and digital channels, resulting in lost business and incorrect recommendations.
Innovation Solution
A user device that determines user intent based on image-captured actions using a machine learning model, identifying users, pose points, and actions to generate intent identifiers and provide them to conversational AI services, enabling proactive assistance and resource conservation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the system waits for user-initiated conversations to determine intent, then the chatbot can respond accurately to user queries, but it misses opportunities to proactively assist users and wastes resources on inactive users
Solution Approach 1:
The system performs preliminary actions by capturing images of users and their actions, processing these images through machine learning models to identify pose points and determine user actions before users initiate conversations. This allows the system to proactively determine user intent and initiate assistance, rather than waiting passively for user input.
Solution Approach 2:
The patent replaces the mechanical system of waiting for user input (audio/text) with an automated vision-based system using cameras and machine learning models. The camera captures user actions, the ML model processes images to identify pose points and determine actions, and the system automatically maps these to intents, substituting passive waiting with active automated detection.
2Productivity
If the system continuously monitors users to enable proactive assistance, then user assistance timeliness improves, but computing and networking resources are wasted on users who do not require assistance
Solution Approach 1:
The system applies partial monitoring by focusing computational resources only on users who exhibit specific actions indicating potential assistance needs. Rather than continuously analyzing all user activities in detail, the system monitors for key pose points and actions that trigger intent determination, performing comprehensive analysis only when relevant actions are detected.
Solution Approach 2:
The system enables self-service by using automated machine learning models to process images and determine user actions without requiring continuous human supervision or intervention. The ML models autonomously identify pose points, determine actions, and map them to intents, reducing the need for extensive computing resources while maintaining effective monitoring.
3Adaptability or versatility
If the system uses image processing and machine learning to determine user intent, then proactive assistance capability improves, but the device complexity increases
Solution Approach 1:
The system segments the intent determination process into distinct modular components: image capture by camera, image processing to identify pose points using ML models, action determination based on pose points, and intent mapping. This segmentation allows each component to be independently optimized and managed, reducing overall system complexity while maintaining proactive assistance capability.
Data Source
AI summary
A device may process an image, with a machine learning model, to identify a user in the image and may generate a user identifier for the user. The device may process a portion of the image that includes the user, with the machine learning model, to identify pose points for the user. The device may calculate pose angles of the user based on the pose points and may determine an action of the user based on the pose angles. The device may associate the action with an action identifier, may map the action to an intent, and may generate an intent identifier for the intent. The device may provide the intent to a conversational service and may receive a response from the conversational service. The device may map the response to the intent based on the intent identifier and may provide the response to the user.


