User Intention Prediction via Spatial-Temporal Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users with limited mobility, such as those with quadriplegia, face challenges in transmitting their intentions to muscles, leading to cumbersome equipment setups and additional operation requirements for intention detection, particularly when using bio-signal detection methods.
Innovation Solution
A method that predicts user intention through image analysis using spatial and temporal information from a first-person point-of-view camera, eliminating the need for bio-signal sensors and additional operations, by employing a deep learning network to interpret images and generate driving signals for motion assistance devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If bio-signal detection methods are used to identify user intention, then intention detection accuracy is improved, but device complexity and ease of operation deteriorate due to cumbersome equipment and additional operations required
Solution Approach 1:
The patent replaces bio-signal detection methods (electrical/physiological measurement systems) with an image processing system that uses a camera to capture visual information of user actions. This substitution eliminates the need for complex bio-signal sensors, EEG/EMG equipment, and related processing hardware, thereby reducing device complexity while maintaining intention detection capability through visual analysis of user movements and gestures.
Solution Approach 2:
The patent creates a visual copy (image) of the user's physical actions through camera capture, and then processes this visual information to infer user intention. Instead of directly measuring physiological signals, the system captures and analyzes visual representations of user behavior, simplifying the detection mechanism while preserving the ability to accurately identify user intent through image analysis algorithms.
2Reliability
If bio-signal sensors and additional operations are used for intention detection, then intention identification reliability is improved, but ease of operation worsens due to additional user actions required
Solution Approach 1:
The patent enables the system to automatically detect and analyze user intention through passive image capture and processing, without requiring the user to perform additional operations such as pressing buttons, wearing sensors, or consciously activating detection modes. The camera continuously captures visual information, and the processing unit automatically analyzes it to identify user intent, making the system self-sufficient in detecting user needs without adding operational burden.
Solution Approach 2:
The patent replaces active user operations (button pressing, sensor attachment) with passive visual observation. The system substitutes mechanical/physiological measurement requirements with optical capture and image analysis, allowing users to naturally perform their intended actions while the system observes and interprets these actions through imaging, thereby maintaining reliability without compromising ease of operation.
3Ease of operation
If image analysis method is used to predict user intention, then ease of operation is improved by eliminating additional operations, but measurement precision must be maintained through effective spatial and temporal information processing
Solution Approach 1:
The patent segments the image processing task into distinct functional components: capturing spatial information (object position, user posture, gesture) and temporal information (motion sequences, action progression) from video frames. By dividing the complex analysis into separate spatial and temporal processing streams, the system can effectively extract meaningful features while maintaining operational simplicity, ensuring that intention prediction accuracy is preserved through structured information processing.
Solution Approach 2:
The patent performs preliminary processing of image data by extracting and organizing spatial and temporal features before final intention prediction. The system pre-processes captured images to identify key visual elements, motion patterns, and contextual information in advance, preparing structured data that facilitates accurate intention inference. This preliminary action ensures that when user intention needs to be predicted, the system already has processed information ready, maintaining both simplicity and accuracy.
Data Source
AI summary
A method for predicting the intention of a user through an image acquired by capturing the user includes: receiving an image acquired by capturing a user; and predicting the intention of the user for the next motion by using spatial information and temporal information about the user and a target object included in the image.


