AR Control via Action Prediction and Feature Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality (AR) systems require significant user effort for action prediction, as they rely on manual control and may struggle to distinguish similar actions due to limited information from color images and optical flow, and are not suitable for real-time interactive scenarios.
Innovation Solution
A method and apparatus for controlling AR devices by acquiring video, detecting human bodies, performing action prediction using a combination of frame-based and video-based feature images, and mapping human body actions to AR functions, enabling automatic execution of AR functions based on predicted actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If action prediction is performed using only color image and optical flow information, then the system complexity is reduced, but the accuracy of distinguishing similar actions deteriorates
Solution Approach 1:
The patent segments the action recognition system into multiple independent modules: color image processing module, optical flow processing module, and human body part model module. Each module processes specific features separately, then results are fused for final action prediction. This segmentation allows the system to maintain low complexity in individual modules while achieving high accuracy through comprehensive feature fusion.
Solution Approach 2:
The patent combines multiple types of information (color images, optical flow data, and human body part models) into a composite feature representation. This composite approach integrates diverse data sources to improve action recognition accuracy while managing system complexity through modular architecture.
2Speed
If only a ROI containing a user is used for action prediction, then the processing speed is improved, but the sufficiency of human interaction and context information deteriorates
Solution Approach 1:
The patent extends the analysis from a single ROI dimension to multiple dimensions by incorporating spatial context around the user and temporal context through video sequences. This multi-dimensional approach captures both local user actions and global environmental context simultaneously, maintaining processing efficiency through optimized feature extraction.
3Measurement precision
If optical flow calculation is performed to improve action prediction accuracy, then the measurement precision is improved, but the time consumption deteriorates
Solution Approach 1:
The patent performs preliminary processing of video frames to extract key features before optical flow calculation. By pre-processing frames to identify regions of interest and extract salient features in advance, the system reduces the computational burden of subsequent optical flow calculations, thereby decreasing overall processing time while maintaining accuracy.
4Ease of operation
If manual control operations are required for AR functions, then the ease of operation deteriorates, but the extent of automation is reduced
Solution Approach 1:
The patent implements self-service functionality where the AR system automatically detects user actions, predicts intended operations, and executes corresponding AR functions without manual input. The system serves itself by interpreting user behavior and autonomously triggering appropriate functions, thereby improving ease of operation while maximizing automation.
Data Source
AI summary
A method and apparatus for controlling an augmented reality (AR) apparatus are provided. The method includes acquiring a video, detecting a human body from the acquired video, performing an action prediction with regard to the detected human body, and controlling the AR apparatus based on a result of the action prediction and a mapping relationship between human body actions and AR functions.


