Single Camera AR Touch Interaction via Joint Coordinate Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current augmented reality technologies face challenges in accurately rendering virtual objects on user hands and require multiple cameras for interaction, leading to inefficient real-time image processing and incorrect hand positioning.
Innovation Solution
An electronic apparatus with a single camera and processor uses a pre-trained learning model to estimate joint coordinates of the user's body, rendering virtual objects and generating AR images, allowing for real-time interaction and transparent display changes based on touch detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple cameras are used to capture user and space from various viewpoints, then interaction between user and virtual object is improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent segments the complex multi-camera system into a simplified single-camera system by using a pre-trained learning model to extract and process only the necessary features (joint coordinates) from the captured image, eliminating the need for multiple cameras while maintaining interaction capability
Solution Approach 2:
The patent replaces the mechanical approach of using multiple physical cameras with a computational approach using a single camera combined with a pre-trained learning model (convolutional neural network) to achieve the same functional outcome of capturing and processing spatial information
2Productivity
If multiple cameras are used for real-time image processing, then interaction between user and virtual object is improved, but processing power requirements and energy consumption increase
Solution Approach 1:
The patent extracts only the essential information (joint coordinates of user body) from the captured image using a pre-trained learning model, discarding unnecessary data processing steps that would be required in a multi-camera system, thereby reducing processing power requirements while maintaining real-time capability
Solution Approach 2:
The patent uses a pre-trained learning model that has been previously trained offline with大量 hand images and 3D coordinate data, so that during real-time operation, only inference is needed rather than full training and processing, enabling real-time performance with lower power consumption
3Speed
If virtual object is rendered without accurate depth perception, then rendering speed is improved, but positioning accuracy of virtual object on user hand deteriorates
Solution Approach 1:
The patent replaces complex depth sensing hardware with a computational depth estimation approach using a pre-trained convolutional neural network that infers 3D joint coordinates from 2D captured images, achieving accurate hand positioning without additional depth sensing equipment
Solution Approach 2:
The patent transforms the 2D image data from a single camera into 3D spatial information (joint coordinates with x, y, z positions) through the pre-trained learning model, enabling accurate depth perception and virtual object positioning on the user's hand while maintaining rendering speed
Data Source
AI summary
An electronic apparatus is provided. The electronic apparatus includes a display, a camera configured to capture a rear of the electronic apparatus facing a front of the electronic apparatus in which the display displays an image, and a processor configured to render a virtual object based on the image captured by the camera, based on a user body being detected from the captured image, estimate a plurality of joint coordinates with respect to the detected user body using a pre-trained learning model, generate an augmented reality image using the estimated plurality of joint coordinates, the rendered virtual object, and the captured image, and control the display to display the generated augmented reality image, wherein the processor is configured to identify whether the user body touches the virtual object based on the plurality of estimated joint coordinates, and change a transmittance of the virtual object based on the touch being identified.


