Single Camera AR Touch Interaction via Joint Coordinate Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality technologies face challenges in accurately rendering virtual objects on user hands and require multiple cameras for interaction, leading to inefficient real-time image processing and incorrect hand positioning.

Innovation Solution

An electronic apparatus with a single camera and processor uses a pre-trained learning model to estimate joint coordinates of the user's body, rendering virtual objects and generating AR images, allowing for real-time interaction and transparent display changes based on touch detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple cameras are used to capture user and space from various viewpoints, then interaction between user and virtual object is improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improveinteraction capabilityVSAvoidcamera system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex multi-camera system into a simplified single-camera system by using a pre-trained learning model to extract and process only the necessary features (joint coordinates) from the captured image, eliminating the need for multiple cameras while maintaining interaction capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces the mechanical approach of using multiple physical cameras with a computational approach using a single camera combined with a pre-trained learning model (convolutional neural network) to achieve the same functional outcome of capturing and processing spatial information

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If multiple cameras are used for real-time image processing, then interaction between user and virtual object is improved, but processing power requirements and energy consumption increase

Engineering Contradiction:
Improvereal-time processing capabilityVSAvoidprocessing power requirement
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

The patent extracts only the essential information (joint coordinates of user body) from the captured image using a pre-trained learning model, discarding unnecessary data processing steps that would be required in a multi-camera system, thereby reducing processing power requirements while maintaining real-time capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses a pre-trained learning model that has been previously trained offline with大量 hand images and 3D coordinate data, so that during real-time operation, only inference is needed rather than full training and processing, enabling real-time performance with lower power consumption

Inventive Principle:
Principle #10Preliminary action

3Speed

If virtual object is rendered without accurate depth perception, then rendering speed is improved, but positioning accuracy of virtual object on user hand deteriorates

Engineering Contradiction:
Improverendering speedVSAvoidhand positioning accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent replaces complex depth sensing hardware with a computational depth estimation approach using a pre-trained convolutional neural network that infers 3D joint coordinates from 2D captured images, achieving accurate hand positioning without additional depth sensing equipment

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the 2D image data from a single camera into 3D spatial information (joint coordinates with x, y, z positions) through the pre-trained learning model, enabling accurate depth perception and virtual object positioning on the user's hand while maintaining rendering speed

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11514650B2Electronic apparatus and method for controlling thereof
Publication Date: 2022.11.29 SAMSUNG ELECTRONICS CO LTD
  • US11514650B2 patent drawing
  • US11514650B2 patent drawing
  • US11514650B2 patent drawing

AI summary

An electronic apparatus is provided. The electronic apparatus includes a display, a camera configured to capture a rear of the electronic apparatus facing a front of the electronic apparatus in which the display displays an image, and a processor configured to render a virtual object based on the image captured by the camera, based on a user body being detected from the captured image, estimate a plurality of joint coordinates with respect to the detected user body using a pre-trained learning model, generate an augmented reality image using the estimated plurality of joint coordinates, the rendered virtual object, and the captured image, and control the display to display the generated augmented reality image, wherein the processor is configured to identify whether the user body touches the virtual object based on the plurality of estimated joint coordinates, and change a transmittance of the virtual object based on the touch being identified.