Surface Normal Tensor for Pixel-Aligned AR Object Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current augmented reality systems that modify images without depth sensors struggle to accurately recognize and replace the background of a user's entire body, leading to poor image quality and failure in applying visual effects to multiple objects, as they rely on specialized techniques optimized for recognizing specific object portions like faces, which are inadequate for recognizing the entirety of the user or multiple objects in the image.
Innovation Solution
The system improves image processing by cropping out a portion of the image depicting a user's body and applying a machine learning model to estimate both segmentation and a surface normal tensor, allowing for the application of augmented reality effects only to the object without affecting the background, enabling more realistic AR displays and reducing system resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specialized techniques optimized for recognizing specific object portions (like faces) are used, then recognition accuracy for that specific portion is improved, but the ability to recognize the entirety of the user or multiple objects deteriorates
Solution Approach 1:
The patent applies a universal machine learning model that can recognize and process multiple types of objects (faces, bodies, clothing, accessories) within a single unified framework, eliminating the need for separate specialized techniques for each object type while maintaining high recognition accuracy across all object categories
Solution Approach 2:
The system segments the image into multiple object categories (user body, clothing items, accessories, background) and applies appropriate processing to each segment, enabling comprehensive recognition of the entire user and multiple objects simultaneously while maintaining specialized processing quality for each category
2Manufacturing precision
If depth sensors are used to improve AR effect accuracy, then image quality and AR effect realism are improved, but device cost and complexity increase
Solution Approach 1:
The patent replaces the mechanical depth sensing system with a computational approach using machine learning models that analyze 2D images to infer 3D surface properties, eliminating the need for physical depth sensors while achieving comparable or superior AR effect accuracy
Solution Approach 2:
The system changes the processing parameters by estimating surface normal tensors and depth information from 2D image data through machine learning, rather than directly capturing depth information, enabling accurate AR effects without additional hardware complexity
3Adaptability or versatility
If the entire image is processed to apply AR effects to all objects, then comprehensive AR effect application is improved, but system resource usage increases
Solution Approach 1:
The patent applies different processing qualities and levels of detail to different regions of the image based on their importance - full processing for user body and clothing, selective processing for accessories, and minimal processing for background, optimizing resource usage while maintaining comprehensive AR effect application
Solution Approach 2:
The system performs partial processing by focusing computational resources on key objects (user body, clothing) that require accurate AR effects, while using simplified processing for less critical elements, achieving comprehensive coverage with reduced overall resource consumption
Data Source
AI summary
Methods and systems are disclosed for performing operations for applying augmented reality elements to a person depicted in an image. The operations include receiving an image that includes data representing a depiction of a person; generating a segmentation of the data representing the person depicted in the image; extracting a portion of the image corresponding to the segmentation of the data representing the person depicted in the image; applying a machine learning model to the portion of the image to predict a surface normal tensor for the data representing the depiction of the person, the surface normal tensor representing surface normals of each pixel within the portion of the image; and applying one or more augmented reality (AR) elements to the image based on the surface normal tensor.


