Whole-Body AR Stylization Without Depth Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AR systems require depth sensors to modify images, increasing device cost and complexity, struggle to recognize and stylize whole bodies, and inefficiently apply visual effects due to processing complexities and power requirements, especially on mobile devices.
Innovation Solution
A machine learning model estimates a stylized version of a whole body in real-time without depth sensors, applying visual effects efficiently by training on synthetically rendered images to generate a target style, reducing processing complexities and power requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If depth sensors are used to modify images in AR systems, then image modification capability is improved, but device cost and complexity increase
Solution Approach 1:
The patent extracts the depth sensing requirement from the AR system by using machine learning models that can process images without depth sensor input. The system uses a camera-only approach with neural networks to achieve body segmentation and stylization, removing the need for additional depth sensors while maintaining image modification capability.
Solution Approach 2:
The patent replaces the mechanical depth sensing system with a computational approach using machine learning models. Instead of using physical depth sensors to capture spatial information, the system uses neural networks trained on image data to infer depth and apply stylization effects, substituting mechanical sensing with computational intelligence.
2Reliability
If existing AR systems process images for body stylization, then visual effects are applied, but processing complexity and power requirements increase
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models on large datasets of stylized images before actual use. The neural networks are trained in advance to recognize body parts, poses, and stylization patterns, so that during real-time operation the system only needs to infer from input images without performing complex processing, significantly reducing power requirements during execution.
Solution Approach 2:
The patent changes the processing parameters by using optimized neural network architectures and data augmentation techniques during training to improve model efficiency. The models are trained with augmented data that simulates various lighting conditions, angles, and body poses, allowing the system to handle diverse real-world images with reduced computational effort during inference.
3Measurement precision
If machine learning models are trained on large datasets for body stylization, then stylization accuracy is improved, but training time and computational resources increase
Solution Approach 1:
The patent applies preliminary action by performing all heavy computational work during the training phase before deployment. Large datasets are processed and models are trained in advance using high-performance computing clusters, so that the actual mobile device only needs to execute the pre-trained model inference, which is much faster and less resource-intensive.
Solution Approach 2:
The patent uses copying by creating synthetic training data through image generation and transformation techniques. Instead of requiring actual photographs of diverse bodies for training, the system generates synthetic images and stylized variations through neural networks, reducing the need for manual data collection and processing while maintaining training accuracy.
Data Source
AI summary
Methods and systems are disclosed for performing real-time stylizing operations. The system receives an image that includes a depiction of a whole body of a real-world person. The system applies a machine learning model to the image to generate a stylized version of the whole body of the real-world person corresponding to a given style, the machine learning model being trained using training data to establish a relationship between a plurality of training images depicting synthetically rendered whole bodies of persons and corresponding ground-truth stylized versions of the whole bodies of the persons of the given style. The system replaces the depiction of the whole body of the real-world person in the image with the generated stylized version of the whole body of the real-world person.


