Machine Learning System for Augmenting Images with Contextual Pose and Shape
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image manipulation technologies require specialized devices for estimating 3D human body shape and lack contextual information, such as clothing and lighting conditions, which limits their ability to generate semantically meaningful augmented images.
Innovation Solution
An artificially intelligent machine learning system that identifies human body pose, shape, and environmental parameters in images, using deep learning models to generate augmented images that are contextually appropriate and realistic, without the need for human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If specialized devices are used for estimating 3D human body shape, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent uses 2D image copies from standard cameras to create 3D body shape models, avoiding the need for specialized 3D scanning devices. The machine learning system processes multiple 2D images to infer 3D structure, effectively copying information from simple 2D sources to generate accurate 3D representations.
Solution Approach 2:
The patent replaces mechanical 3D scanning systems with a machine learning-based computational approach. Instead of using specialized optical devices and mechanical sensors, the system uses deep learning models to estimate 3D body shape from 2D images, substituting mechanical complexity with algorithmic processing.
2Loss of information
If specialized scanning systems are used, then measurement precision is improved, but loss of information decreases, but device complexity and cost increase
Solution Approach 1:
The patent makes a single standard camera perform multiple functions: capturing 2D images, providing depth information through machine learning, detecting clothing presence, and inferring lighting conditions. This universal approach eliminates the need for separate specialized sensors for each type of information.
Solution Approach 2:
The system copies contextual information (clothing, lighting, pose) from 2D images using machine learning models, rather than requiring specialized sensors to directly measure these properties. The AI system infers this information by analyzing patterns in standard 2D photographs.
3Reliability
If multiple specialized sensors are used to capture contextual information, then reliability is improved, but device complexity increases
Solution Approach 1:
The patent employs a unified machine learning framework that processes 2D images to simultaneously extract body shape, pose, clothing, and lighting information. This single multi-functional system replaces multiple specialized sensors, maintaining reliability through comprehensive AI analysis rather than hardware redundancy.
4Manufacturing precision
If expensive scanning systems are used, then manufacturing precision is improved, but ease of manufacture worsens
Solution Approach 1:
The patent replaces expensive, complex scanning systems with inexpensive standard cameras and software-based machine learning models. The solution uses affordable 2D image capture devices combined with computational algorithms to achieve high-precision body shape modeling, making the technology accessible and easy to deploy.
Solution Approach 2:
The patent substitutes mechanical scanning hardware with software-based image processing and machine learning. This replacement dramatically reduces manufacturing complexity and cost while maintaining or improving measurement accuracy, as the solution relies on algorithmic processing rather than precision mechanical components.
Data Source
AI summary
Disclosed is a method including receiving visual input comprising a human within a scene, detecting a pose associated with the human using a trained machine learning model that detects human poses to yield a first output, estimating a shape (and optionally a motion) associated with the human using a trained machine learning model associated that detects shape (and optionally motion) to yield a second output, recognizing the scene associated with the visual input using a trained convolutional neural network which determines information about the human and other objects in the scene to yield a third output, and augmenting reality within the scene by leveraging one or more of the first output, the second output, and the third output to place 2D and/or 3D graphics in the scene.


