Electronic Device 3D Pose Learning for Synthetic Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing motion capture methods for generating learning data in image recognition systems face physical limitations and high costs, particularly when photographing large or difficult-to-photograph objects, and require extensive time and resources.
Innovation Solution
An electronic device uses a learning network model trained with 3D human models to convert 2D images into 3D modeling images, allowing for the generation of learning data through virtual simulations, reducing the need for physical photography and minimizing time and cost.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a motion capture method using direct photography is used to generate learning data, then the recognition accuracy can be improved, but the cost and time consumption increase significantly
Solution Approach 1:
The patent uses pre-generated 3D modeling data as virtual copies of real objects to train the learning network model, eliminating the need for time-consuming direct photography of actual objects. The system creates synthetic training data by rendering 3D models in various poses and lighting conditions, which then serves as learning data for the AI model.
Solution Approach 2:
The patent performs preliminary actions by pre-generating comprehensive 3D modeling data covering multiple poses, angles, and conditions before the actual recognition task. This advance preparation of training data through virtual simulation reduces the time required during deployment, as the learning model is already trained on diverse synthetic data.
2Measurement precision
If a motion capture method using direct photography is used to generate learning data, then the recognition accuracy can be improved, but the cost increases significantly
Solution Approach 1:
The patent replaces expensive physical motion capture sessions with cost-effective virtual 3D model rendering. By using computer-generated 3D models and synthetic imaging, the system eliminates costs associated with hiring actors, equipment rental, and studio time, while maintaining high-quality training data through controlled virtual environments.
Solution Approach 2:
The patent uses inexpensive virtual 3D models as substitutes for expensive physical objects or actors. These digital models can be freely manipulated, reproduced, and modified without the costs associated with physical production, making the training data generation process economically efficient.
3Quantity of substance
If traditional motion capture methods are used, then learning data can be obtained, but physical limitations prevent photographing large or difficult-to-photograph objects
Solution Approach 1:
The patent creates virtual copies of objects in 3D space that can be photographed from any angle or scale without physical constraints. These digital models can represent objects of any size or complexity, allowing the generation of diverse learning data for objects that would be impractical or impossible to photograph in the real world.
Solution Approach 2:
The patent transitions from 2D physical photography to 3D virtual modeling, adding a dimensional aspect that overcomes physical limitations. By working in three-dimensional virtual space, the system can capture and analyze objects from any perspective, scale, or configuration without being constrained by real-world physical boundaries.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An electronic device is disclosed. The present electronic device comprises: a display; a processor electronically connected to the display so as to control the display; and a memory electronically connected to the processor, wherein the memory stores instructions causing the processor to control the display to display a 3D modeling image acquired by applying an input 2D image to a learning network model configured to convert the input 2D image into a 3D modeling image, and the learning network model is obtained by learning using a 3D pose acquired by rending virtual 3D modeling data and a 2D image corresponding to the 3D pose.