Synthetic Training Data for Movement Recognition Using Virtual Actors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models struggle to accurately recognize and classify body positions and movements from various angles and perspectives, especially in real-world applications where users may not be centrally positioned or rotated, due to the limitations of traditional training data captured by physical cameras.
Innovation Solution
The use of virtual actors and virtual cameras in 3D environments to generate synthetic training data, allowing for a full range of scene variability and body positioning, which is then used to train machine learning models to recognize and classify movements from any angle or perspective.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple videos and images are captured to account for each variance in body positioning and orientation, then the machine learning model can recognize more scenarios, but the cost and time required for data capture and labeling increases significantly
Solution Approach 1:
The patent uses virtual actors (3D models, wireframes) to create synthetic training data instead of capturing images of real people. This copying approach generates unlimited variations of body positions and orientations without requiring additional physical photo sessions, thereby maintaining high recognition coverage while dramatically reducing data capture time and cost.
Solution Approach 2:
The patent transitions from 2D image capture to 3D virtual environment modeling. By creating virtual actors in 3D space and capturing them from multiple angles and perspectives within a virtual environment, the system generates comprehensive training data that covers all possible body orientations and positions without the constraints of physical photography.
2Adaptability or versatility
If traditional physical cameras are used to capture training data, then the data capture process is simple, but the camera is limited in what it captures and cannot provide full range of scene variability
Solution Approach 1:
Instead of using physical cameras to capture real scenes, the patent creates virtual copies of cameras (virtual cameras) within a 3D environment. These virtual cameras can be positioned anywhere and oriented in any direction to capture the virtual actor from any perspective, providing unlimited scene variability without the physical constraints of real camera equipment.
Solution Approach 2:
The patent introduces a 3D virtual environment as an intermediary between the training data generation process and the machine learning model. This virtual environment mediates by allowing virtual actors and virtual cameras to interact in a controlled digital space, generating diverse training data without requiring complex physical setups or multiple physical cameras.
3Measurement precision
If training data is captured with live actors in specific poses, then the data can be labeled correctly, but some variances may be missed causing the model to not recognize or misidentify body positions
Solution Approach 1:
The patent uses 3D virtual modeling to represent the actor, allowing precise control over body pose and orientation. This 3D approach enables the generation of training data covering the complete range of motion and all possible angles, ensuring both accurate labeling and comprehensive variance coverage that cannot be achieved with traditional 2D photo capture of live actors.
Solution Approach 2:
The patent pre-defines the complete range of body positions and orientations in the 3D virtual environment before generating training data. By systematically exploring all possible poses and angles in advance, the system ensures that no variances are missed and that the training data comprehensively covers all scenarios the machine learning model may encounter in real-world applications.
Data Source
AI summary
The present disclosure describes generating synthetic training data for a machine learning model to detect one or more movements. The training data comprises a virtual actor (e.g., a 3D model of a skeleton, a 3D model of human, or a wireframe) posed in a plurality of positions and one or more virtual cameras may be used to capture each of the plurality of positions. The images of each of the plurality of positions may be used as synthetic training data for a machine learning model. The machine learning model may be trained to recognize a repetitive motion being performed by a non-virtual actor (e.g., a human) and count repetitions and provide feedback with respect to form.


