Learning Apparatus for Joint Estimation in Occluded Scenes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for estimating joint information in images with occlusion regions, such as those found in crowded scenes, suffer from decreased accuracy, often leading to erroneous or omitted estimations, which hinder applications in computer graphics and animations.
Innovation Solution
A learning device and method that utilize time-series information generation and machine learning, specifically deep neural networks, to estimate depth and silhouette information, which are then used to accurately determine joint information even in occlusion regions by accounting for the relative positional relationships between subjects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If deep learning is used to estimate joint information in images with occlusion regions, then the technique can robustly estimate joint information even when multiple persons are captured, but the accuracy of estimation decreases in occlusion regions where persons overlap
Solution Approach 1:
The system performs preliminary actions by capturing multiple time-series images before the actual estimation is needed. By accumulating temporal information in advance, the system builds up contextual data about occluded regions from different time points, enabling more accurate reconstruction of joint information even when persons overlap in individual frames.
Solution Approach 2:
The invention utilizes continuous temporal information from multiple successive images to maintain estimation accuracy throughout the sequence. By processing time-series data continuously and leveraging the continuity of motion across frames, the system compensates for the information loss in occlusion regions through temporal coherence and motion consistency.
2Measurement precision
If joint information estimation is omitted in occlusion regions to avoid erroneous estimation, then accuracy is maintained, but the joint information becomes difficult to apply to production of computer graphics or animations
Solution Approach 1:
The invention converts the harmful effect of occlusion (which causes estimation errors) into a beneficial situation by utilizing the temporal context from multiple images. The occlusion problem is transformed into an opportunity to leverage motion continuity and contextual information from surrounding time points, enabling accurate estimation that would otherwise be impossible in single-frame approaches.
Solution Approach 2:
Time-series information acts as an intermediary that bridges the gap between individual image frames and accurate joint estimation. This temporal intermediary provides the additional context needed to infer joint positions in occluded regions by analyzing motion patterns and spatial relationships across successive images, thereby enabling both accuracy and applicability.
3Measurement precision
If motion capture technique is used to acquire joint information, then accurate joint information can be obtained, but it requires a person to wear a special suit and involves complicated calibration tasks
Solution Approach 1:
The system creates a computational copy of the motion capture process by using deep learning models trained on labeled data. Instead of requiring physical motion capture suits and complex calibration hardware, the invention uses learned patterns from training data to replicate the measurement function, achieving similar accuracy through software-based inference rather than physical measurement systems.
Solution Approach 2:
The invention replaces the mechanical motion capture system (with suits and markers) with a computational vision system based on deep learning. By substituting the mechanical measurement approach with neural network-based estimation trained on image data, the system eliminates the need for special equipment while maintaining measurement precision through learned spatial relationships and motion patterns.
Data Source
AI summary
A learning device includes: a time-series information generation unit that obtains a first image group including a plurality of successive time-series images including a reference image and generates first time-series information based on a difference between the reference image and each of the images in the first image group other than the reference image; and a first learning unit that performs machine learning using the reference image and the first time-series information, thereby obtaining a first learning result used for estimating depth information on a target image, which is an image to be processed, and silhouette information on a subject captured in the target image based on the target image and second time-series information generated from a second image group including a plurality of successive time-series images including the target image.


