Learning Apparatus for Joint Estimation in Occluded Scenes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for estimating joint information in images with occlusion regions, such as those found in crowded scenes, suffer from decreased accuracy, often leading to erroneous or omitted estimations, which hinder applications in computer graphics and animations.

Innovation Solution

A learning device and method that utilize time-series information generation and machine learning, specifically deep neural networks, to estimate depth and silhouette information, which are then used to accurately determine joint information even in occlusion regions by accounting for the relative positional relationships between subjects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If deep learning is used to estimate joint information in images with occlusion regions, then the technique can robustly estimate joint information even when multiple persons are captured, but the accuracy of estimation decreases in occlusion regions where persons overlap

Engineering Contradiction:
Improveability to handle multiple persons in imageVSAvoidaccuracy of joint information estimation in occlusion region
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary actions by capturing multiple time-series images before the actual estimation is needed. By accumulating temporal information in advance, the system builds up contextual data about occluded regions from different time points, enabling more accurate reconstruction of joint information even when persons overlap in individual frames.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention utilizes continuous temporal information from multiple successive images to maintain estimation accuracy throughout the sequence. By processing time-series data continuously and leveraging the continuity of motion across frames, the system compensates for the information loss in occlusion regions through temporal coherence and motion consistency.

Inventive Principle:
Principle #20Continuity of useful action

2Measurement precision

If joint information estimation is omitted in occlusion regions to avoid erroneous estimation, then accuracy is maintained, but the joint information becomes difficult to apply to production of computer graphics or animations

Engineering Contradiction:
Improveaccuracy of joint information estimationVSAvoidapplicability to computer graphics and animations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The invention converts the harmful effect of occlusion (which causes estimation errors) into a beneficial situation by utilizing the temporal context from multiple images. The occlusion problem is transformed into an opportunity to leverage motion continuity and contextual information from surrounding time points, enabling accurate estimation that would otherwise be impossible in single-frame approaches.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

Time-series information acts as an intermediary that bridges the gap between individual image frames and accurate joint estimation. This temporal intermediary provides the additional context needed to infer joint positions in occluded regions by analyzing motion patterns and spatial relationships across successive images, thereby enabling both accuracy and applicability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If motion capture technique is used to acquire joint information, then accurate joint information can be obtained, but it requires a person to wear a special suit and involves complicated calibration tasks

Engineering Contradiction:
Improveaccuracy of joint informationVSAvoidcomplexity of measurement process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system creates a computational copy of the motion capture process by using deep learning models trained on labeled data. Instead of requiring physical motion capture suits and complex calibration hardware, the invention uses learned patterns from training data to replicate the measurement function, achieving similar accuracy through software-based inference rather than physical measurement systems.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The invention replaces the mechanical motion capture system (with suits and markers) with a computational vision system based on deep learning. By substituting the mechanical measurement approach with neural network-based estimation trained on image data, the system eliminates the need for special equipment while maintaining measurement precision through learned spatial relationships and motion patterns.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11423561B2Learning apparatus, estimation apparatus, learning method, estimation method, and computer programs
Publication Date: 2022.08.23 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11423561B2 patent drawing
  • US11423561B2 patent drawing
  • US11423561B2 patent drawing

AI summary

A learning device includes: a time-series information generation unit that obtains a first image group including a plurality of successive time-series images including a reference image and generates first time-series information based on a difference between the reference image and each of the images in the first image group other than the reference image; and a first learning unit that performs machine learning using the reference image and the first time-series information, thereby obtaining a first learning result used for estimating depth information on a target image, which is an image to be processed, and silhouette information on a subject captured in the target image based on the target image and second time-series information generated from a second image group including a plurality of successive time-series images including the target image.