3D Human Skeleton Image Segmentation via 2D Video Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image segmentation methods lack automatic segmentation capabilities, have poor robustness, and are costly due to reliance on two-dimensional approaches or require depth cameras, limiting their effectiveness, especially in outdoor conditions.

Innovation Solution

An image segmentation method that extracts both two-dimensional and three-dimensional skeleton estimations of a human skeleton from a video frame using neural networks, calculates errors, and adjusts node positions to minimize an error function, enabling accurate image segmentation with a monocular camera.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If deep learning-based image segmentation methods are used, then segmentation automation is improved, but robustness deteriorates due to reliance on two-dimensional images

Engineering Contradiction:
Improvesegmentation automationVSAvoidrobustness
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent transforms two-dimensional image data into three-dimensional skeleton representations by introducing temporal dimension (video sequences) and spatial depth estimation. The system constructs 3D human skeletons from 2D video frames using neural networks, enabling automatic segmentation while improving robustness through three-dimensional spatial reasoning that captures depth relationships and temporal motion patterns.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If three-dimensional skeleton methods based on depth cameras are used, then segmentation accuracy is improved, but application cost deteriorates

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidapplication cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual three-dimensional copy of the human skeleton from two-dimensional video images using neural network-based depth estimation. Instead of requiring physical depth cameras, the system synthesizes 3D skeleton data by processing sequences of 2D frames, thereby achieving accurate segmentation while eliminating the need for expensive depth-sensing hardware.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical/optical depth sensing system (depth camera) with a computational system using neural networks. The system substitutes physical depth measurement hardware with algorithmic depth estimation that processes 2D image sequences to infer three-dimensional skeleton positions, reducing hardware complexity and cost.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If conventional image segmentation methods are used, then application cost is reduced, but automation capability deteriorates requiring manual box-selection

Engineering Contradiction:
Improveapplication costVSAvoidautomation capability
Core Design Contradiction:
Device complexityVSExtent of automation

Solution Approach 1:

The patent enables the system to automatically perform segmentation tasks without manual intervention. The neural network-based three-dimensional skeleton extraction system autonomously identifies and segments human subjects from video sequences, eliminating the need for manual box-selection while maintaining cost-effectiveness by using standard 2D cameras and computational algorithms.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11367195B2Image segmentation method, image segmentation apparatus, image segmentation device
Publication Date: 2022.06.21 BEIJING BOE OPTOELECTRONCIS TECH CO LTD
  • US11367195B2 patent drawing
  • US11367195B2 patent drawing
  • US11367195B2 patent drawing

AI summary

An image segmentation method, an image segmentation apparatus, an image segmentation device are provided, the image segmentation method including: extracting, from a current frame of a video image, a skeleton two-dimensional estimation and a skeleton three-dimensional estimation of a human three-dimensional skeleton; obtaining a target three-dimensional skeleton based on the skeleton two-dimensional estimation and the skeleton three-dimensional estimation; implementing image segmentation based on the target three-dimensional skeleton; wherein the human three-dimensional skeleton has multiple nodes, and the video image is a two-dimensional image, when calculating the target three-dimensional skeleton in the current frame, by comprehensively considering the skeleton two-dimensional estimation and skeleton three-dimensional skeleton estimation of the human three-dimensional skeleton in the current frame, the accuracy and robustness of the obtained target three-dimensional skeleton can be improved, thereby improving the accuracy of image segmentation.