Digital Person Pose Estimation Network Jitter Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy of pose estimation in digital person training is decreased due to jitter caused by occlusion of various body parts in the training data.

Innovation Solution

A digital person training method and system that involves obtaining training data, extracting human-body pose estimation data, and inputting position, speed, and acceleration estimation data into an optimized pose estimation network to calculate generation losses and update network parameters, thereby reducing jitter and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If motion video is used for training digital person, then action information can be extracted, but body parts are blocked causing estimation deviations and jitter

Engineering Contradiction:
Improveaction informationVSAvoidpose estimation accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent uses image data with pose labels as training copies instead of motion videos. These image-based training samples contain pose information without the occlusion problems of video, allowing the model to learn accurate pose estimation from clean, unobstructed images while still capturing action information through the pose labels.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts pose estimation data from image data by inputting images into a pose estimation network. This extraction process isolates the pose information from the potentially occluding background and other body parts present in video data, obtaining clean pose estimates that can be used for training without the jitter and estimation deviations caused by occlusion in motion videos.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of information

If motion video is used for training, then pose information can be obtained, but occlusion causes jitter during training process

Engineering Contradiction:
Improvepose informationVSAvoidtraining stability
Core Design Contradiction:
Loss of informationVSStability of the object's composition

Solution Approach 1:

The patent replaces motion video data with image data copies that have associated pose labels. These image-based copies provide stable, non-jittery pose information because they are static observations rather than dynamic video sequences prone to occlusion-induced jitter, thereby stabilizing the training process while preserving pose information.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent converts the limitation of image data (lack of temporal information) into a benefit by using pose-labeled images that directly provide ground truth pose estimates. This approach eliminates the harmful jitter and occlusion issues of video data while still enabling the model to learn pose information effectively through the labeled data.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Adaptability or versatility

If various body parts are included in training data, then comprehensive action information is obtained, but occlusion between parts reduces estimation accuracy

Engineering Contradiction:
Improveaction information completenessVSAvoidkey point position accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent uses image data copies with pose labels that represent complete body actions. These images capture comprehensive action information across all body parts while avoiding the occlusion problems of video, as the pose labels provide direct annotations for all key points regardless of whether parts are visible in the image.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces pose labels as intermediary information that mediates between the image data and the training process. These pose labels serve as ground truth that bridges the gap between visual input and training objectives, allowing the model to learn accurate key point positions without being affected by occlusion between body parts in the image.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250086826A1Digital person training method and system, and digital person driving system
Publication Date: 2025.03.13 NANJING SILICON INTELLIGENCE TECH CO LTD
  • US20250086826A1 patent drawing
  • US20250086826A1 patent drawing
  • US20250086826A1 patent drawing

AI summary

This application provides a digital person training method and system, and a digital person driving system. According to the method, human-body pose estimation data in training data is extracted, and the human-body pose estimation data is input into an optimized pose estimation network to obtain human-body pose optimization data. Generation losses of position optimization data and acceleration optimization data in the human-body pose optimization data are calculated based on a loss function of the optimized pose estimation network, so as to minimize errors between position estimation data and acceleration estimation data and a real value. In this way, the optimized pose estimation network is driven to update a network parameter to obtain an optimal driving model that is based on the optimized pose estimation network. The errors between the position estimation data and the acceleration estimation data and the real value are minimized.