Unsupervised Discretized Motion Model for Human Sequence Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image generation systems face inefficiencies, lack flexibility, and struggle with accurately representing human motions with significant variability in body types and motion states, requiring manual labeling and being limited to specific motion segments.

Innovation Solution

The system employs an unsupervised learning approach using a discretized motion model with an encoder-decoder architecture to extract and reconstruct human motion sequences from unlabeled digital scenes, leveraging a codebook of discretized feature representations and combining reconstruction and distribution losses for parameter learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional image generation systems are used, then manual labeling and specific motion segments can be processed, but efficiency is reduced and flexibility is limited

Engineering Contradiction:
ImproveefficiencyVSAvoidmanual labeling requirement
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system performs unsupervised learning where the neural network automatically learns to encode and decode motion sequences without requiring manual labels or human intervention during the learning process. The model self-adjusts parameters through reconstruction loss and distribution loss to capture motion patterns autonomously.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual labeling mechanisms with an automated neural network-based system. The encoder-decoder architecture with codebook discretization automatically processes and labels motion sequences through learned representations, substituting mechanical manual annotation with computational automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If conventional systems process specific motion segments, then accuracy for those segments is maintained, but adaptability to varied body types and motion states is reduced

Engineering Contradiction:
ImproveflexibilityVSAvoidaccuracy representation
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The neural network encoder-decoder system is designed to universally process diverse motion sequences featuring different body types, ages, genders, and motion states. The codebook of discretized feature representations serves as a universal vocabulary that can represent any motion pattern, making the system adaptable to various motion scenarios while maintaining accuracy through reconstruction loss optimization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If manual labeling is used, then specific motion segments can be accurately processed, but device complexity and time consumption increase

Engineering Contradiction:
Improvetime consumptionVSAvoidmanual intervention requirement
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The system performs preliminary action by pre-training the neural network on large datasets of motion sequences, learning general motion patterns and representations in advance. This pre-learning enables the model to process new motion sequences quickly without requiring time-consuming manual labeling during actual application.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If conventional image generation systems are used, then specific motion segments can be generated, but scalability is limited

Engineering Contradiction:
ImprovescalabilityVSAvoidsystem architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the motion representation into discrete codebook entries, where each entry represents a specific motion feature or pattern. This segmentation allows the system to scale by simply adding more codebook entries or expanding the codebook size, enabling representation of increasingly complex and diverse motion patterns without proportionally increasing system complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240346737A1Generating human motion sequences utilizing unsupervised learning of discretized features via a neural network encoder-decoder
Publication Date: 2024.10.17 ADOBE INC
  • US20240346737A1 patent drawing
  • US20240346737A1 patent drawing
  • US20240346737A1 patent drawing

AI summary

Methods, systems, and non-transitory computer readable storage media are disclosed for utilizing unsupervised learning of discrete human motions to generate digital human motion sequences. The disclosed system utilizes an encoder of a discretized motion model to extract a sequence of latent feature representations from a human motion sequence in an unlabeled digital scene. The disclosed system also determines sampling probabilities from the sequence of latent feature representations in connection with a codebook of discretized feature representations associated with human motions. The disclosed system converts the sequence of latent feature representations into a sequence of discretized feature representations by sampling from the codebook based on the sampling probabilities. Additionally, the disclosed system utilizes a decoder to reconstruct a human motion sequence from the sequence of discretized feature representations. The disclosed system also utilizes a reconstruction loss and a distribution loss to learn parameters of the discretized motion model.