State Space Model for 3D Human Motion from Continuous Light Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for evaluating human motion from video streams struggle to adapt effectively to new frame rates during real-time processing of continuous video frames, impacting applications like pose estimation, mesh recovery, and action recognition.

Innovation Solution

The implementation of a state space model that processes 2D information from continuous time light signals, combined with discretization information about frame rates, to generate accurate 3D information about a user, enhancing adaptability without the need for retraining.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing techniques process video streams at fixed frame rates, then processing is simple, but adaptability to new frame rates requires retraining

Engineering Contradiction:
Improveadaptability to new frame ratesVSAvoidmodel retraining requirement
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by using discretization information (frame rate parameters) as inputs to the state space model, allowing the model to adapt to different frame rates through parameter variation rather than retraining. The model receives discretization information about the frame rate and adjusts its processing accordingly, enabling versatility across different video frame rates without changing the model architecture or requiring retraining.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If state space model processes continuous time light signals, then accuracy and efficiency improve, but processing complexity increases

Engineering Contradiction:
Improveaccuracy of 3D information generationVSAvoidmodel processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses discretization information as an intermediary element that bridges the continuous time light signals and the state space model processing. This intermediary parameter allows the model to handle continuous signals efficiently by converting temporal information into discrete representations that can be processed through the state space model, achieving accurate 3D information generation without excessive computational complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Speed

If real-time processing is performed at high frame rates, then processing speed increases, but computational resources consume more

Engineering Contradiction:
Improveprocessing speedVSAvoidcomputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent applies dynamics by making the processing approach adaptive to the actual frame rate through discretization information. The state space model dynamically adjusts its processing based on the input frame rate, allowing efficient real-time processing at varying speeds without consuming excessive computational resources. The model can operate effectively whether the input is at 30fps, 60fps, or other frame rates, optimizing resource usage according to the actual processing needs.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250148623A1Human motion understanding using state space models
Publication Date: 2025.05.08 APPLE INC
  • US20250148623A1 patent drawing
  • US20250148623A1 patent drawing
  • US20250148623A1 patent drawing

AI summary

Various implementations disclosed herein include devices, systems, and methods that generate 3-dimensional (3D) information related to a user from a continuous time light signal. For example, a process may obtain two-dimensional (2D) information corresponding to a continuous time light signal providing information about a user in a 3D environment. The 2D information may be based on frames comprising images capturing the continuous time light signal at one or more frame rates. The process may further obtain discretization information corresponding to the one or more frame rates. The process may further determine 3D information about the user by inputting the 2D information and the discretization information into a state space model. The state space model may be a continuous time learnable framework for mapping between continuous time 2D scalar inputs and continuous time scalar 3D outputs.