Motion Capture Intent Recognition via Depth Camera Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current motion capture systems require explicit user actions for login and logout, disrupting the natural interaction experience and lacking seamless continuity when re-engaging with applications.

Innovation Solution

A processor-implemented method that uses depth camera systems to track a user's location, movement, and voice data to automatically recognize intent to engage or disengage with applications, generating avatars and profiles for seamless interaction without manual inputs, allowing users to walk in and out of experiences without interruption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual login and logout procedures are required, then user identification and session management are ensured, but user interaction becomes cumbersome and discontinuous

Engineering Contradiction:
Improveuser interactionVSAvoidlogin/logout time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically detecting user presence through motion capture technology before interaction is needed. The skeletal model tracking and intent recognition systems are pre-configured to identify users as they enter the field of view, eliminating the need for manual login procedures when the user actually wants to interact with the application.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The motion capture system provides self-service functionality by automatically performing user identification, profile association, and session management without requiring manual input. The system uses the user's own body movements and skeletal data to trigger login/logout events, allowing users to simply walk in and out of the interaction zone to control their session state.

Inventive Principle:
Principle #25Self-service

2Extent of automation

If automatic intent recognition is implemented, then user engagement detection becomes seamless, but system complexity increases

Engineering Contradiction:
Improveintent recognitionVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces a skeletal model as an intermediary representation between the physical user and the digital application. The motion capture system converts complex user movements into simplified skeletal joint data, which then serves as the basis for intent recognition. This intermediary layer abstracts the complexity, allowing automated decision-making about user engagement without requiring the system to directly interpret raw motion data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If motion capture tracking is used to detect user intent, then natural interaction is enabled, but measurement and detection difficulty increases

Engineering Contradiction:
Improveinteraction naturalnessVSAvoiduser intent detection
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system applies local quality analysis by focusing intent detection on specific key points of the skeletal model rather than analyzing the entire body movement. Critical joints such as the head, torso, and hands are monitored with higher priority to detect intent signals like facing the display or raising hands, while less critical areas receive minimal attention. This localized approach simplifies the detection process while maintaining natural interaction capabilities.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP2524350B1Recognizing user intent in motion capture system
Publication Date: 2016.05.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP2524350B1 patent drawingFigure 1a
  • EP2524350B1 patent drawingFigure 1b
  • EP2524350B1 patent drawingFigure 2

AI summary

Techniques for facilitating interaction with an application in a motion capture system allow a person to easily begin interacting without manual setup. A depth camera system tracks a person in physical space and evaluates the person's intent to engage with the application. Factors such as location, stance, movement and voice data can be evaluated. Absolute location in a field of view of the depth camera, and location relative to another person, can be evaluated. Stance can include facing a depth camera, indicating a willingness to interact. Movements can include moving toward or away from a central area in the physical space, walking through the field of view, and movements which occur while standing generally in one location, such as moving one's arms around, gesturing, or shifting weight from one foot to another. Voice data can include volume as well as words which are detected by speech recognition.