Robot Motion Planning With Multimodal Inputs And Iterative Denoising

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robot motion planning techniques relying on human guidance are physically demanding and limited to unimodal interaction, which can lead to user fatigue and restrict the richness of user inputs.

Innovation Solution

A computer-implemented method that receives multi-modal user inputs, generates estimated noise, iteratively denoises candidate motion plans, and selects a robot motion plan, allowing for flexible and less physically demanding user interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional robot motion planning techniques using human guidance are employed, then robot motion plans can be acquired and refined through human input, but users experience physical fatigue and performance degradation due to repetitive demonstrations

Engineering Contradiction:
Improverobot motion plan qualityVSAvoiduser physical burden
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent uses motion capture technology to record human demonstrations and creates digital copies of motion data that can be processed and reused multiple times without requiring repeated physical demonstrations. This allows the robot to learn from a single set of human demonstrations through automated processing pipelines

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces manual repetitive physical demonstrations with automated motion processing systems that use machine learning algorithms to extract motion patterns from recorded data, eliminating the need for users to repeatedly perform physically demanding tasks

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If unimodal interaction techniques are used for robot control, then the interaction interface is simple and straightforward, but the richness and variety of user inputs are limited

Engineering Contradiction:
Improveinteraction simplicityVSAvoiduser input variety
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent combines multiple input modalities including voice commands, gestures, and tactile interactions into a unified control system that processes and integrates inputs from different channels to generate comprehensive robot motion plans

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a multi-modal interaction framework where a single system can handle various types of user inputs (speech, gestures, touch) and translate them into robot control commands, making the interface adaptable to different user preferences and task requirements

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250249574A1Techniques for robot control using multi-modal user inputs
Publication Date: 2025.08.07 NVIDIA CORP
  • US20250249574A1 patent drawing
  • US20250249574A1 patent drawing
  • US20250249574A1 patent drawing

AI summary

Techniques for robot control using multi-modal user inputs include receiving one or more multi-modal inputs from a user, extracting a motion hint from the one or more multi-modal inputs, generating estimated noise based on a current motion scene for the robot, generating a plurality of candidate motion plans, iteratively denoising the plurality of candidate motion plans based on the estimated noise and the motion hint to generate a plurality of revised robot motion plans, selecting a robot motion plan from the plurality of revised robot motion plans, generating a robot trajectory from the selected robot motion plan; and commanding the robot to perform a first step of the robot trajectory.