Robot Action Planning via Visual Feature Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotics technologies face challenges in executing long-horizon manipulation tasks, which require multiple sequential subtasks, due to complexities such as robust error recovery, high accuracy at specific steps, and subtle variations in observations across subtasks.

Innovation Solution

A method and apparatus for action planning that generate a sequence of images for an action execution plan based on description information and a reference image, extract visual feature representations, and determine control information for each step, incorporating observed information and reference visual features to ensure accurate execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a robot executes long-horizon manipulation tasks involving multiple sequential subtasks, then the robot can complete complex tasks, but the complexity of error recovery and maintaining high accuracy at specific steps increases

Engineering Contradiction:
Improveability to complete long-horizon manipulation tasksVSAvoidcomplexity of error recovery and control
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments long-horizon manipulation tasks into multiple short-horizon subtasks, each corresponding to a generated image frame. The action planning is divided into sequential steps where each step processes one subtask independently, allowing for localized error recovery without affecting the entire task sequence. This segmentation enables the robot to handle complex tasks by breaking them down into manageable units with dedicated control strategies for each.

Inventive Principle:
Principle #1Segmentation

2Stability of the object's composition

If the robot predicts actions over multiple timesteps to complete tasks smoothly, then task completion stability improves, but the accuracy and adaptability to real-time observations decrease

Engineering Contradiction:
Improvetask completion stabilityVSAvoidaction prediction accuracy
Core Design Contradiction:
Stability of the object's compositionVSMeasurement precision

Solution Approach 1:

The patent generates a sequence of future image frames that represent preliminary planned actions before execution. These generated images serve as predictive visual goals that guide the robot through multiple timesteps, ensuring smooth and stable task completion. The preliminary action planning allows the robot to anticipate future states while maintaining the flexibility to adjust based on real-time observations during execution.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates a feedback mechanism where the robot compares real-time observations with the generated future image frames during execution. This feedback loop allows the system to detect deviations from the planned trajectory and make corrective adjustments, thereby maintaining both stability and accuracy throughout the long-horizon task execution.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the robot focuses on short-horizon manipulation tasks, then action prediction accuracy is high, but the ability to handle long-horizon tasks with multiple sequential subtasks is limited

Engineering Contradiction:
Improveaction prediction accuracyVSAvoidtask horizon duration
Core Design Contradiction:
Measurement precisionVSDuration of action of moving object

Solution Approach 1:

The patent extends the action planning from short-horizon to long-horizon by adding a temporal dimension through image sequence generation. Instead of predicting actions for a single timestep, the system generates a sequence of future images representing multiple timesteps, thereby extending the planning horizon while maintaining accuracy through the visual guidance provided by each generated frame.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Ease of manufacture

If the robot executes actions based on pre-generated image sequences, then task planning is comprehensive, but the adaptability to subtle variations in real-time observations is reduced

Engineering Contradiction:
Improvetask planning comprehensivenessVSAvoidadaptability to real-time observations
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic execution framework where the robot continuously adjusts its actions based on real-time observations compared against the pre-generated image sequence. The system is not rigidly bound to the planned trajectory but dynamically adapts by detecting deviations and making corrective adjustments, thereby maintaining both comprehensive planning and real-time adaptability throughout task execution.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250162150A1Action planning for robot control
Publication Date: 2025.05.22 BYTEDANCE TECHNOLOGY LTD
  • US20250162150A1 patent drawing
  • US20250162150A1 patent drawing
  • US20250162150A1 patent drawing

AI summary

Embodiments of the disclosure provide a solution for action planning. A method includes: generating a sequence of images for an action execution plan based on description information and a reference image related to an environment with an action executor located, the description information describing the action execution plan to be executed by the action executor; extracting a sequence of visual feature representations from the sequence of images, respectively; and for a respective visual feature representation of the sequence of visual feature representations, determining control information for controlling an action to be executed by the action executor in the environment to complete the action execution plan at least based on the respective visual feature representation, a reference visual feature representation prior to the respective visual feature representation in the sequence and observed information of the action executor in the environment during execution of a reference action prior to the action.