Robotic Visual Embeddings for Keyframe-Guided Task Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robotic vision systems rely on feature-based approaches for object detection, which are limited in accuracy when images lack distinctive features, and struggle to perform tasks consistently across varying orientations and locations.

Innovation Solution

The method involves capturing images and identifying keyframe pixels, using pixel-level descriptors that include RGB and depth information, to match current images with trained keyframes, allowing the robotic device to perform tasks based on these matches, even with changes in pose or location.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If feature-based approaches are used for object detection, then the system complexity is reduced, but the measurement precision deteriorates when images lack distinctive features

Engineering Contradiction:
Improvesystem complexityVSAvoidobject detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces the conventional feature-based mechanical approach with a pixel-based neural network approach. Instead of extracting and comparing discrete features, the system uses a trained neural network to directly compare pixel values and their contextual relationships, achieving higher precision without proportionally increasing system complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters used for comparison from hand-crafted features to raw pixel values processed through a neural network. This parameter transformation allows the system to capture subtle visual differences that feature-based methods miss, improving detection accuracy while the neural network handles the complexity automatically

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If feature-based approaches are used for task performance, then the ease of operation is improved, but the reliability deteriorates across varying orientations and locations

Engineering Contradiction:
Improvetask execution simplicityVSAvoidtask performance consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent adds contextual dimensionality to pixel comparison by analyzing not just individual pixel values but also the spatial relationships and patterns surrounding each pixel. This multi-dimensional approach allows the system to recognize objects and tasks reliably regardless of orientation or location, as the contextual patterns remain consistent even when individual pixel positions change

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent replaces traditional image processing mechanics with neural network-based pattern recognition. The neural network learns robust task representations that are invariant to orientation and location changes, maintaining reliability across varying conditions while keeping the operation interface simple for users

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If pixel-level descriptors with RGB and depth information are used, then the measurement precision is improved, but the loss of information increases due to larger data processing requirements

Engineering Contradiction:
Improvepixel matching accuracyVSAvoiddata processing overhead
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts only the most relevant pixel information by using a neural network to identify and process key pixels and their immediate contexts. Instead of processing all pixels uniformly, the system selectively focuses on discriminative regions, reducing the effective data volume while maintaining high precision in pixel matching

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary processing by pre-training the neural network on large datasets before deployment. This preliminary training embeds knowledge about which pixel features are most important, allowing the system to efficiently process only relevant information during actual operation, thereby reducing processing overhead while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11741701B2Autonomous task performance based on visual embeddings
Publication Date: 2023.08.29 TOYOTA JIDOSHA KK
  • US11741701B2 patent drawing
  • US11741701B2 patent drawing
  • US11741701B2 patent drawing

AI summary

A method for controlling a robotic device is presented. The method includes capturing an image corresponding to a current view of the robotic device. The method also includes identifying a keyframe image comprising a first set of pixels matching a second set of pixels of the image. The method further includes performing, by the robotic device, a task corresponding to the keyframe image.