Metric Learning Embeddings for Multi-Sensor Robot Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training artificial intelligence models to control complex dynamic mechanical systems, such as robots, is time-consuming and challenging due to the need for robot-specific learning and teaching periods to account for unique characteristics and environments.

Innovation Solution

A process involving obtaining data from multiple sensors, forming a training set by segmenting data by time and grouping segments across channels, and training a metric learning model to encode inputs as vectors in an embedding space with self-supervised learning, ensuring temporal, spatial, and tactile consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional machine learning models are trained to control dynamic mechanical systems, then the models can learn system characteristics and environment interactions, but the training process is time-consuming and requires extensive robot-specific learning periods

Engineering Contradiction:
Improvecontrol accuracyVSAvoidtraining period
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the control system into multiple independent metric learning models, each specialized for a specific sensor modality (e.g., one model for camera data, another for LIDAR, another for tactile sensors). This segmentation allows each model to be trained independently and efficiently on its specific data type, reducing the overall training time while maintaining comprehensive control capability across all sensor inputs

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the training approach by changing the objective function parameters to use metric learning with distance-based metrics (such as cosine similarity or Euclidean distance) instead of traditional supervised learning loss functions. This parameter change enables the system to learn effective representations with shorter training periods while maintaining control reliability

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multiple sensor channels are processed to capture comprehensive system state, then the control accuracy improves, but the data processing complexity and computational burden increase

Engineering Contradiction:
Improvestate estimation accuracyVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the multi-modal data processing into separate metric learning models for each sensor channel. Each model independently processes its specific modality (visual, depth, tactile, etc.) and produces a distance metric representation. This segmentation reduces processing complexity by avoiding the need to handle all sensor data jointly, while still achieving comprehensive state estimation through fusion of the independent metric representations

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces distance metric representations as intermediary variables between raw sensor data and control decisions. Each sensor modality is transformed into a distance metric space that captures essential features while reducing dimensionality. These intermediary metric representations serve as a simplified interface that reduces computational burden while preserving the information needed for accurate control

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12277483B2Spatio-temporal consistency embeddings from multiple observed modalities
Publication Date: 2025.04.15 SANCTUARY COGNITIVE SYST CORP
  • US12277483B2 patent drawing
  • US12277483B2 patent drawing
  • US12277483B2 patent drawing

AI summary

Provided is a process that includes obtaining data indicative of state of a dynamic mechanical system and an environment of the dynamic mechanical system, the data comprising a plurality of channels of data from a plurality of different sensors including a plurality of cameras and other sensors indicative of state of actuators of the dynamic mechanical system; forming a training set from the obtained data by segmenting the data by time and grouping segments from the different channels by time to form units of training data that span different channels among the plurality of channels; training a metric learning model to encode inputs corresponding to the plurality of channels as vectors in an embedding space with self-supervised learning based on the training set; and using the trained metric learning model to control the dynamic mechanical system or another dynamic mechanical system.