Multi-Modal 3D Point Cloud for Robotic Functional Part Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current robotic systems for manipulating objects are limited by their reliance on vision-based 3D point cloud representations, which fail to encode multi-modal information such as tactile and auditory feedback, making them unable to adapt to novel objects and requiring hard-coded target locations, thus restricting their ability to detect functional parts effectively.

Innovation Solution

A system that uses a combination of tactile and auditory sensory feedback to generate a 3D multi-modal representation of objects, allowing robots to detect functional parts by recording and annotating 3D locations of audio and tactile events during object manipulation, merging these with a 3D model point cloud to estimate the probability of detecting audio events, enabling the identification of functional elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If vision-based 3D point cloud representations are used, then the system can capture object appearance, but it fails to encode multi-modal information (tactile and auditory feedback)

Engineering Contradiction:
Improvemulti-modal information encodingVSAvoidsensory system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges vision-based 3D point cloud data with tactile sensor data and auditory sensor data into a unified multi-modal representation. The tactile events and auditory events are integrated with the visual 3D model to create a comprehensive object representation that encodes appearance, texture, and functional properties simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a universal 3D object representation that serves multiple functions: visual identification, tactile property detection, and auditory event localization. This multi-functional representation allows the robot to perceive objects through multiple sensory modalities within a single unified framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If hard-coded target locations are used, then the robot can perform predefined tasks, but it cannot adapt to novel objects

Engineering Contradiction:
Improveadaptability to novel objectsVSAvoidtask execution reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system uses tactile feedback and auditory feedback during object manipulation to identify functional parts. The robot manipulates the object while sensors detect tactile events (contact points) and auditory events (sounds produced), providing real-time feedback that enables the robot to learn and adapt to novel objects without hard-coded programming.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary manipulation actions to explore the object and gather sensory data before executing the actual task. By repeatedly manipulating the object and recording tactile and auditory events during exploration, the robot builds a multi-modal understanding of the object's functional parts that guides subsequent task execution.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If 2D representations with fixed frame of reference are used, then the system can detect simple objects, but it cannot handle objects with degrees of freedom in 3D space

Engineering Contradiction:
Improvehandling of 3D objects with degrees of freedomVSAvoidevent location precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transitions from 2D representations to 3D point cloud representations, adding spatial dimensions to the object model. This enables the system to accurately locate tactile and auditory events in three-dimensional space, accounting for the object's geometry and degrees of freedom while maintaining precise event location measurement.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9144905B1Device and method to identify functional parts of tools for robotic manipulation
Publication Date: 2015.09.29 HRL LAB
  • US9144905B1 patent drawing
  • US9144905B1 patent drawing
  • US9144905B1 patent drawing

AI summary

Described is system for identifying functional elements of objects for robotic manipulation. The system causes a robot to manipulate an object, the robot having an audio sensor and touch sensors. A three-dimensional (3D) location of each audio event that produces a response during manipulation of the object is recorded. Additionally, a 3D location of each tactile event that produces a response during manipulation of the object is recorded. A 3D location of each audio and tactile event in 3D space is then determined. A 3D audio point cloud and a 3D tactile point cloud are generated. Then, the 3D audio point cloud and the 3D tactile point cloud are registered with a 3D model point cloud of the object. Finally, an annotated 3D model point cloud of the object is generated, which encodes the location of a functional element of the object.