Robot Action Control Using Fused Visual and Tactile Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing autonomous control systems for robots face challenges in accurately determining the posture of a target object and deciding on actions due to occlusion issues when using visual information alone, and current fusion of visual and tactile senses through machine learning is insufficient for accomplishing objective tasks.
Innovation Solution
An autonomous control system that acquires state, visual, and tactile data, generates compressed data using neural networks to combine and dimensionally reduce the data, and employs reinforcement learning to decide on robot actions, thereby overcoming occlusion and improving task accomplishment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual information alone is used to control robot actions, then the system complexity is low, but the robot cannot accurately determine object posture when occlusion occurs
Solution Approach 1:
The patent combines visual data from cameras and tactile data from sensors into a unified data structure that represents both visual and tactile information. This merging allows the robot to determine object posture accurately even when visual information is occluded, as tactile data provides complementary information about contact points and forces.
Solution Approach 2:
The patent introduces a compressed representation as an intermediary that transforms high-dimensional visual and tactile data into a lower-dimensional latent space. This compressed representation serves as a mediator that preserves essential information for posture determination while reducing computational complexity.
2Measurement precision
If visual and tactile data are combined by fusing and dimensionally compressing, then the object posture determination accuracy improves, but the data processing complexity increases
Solution Approach 1:
The patent segments the data processing into distinct components: visual data processing, tactile data processing, and fusion/compression processing. Each component handles specific types of data with appropriate processing methods, making the overall complex system more manageable and efficient.
Solution Approach 2:
The patent changes the dimensional parameters of the data by compressing high-dimensional visual and tactile data into a lower-dimensional latent representation. This parameter transformation reduces the complexity of subsequent processing while preserving the essential information needed for accurate posture determination.
3Productivity
If high-dimensional visual and tactile data are processed directly, then the information completeness is high, but the processing efficiency and task accomplishment speed decrease
Solution Approach 1:
The patent extracts the essential information from high-dimensional visual and tactile data by projecting it into a compressed latent representation. This extraction process removes redundant information while preserving the critical features needed for task accomplishment, thereby improving processing efficiency without significant information loss.
Solution Approach 2:
The patent transforms the data from high-dimensional space to lower-dimensional space through compression, changing the parameter dimensions while maintaining the essential information structure. This parameter transformation enables faster processing while preserving the information necessary for accurate robot control.
Data Source
AI summary
An autonomous control system includes an acquirer configured to acquire state data of a robot, visual data of the robot, and tactile data of the robot and a processor configured to decide on an action of the robot capable of accomplishing a task given to the robot on the basis of the state data, the visual data, and the tactile data. The processor generates first compressed data having a smaller number of dimensions than data obtained by combining the visual data and the tactile data by fusing and dimensionally compressing the visual data and the tactile data. The processor generates second compressed data having a smaller number of dimensions than the tactile data by dimensionally compressing the tactile data. The processor decides on the action on the basis of combined state data obtained by combining the state data, the first compressed data, and the second compressed data into one.


