Robot Vision Search Using A Posteriori Object Location Knowledge
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional robots are inefficient in locating and identifying objects not directly in view, requiring time-consuming simultaneous localization and mapping (SLAM) operations, which consume resources and can be disruptive, unlike humans who use posteriori knowledge to efficiently find objects in predictable locations.
Innovation Solution
A machine learning model, such as a convolutional neural network, is trained to generate output indicating the likely locations of objects of interest based on visual data, allowing robots to quickly locate objects by maneuvering to positions that provide a direct view of concealed objects, reducing the need for exhaustive environmental mapping.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional robots perform exhaustive SLAM operations to locate objects not directly in view, then they can eventually find the objects, but they consume excessive resources (power, processing cycles, memory) and time
Solution Approach 1:
The system performs preliminary action by training a machine learning model in advance with a posteriori knowledge about object locations. This pre-trained model enables the robot to quickly predict object locations without performing exhaustive SLAM operations during actual search tasks, thereby resolving the contradiction between detection precision and search time
Solution Approach 2:
The patent replaces the mechanical exhaustive scanning system (SLAM operations) with an intelligent prediction system based on machine learning. The trained model substitutes the resource-intensive mechanical search process, allowing the robot to locate objects efficiently without consuming excessive time and resources
2Loss of information
If conventional robots perform exhaustive SLAM operations to map the environment, then they gain complete knowledge of object locations, but they expend excessive resources (power, processing cycles, memory)
Solution Approach 1:
The system extracts only the necessary information (a posteriori knowledge about where objects are typically located) from complete environmental mapping. Instead of performing full SLAM operations to map the entire environment, the robot uses the trained model to extract and utilize only the relevant location patterns, thereby reducing power consumption while maintaining effective environmental knowledge
Solution Approach 2:
The machine learning model is pre-trained with environmental knowledge in advance, allowing the robot to access this information without performing exhaustive mapping operations during actual tasks. This preliminary preparation stores environmental knowledge in an easily accessible format, eliminating the need for continuous resource-intensive SLAM operations
3Measurement precision
If conventional robots perform exhaustive environmental mapping to locate objects, then they can find objects not directly in view, but the operations are disruptive to the environment
Solution Approach 1:
The patent replaces the disruptive mechanical exhaustive mapping system with a non-intrusive machine learning-based prediction system. The trained model allows the robot to predict object locations without performing disruptive SLAM operations, thereby maintaining measurement precision while eliminating environmental disruption
Data Source
AI summary
Techniques described herein relate to generating a posteriori knowledge about where objects are typically located within environments to improve object location. In various implementations, output from vision sensor(s) of a robot may include visual frame(s) that capture at least a portion of an environment in which a robot operates/will operate. The visual frame(s) may be applied as input across a machine learning model to generate output that identifies potential location(s) of an object of interest. The robot's position/pose may be altered based on the output to relocate one or more of the vision sensors. One or more subsequent visual frames that capture at least a not-previously-captured portion of the environment may be applied as input across the machine learning model to generate subsequent output identifying the object of interest. The robot may perform task(s) that relate to the object of interest.


