Reinforcement Learning Sensor Planning for Robotic Sampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing robotic sensing systems face challenges in capturing samples of unknown or changing areas of interest in sufficient detail and completeness while minimizing redundancy and time, as manually programmed plans often result in errors and inefficiencies.

Innovation Solution

A computer-implemented method using reinforcement learning to determine sampling combinations that encourage overlap between samples, calculating a score function based on the total area covered and overlap perimeter, and selecting the combination that maximizes this score to optimize sampling, thereby minimizing redundant samples and ensuring comprehensive coverage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manually programmed sensor plans are used, then the sensing system can operate with simple programming, but the system fails to adapt to unknown or changing areas of interest resulting in sampling errors and omissions

Engineering Contradiction:
Improveadaptability to unknown or changing areas of interestVSAvoidprogramming complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The sensor plan transitions from a static manually programmed path to a dynamic reinforcement learning-based plan that adapts in real-time to unknown or changing areas of interest. The RL agent continuously learns optimal sampling trajectories by receiving rewards for covering new areas and penalties for redundant sampling, enabling the system to dynamically adjust to environmental changes without requiring complex manual reprogramming.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system employs self-service through autonomous reinforcement learning where the sensor planning system automatically improves its own performance over time. The RL agent learns optimal sampling strategies through trial and error, using reward signals to self-correct and adapt to different areas of interest without human intervention, thereby eliminating the need for complex manual programming while maintaining high adaptability.

Inventive Principle:
Principle #25Self-service

2Reliability

If large quantities of samples are captured to ensure completeness, then the area of interest is captured in sufficient detail, but the process becomes time consuming and produces inefficient redundant samples

Engineering Contradiction:
Improvecompleteness of area coverageVSAvoidsampling efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The reinforcement learning system implements continuous feedback loops where the sensor planning agent receives reward signals based on sampling effectiveness. The reward function provides positive reinforcement for capturing new unique areas and negative reinforcement for redundant sampling, enabling the system to learn optimal sampling quantities that ensure completeness while minimizing redundancy. This feedback mechanism allows the system to adaptively determine the precise number of samples needed without excessive overhead.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically changes sampling parameters including the number of samples, sampling density, and trajectory based on real-time feedback from the environment. The reinforcement learning agent adjusts these parameters adaptively, increasing sampling density in complex or changing areas while reducing sampling in stable, well-understood regions, thereby optimizing the balance between completeness and efficiency rather than using a fixed large sample quantity throughout.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If manually programmed sensor plans are used for certain areas, then the sampling process is straightforward, but the system cannot handle deviations from the preprogrammed plan resulting in errors and omissions

Engineering Contradiction:
Improvesimplicity of sensor planningVSAvoidaccuracy in capturing deviating areas
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent replaces the mechanical system of manual programming with an intelligent reinforcement learning-based sensor planning system. Instead of relying on pre-programmed trajectories that require careful manual configuration, the RL agent autonomously generates and adjusts sampling plans based on real-time environmental perception, maintaining ease of operation while significantly improving reliability for handling unexpected deviations and unknown areas.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Productivity

If reinforcement learning is used to determine sampling combinations, then the system achieves autonomous and efficient capture with minimal samples, but the computational complexity increases

Engineering Contradiction:
Improvesampling efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-training the reinforcement learning agent offline or during idle periods to learn general sampling strategies and environmental patterns. This preliminary learning phase allows the agent to develop robust policies that can be quickly executed during actual sensing operations, reducing the computational burden during real-time sampling while maintaining high efficiency and adaptability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3535096B1Robotic sensing apparatus and methods of sensor planning
Publication Date: 2020.09.02 GENERAL ELECTRIC CO
  • EP3535096B1 patent drawingFigure 1~2
  • EP3535096B1 patent drawingFigure 3
  • EP3535096B1 patent drawingFigure 4

AI summary

The present disclosure is directed to a computer-implemented method of sensor planning for acquiring samples via an apparatus including one or more sensors. The computer- implemented method includes defining, by one or more computing devices, an area of interest; identifying, by the one or more computing devices, one or more sensing parameters for the one or more sensors; determining, by the one or more computing devices, a sampling combination for acquiring a plurality of samples by the one or more sensors based at least in part on the one or more sensing parameters; and providing, by the one or more computing devices, one or more command control signals to the apparatus including the one or more sensors to acquire the plurality of samples of the area of interest using the one or more sensors based at least on the sampling combination.