Control Learning Feedback for Efficient Training Data Acquisition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing learning systems for controlling robots lack the ability to determine when to stop learning, leading to inefficient and potentially unnecessary training processes.

Innovation Solution

A learning device that selects search points for training data acquisition, calculates evaluation information for the executability of operations, and determines whether to continue acquiring training data based on the evaluation of the acquisition status.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If training data acquisition is performed continuously without evaluation, then the learning model may achieve better performance, but the learning time and computational resources are wasted on unnecessary training

Engineering Contradiction:
Improvelearning effectivenessVSAvoidlearning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a feedback mechanism where the evaluation unit continuously assesses the acquisition status of training data and provides feedback to the search point setting unit. This feedback loop enables the system to determine when sufficient training data has been acquired and when learning can be stopped, preventing unnecessary training while ensuring adequate learning effectiveness.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary evaluation of training data acquisition status before continuing the learning process. The evaluation unit assesses whether the acquired training data meets the required standards in advance, allowing the system to stop learning early when the criteria are met, thus avoiding time waste on redundant training iterations.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If training data acquisition is performed extensively to ensure comprehensive learning, then the learning model achieves better performance, but the complexity of the learning system increases

Engineering Contradiction:
Improvelearning completenessVSAvoidlearning system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The evaluation unit provides continuous feedback on the acquisition status of training data, enabling the system to determine when learning objectives have been met. This feedback mechanism ensures learning completeness without requiring overly complex learning systems, as the evaluation-guided approach naturally terminates training when sufficient data is acquired.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The learning system performs self-evaluation through the evaluation unit, which autonomously assesses whether the acquired training data is sufficient. This self-service capability allows the system to ensure learning completeness independently without requiring external intervention or overly complex control mechanisms.

Inventive Principle:
Principle #25Self-service

3Loss of time

If the learning process is terminated early to save time, then computational resources are conserved, but the learning model may not achieve sufficient performance

Engineering Contradiction:
Improvelearning timeVSAvoidlearning performance
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The evaluation unit performs preliminary assessment of training data acquisition status before each learning iteration. This preliminary evaluation ensures that learning is terminated only when sufficient data has been acquired, preventing early termination that would compromise performance while avoiding unnecessary continuation that would waste time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from the evaluation unit to make informed decisions about continuing or terminating learning. The feedback mechanism ensures that learning performance requirements are met by continuously monitoring acquisition status and adjusting the learning process accordingly, preventing premature termination.

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If comprehensive training data is acquired for all possible operations, then the control system becomes more versatile, but the data acquisition process becomes inefficient

Engineering Contradiction:
Improvecontrol coverageVSAvoiddata acquisition efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The evaluation unit provides feedback on the acquisition status of training data for different operations, enabling the search point setting unit to prioritize acquiring data for operations that are more critical or have lower coverage. This feedback-driven approach ensures versatile control coverage while maintaining data acquisition efficiency by focusing resources on high-priority operations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies different acquisition strategies to different search points (operations) based on their individual evaluation status. The evaluation unit assesses each operation's training data coverage locally, allowing the system to allocate data acquisition resources unevenly across different operations, ensuring critical operations receive adequate attention while avoiding redundant data collection for well-covered operations.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250164944A1Learning device, control device, learning method, and storage medium
Publication Date: 2025.05.22 NEC CORP
  • US20250164944A1 patent drawing
  • US20250164944A1 patent drawing
  • US20250164944A1 patent drawing

AI summary

A learning device selects, from among search points indicating an operation of a control target, a search point to be subjected to training data acquisition for learning of a control of the control target. The learning device calculates information indicating an evaluation of whether or not an operation indicated by the selected search point is executable, and an output value for the operation indicated by the selected search point to be output by a controller for controlling the control target. The learning device acquires, based on the selected search point, the information indicating the evaluation of whether or not the operation indicated by the selected search point is executable, and the output value for the operation indicated by the selected search point to be output by the controller, training data for learning a control of the control target that is performed by the controller.