Reinforcement Learning Control Parameters for Adaptive Device Operation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for controlling devices are inadequate in efficiently adjusting control parameters based on real-time data and environmental conditions, leading to suboptimal performance and increased energy consumption.

Innovation Solution

An apparatus and method that utilize reinforcement learning to acquire and adjust control parameters by analyzing measurement data from sensors, using a reward function to optimize control actions, allowing for automatic and data-driven decision-making in controlling devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional control methods are used to control devices, then the control system is simple and easy to implement, but the control performance is suboptimal and energy consumption is increased

Engineering Contradiction:
Improvecontrol performanceVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The control system performs self-learning through reinforcement learning, where the learning unit automatically optimizes control parameters by receiving rewards or penalties based on control outcomes. This self-service mechanism eliminates the need for manual tuning and enables the system to autonomously improve energy efficiency and control performance over time

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements a closed-loop feedback mechanism where the learning unit receives reward values based on control results, and uses this feedback to continuously update and optimize control parameters. The reward function provides directional guidance for parameter adjustment, enabling the system to learn from past performance and improve future control decisions

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If trial-and-error operations are used to adjust control parameters, then the system can adapt to different conditions, but the time consumption and operational complexity increase

Engineering Contradiction:
Improveadaptability to conditionsVSAvoidtime consumption
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary learning during idle periods or off-peak times, accumulating knowledge and optimizing control strategies in advance. This allows the control system to be ready with optimized parameters when actual operational decisions are needed, eliminating the need for time-consuming trial-and-error during critical operations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual trial-and-error adjustment with an automated learning system that uses algorithms to optimize control parameters. This substitution of mechanical/manual operations with computational intelligence dramatically reduces the time and effort required for parameter tuning while improving adaptability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3828651B1Apparatus, method and program
Publication Date: 2023.04.26 YOKOGAWA ELECTRIC CORP
  • EP3828651B1 patent drawingFigure 1
  • EP3828651B1 patent drawingFigure 2
  • EP3828651B1 patent drawingFigure 3

AI summary

An apparatus is provided, which includes a first acquisition unit for acquiring measurement data measured by a sensor and a first learning processing unit for executing, by using learning data including the measurement data acquired by the first acquisition unit and a control parameter indicating a first type of control content of at least one device to be controlled, a learning processing of a first model configured to output a recommended control parameter indicating the first type of control content recommended for increasing a reward value determined by a preset reward function in response to input of the measurement data.