Constrained Machine Control Using CAIS-Guided Policy Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional reinforcement learning methods are not suitable for data-driven control of constrained systems as they do not guarantee satisfaction of state and input constraints in continuous state-action spaces, leading to potential system instability.

Innovation Solution

The development of a method that formulates control problems for constrained machines within a constraint admissible invariant set (CAIS), using reinforcement learning principles to iteratively update control policies and CAIS based on collected data, ensuring constraint satisfaction without requiring a dynamical model of the system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional reinforcement learning methods are used for data-driven control, then the control policy can be learned from operational data, but state and input constraints cannot be guaranteed to be satisfied

Engineering Contradiction:
Improvedata-driven control capabilityVSAvoidconstraint satisfaction guarantee
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The control problem is segmented into two distinct components: (1) learning the system dynamics model from operational data, and (2) solving the constrained optimal control problem using the learned model. This segmentation allows each component to be optimized independently, with the model learning phase focusing on accuracy and the control synthesis phase focusing on constraint satisfaction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary model learning from operational data before executing constrained optimal control. By pre-learning the system dynamics, the controller has accurate predictions available when making control decisions, enabling it to anticipate future states and ensure constraint satisfaction without requiring real-time trial and error.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If model-based control methods are used, then state and input constraints can be directly incorporated into control design, but accurate analytical models are often unavailable or difficult to update in real-time

Engineering Contradiction:
Improveconstraint satisfaction guaranteeVSAvoidmodel availability and maintenance
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the need for analytical mathematical models with data-driven learned models. Instead of relying on physics-based equations that are difficult to derive and update, the system learns system dynamics from operational data, substituting complex model derivation with simpler data collection and machine learning processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-learning by automatically capturing operational data and updating its internal model of system dynamics without requiring external expertise in model derivation. The controller continuously improves its understanding of the system through ongoing operation, making the system self-maintaining and adaptive.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If indirect data-driven control methods are used to construct system models first, then control policies can be designed using learned models, but large quantities of data are required in the model-building phase

Engineering Contradiction:
Improvedata-driven control capabilityVSAvoiddata quantity requirement
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system uses partial model learning, focusing only on the specific dynamics relevant to the control task at hand rather than attempting to learn the complete system model. This selective approach reduces data requirements by concentrating learning efforts on the most critical state transitions and control relationships needed for safe operation.

Inventive Principle:
Principle #16Partial or excessive action

4Quantity of substance

If direct data-driven control methods are used, then less data is required compared to indirect methods, but handling state and input constraints remains difficult

Engineering Contradiction:
Improvedata quantity requirementVSAvoidconstraint satisfaction guarantee
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent introduces a learned system dynamics model as an intermediary between raw operational data and control decisions. This intermediary structure enables the system to use limited data efficiently for learning, then apply the learned model in a constrained optimal control framework that guarantees constraint satisfaction, bridging the gap between data efficiency and reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11106189B2System and method for data-driven control of constrained system
Publication Date: 2021.08.31 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US11106189B2 patent drawing
  • US11106189B2 patent drawing
  • US11106189B2 patent drawing

AI summary

A machine subject to state and control input constraints is control, while the control policy is learned from data collected during an operation of the machine. To ensure satisfaction of the constraints, the state of machine is maintained within a constraint admissible invariant set (CAIS) satisfying the constraints and the machine is controlled with corresponding control policy mapping a state of the system within the CAIS to a control input satisfying the control input constraints. The machine is controlled using a constrained policy iteration, in which a constrained policy evaluation updates CAIS and value function and a constrained policy improvement updates control policy that improves the cost function of operation according to the updated CAIS and the corresponding updated value function.