Robust Constraint Control for Systems With Uncertain Dynamics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Real-world systems, such as autonomous vehicles and robotics, face challenges in controlling uncertain dynamics due to non-stationarity, wear-and-tear, and uncalibrated sensors, leading to difficulties in designing optimal controllers, especially when operating in environments with unknown conditions.

Innovation Solution

The integration of principles from Markov decision processes (MDPs), robust MDPs (RMDPs), and constraint MDPs (CMDPs) into a robust and constraint MDP (RCMDP) framework, which reformulates uncertainties as ambiguity sets to optimize performance and safety costs while enforcing constraints, using a joint multifunction optimization approach and Lyapunov theory for computational simplification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a traditional controller is designed for a system with uncertain dynamics, then the controller structure becomes complex and difficult to design, but the system cannot achieve optimal performance under uncertain conditions

Engineering Contradiction:
Improvecontroller performance under uncertaintyVSAvoidcontroller design complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional model-based control mechanisms with a data-driven reinforcement learning approach. Instead of designing complex controllers based on uncertain system models, the system learns optimal control policies through interaction with the environment, substituting mathematical modeling with empirical learning from state-action-reward transitions

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the control problem from optimizing control parameters directly to optimizing a value function and policy parameters through gradient descent. The controller learns by adjusting parameters of the value function and policy in response to environmental feedback, rather than manually tuning controller parameters based on system models

Inventive Principle:
Principle #35Parameter changes

2Productivity

If the system operates with unknown environmental conditions and uncalibrated sensors, then the system can be deployed faster, but the control accuracy deteriorates

Engineering Contradiction:
Improvesystem deployment speedVSAvoidcontrol accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs self-calibration and self-adaptation through the reinforcement learning process. The controller automatically adjusts its policy based on environmental feedback and observed state transitions, enabling it to adapt to unknown conditions and uncalibrated sensors without manual intervention or pre-calibration

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements continuous feedback through the reward signal and observed state transitions. The controller uses this feedback to update its value function and policy parameters, allowing it to compensate for measurement errors and environmental uncertainties by learning from actual system behavior rather than relying on predetermined accuracy

Inventive Principle:
Principle #23Feedback

3Reliability

If robust MDP is used to optimize performance cost for worst-case conditions, then the system becomes more reliable under uncertainty, but the imposed constraints are violated

Engineering Contradiction:
Improveperformance under worst-case conditionsVSAvoidconstraint violation
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the optimization problem into two separate components: the value function optimizes for performance while the policy function ensures constraint satisfaction. This segmentation allows each component to focus on its specific objective without compromising the other, avoiding the constraint violations that occur when optimizing for worst-case performance alone

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The policy gradient method acts as an intermediary mechanism that translates the value function's performance optimization into constraint-satisfying actions. The policy serves as a mediator between the performance objectives and constraint requirements, ensuring that optimization for robust performance does not lead to constraint violations

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11640162B2Apparatus and method for controlling a system having uncertainties in its dynamics
Publication Date: 2023.05.02 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US11640162B2 patent drawing
  • US11640162B2 patent drawing
  • US11640162B2 patent drawing

AI summary

A controller for controlling a system having uncertainties in its dynamics subject to constraints on an operation of the system is provided. The controller is configured to acquire historical data of the operation of the system, and determine, for the system in a current state, a current control action transitioning a state of the system from the current state to a next state. The current control action is determined according to a robust and constraint Markov decision process (RCMDP) that uses the historical data to optimize a performance cost of the operation of the system subject to an optimization of a safety cost enforcing the constraints on the operation, wherein a state transition for each of state and action pairs in the performance cost and the safety cost is represented by a plurality of state transitions capturing the uncertainties of the dynamics of the system.