Deep MRAC Weight Adaptation for Stable Nonlinear Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep Reinforcement Learning methods lack stability guarantees and boundedness during learning transients, making them unsuitable for safety-critical applications, while traditional adaptive control methods do not effectively leverage the power of deep neural networks for modeling nonlinearities.

Innovation Solution

The development of a Deep Neural Network-based Model Reference Adaptive Controller (DMRAC) that uses a dual time-scale adaptation scheme to update weights, ensuring Uniform Ultimate Boundedness and long-term learning properties, combining the stability of MRAC with the learning capabilities of deep networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If Deep Reinforcement Learning methods are used for control, then learning performance and adaptability are improved, but stability and boundedness during learning transients deteriorate

Engineering Contradiction:
Improvelearning performanceVSAvoidstability guarantee
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The controller is segmented into two distinct components: a deep neural network for adaptive learning and a stabilizing feedback controller. This segmentation allows each component to perform its specialized function independently while working together to achieve both high adaptability and stability guarantees.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A stabilizing feedback controller acts as an intermediary between the deep neural network and the plant. This intermediary ensures stability and boundedness during learning transients while allowing the neural network to learn optimal control policies without compromising system safety.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional adaptive control methods are used, then stability is maintained, but the ability to model complex nonlinearities deteriorates

Engineering Contradiction:
ImprovestabilityVSAvoidnonlinearity modeling capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The invention merges traditional model reference adaptive control with deep neural networks to create a hybrid controller. This combination preserves the stability guarantees of MRAC while incorporating the powerful nonlinearity modeling capabilities of deep learning through the neural network's universal approximation properties.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If deep neural networks are used for control, then learning capability and feature extraction are improved, but computational complexity and training time increase

Engineering Contradiction:
Improvelearning capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The stabilizing feedback controller is designed in advance based on a nominal model of the plant, providing immediate stability without requiring extensive online computation. This preliminary action reduces the computational burden during real-time operation while the neural network learns offline or during transient periods.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12259735B2Deep model reference adaptive controller
Publication Date: 2025.03.25 THE BOARD OF TRUSTEES OF THE UNIV OF ILLINOIS
  • US12259735B2 patent drawing
  • US12259735B2 patent drawing
  • US12259735B2 patent drawing

AI summary

Aspects of the subject disclosure may include, for example, determining, at a slower time-scale, inner layer weights of an inner layer of a deep neural network; providing periodically to an outer layer of the deep neural network from the inner layer, a feature vector based upon the inner layer weights; and determining, at a faster time-scale, outer layer weights of the outer layer, wherein the outer layer weights are determined in accordance with a Model Reference Adaptive Control (MRAC) update law that is based upon the feature vector from the inner layer, and wherein the outer layer weights are determined more frequently than the inner layer weights. Other embodiments are disclosed.