Machine Control via Subordinate RL Skills for Stable Multi-Objective Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manually tuning control algorithms for machines to meet multiple control objectives is complex and costly, often hindered by insufficient data, and reinforcement learning processes can become unstable due to complex global reward functions.

Innovation Solution

A device and method that decompose the control problem into subordinate control skills, each optimized through individual reinforcement learning processes, with a superordinate control skill determined by selecting or combining these skills based on individual reward functions to achieve an optimized overall strategy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If reinforcement learning with a global reward function is used to learn control strategy for multiple control objectives, then automated learning capability is achieved, but the learning process becomes unstable and difficult to implement

Engineering Contradiction:
Improveautomated learning capabilityVSAvoidlearning process stability
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent segments the global reward function into multiple independent local reward functions, each corresponding to a specific control objective. Instead of using one complex global reward function that combines all objectives, the system creates separate reward functions (e.g., one for efficiency, one for emissions, one for vibrations) that can be independently designed and optimized. This segmentation makes the learning process more stable and manageable while maintaining automated learning capability across multiple control objectives.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If manual tuning of control algorithm is performed to meet multiple control objectives, then control precision can be optimized, but the process becomes complex and time-consuming

Engineering Contradiction:
Improvecontrol precisionVSAvoidtuning time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the control system to automatically learn and optimize control strategies through reinforcement learning with local reward functions. Instead of requiring manual expert tuning to achieve precise control across multiple objectives, the system autonomously learns optimal control policies by interacting with the environment and receiving feedback from local reward functions. This eliminates the time-consuming manual tuning process while maintaining or improving control precision.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If manual design of control strategy is performed for all individual control objectives, then comprehensive control coverage is achieved, but the process becomes cumbersome and costly

Engineering Contradiction:
Improvecontrol coverageVSAvoiddesign complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by breaking down the comprehensive control strategy design into multiple independent local reward functions, each handling a specific control objective. This allows the system to achieve comprehensive control coverage across all objectives while simplifying the design process. Each local reward function can be designed independently using domain knowledge, avoiding the need for complex integrated manual design while ensuring all control objectives are addressed.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12050440B2Machine control based on automated learning of subordinate control skills
Publication Date: 2024.07.30 SIEMENS AG
  • US12050440B2 patent drawing
  • US12050440B2 patent drawing

AI summary

Method and device for controlling a machine in accordance with to multiple control objectives in which machine control is based on automated learning of subordinate control skills, wherein the device provides multiple subordinate control skills which are each assigned to a different one of the multiple control objectives, the device provides multiple learning processes that are reinforcement learning processes that are each assigned to a different one of the multiple control objectives and are configured to optimize the corresponding subordinate control skill based on input data received from the machine, and where the device is configured to determine a superordinate control skill based on the subordinate control skills and to control the machine based on the superordinate control skill.