Multi-Agent DRL for Scalable Asset Inspection and Maintenance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for inspection and maintenance (I&M) planning in multi-asset infrastructure environments face optimality, scalability, and uncertainty-induced complexities, often generating sub-optimal solutions due to computational challenges, especially in large networks and long time-horizons, and are not easily extendable to environments with constraints.

Innovation Solution

A multi-agent Deep Reinforcement Learning (DRL) framework utilizing Markov Decision Process (MDP) and Partially Observable Markov Decision Process (POMDP) with deep neural networks for adaptive evaluation and prioritization, incorporating state augmentation and Lagrange multipliers to handle noisy and ambiguous data, enabling decentralized decision-making.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If conventional methods (threshold-based formulations, decision tree analysis, renewal theory, stochastic optimal control) are used for I&M planning, then the approach is relatively simple and easy to implement, but the solutions are widely sub-optimal due to computational challenges in large networks and long time-horizons

Engineering Contradiction:
Improveease of implementationVSAvoidoptimality of solution
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent replaces conventional mechanical/mathematical optimization methods (threshold-based formulations, decision tree analysis, renewal theory, stochastic optimal control) with a deep reinforcement learning system. The DRL agent learns optimal maintenance policies through interaction with the infrastructure environment, substituting complex mathematical programming with neural network-based decision-making that can handle large state and action spaces efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The reinforcement learning agent performs self-learning and self-optimization through continuous interaction with the environment. The agent autonomously explores the infrastructure system, learns from observations and rewards, and adapts its maintenance policies without requiring external reconfiguration or manual programming of maintenance strategies.

Inventive Principle:
Principle #25Self-service

2Use of energy by moving object

If conventional methods are used for I&M planning, then the computational requirements are lower, but the scalability to large networks and long time-horizons is poor

Engineering Contradiction:
Improvecomputational resource consumptionVSAvoidscalability
Core Design Contradiction:
Use of energy by moving objectVSAdaptability or versatility

Solution Approach 1:

The patent segments the infrastructure system into multiple independent assets or components, each managed by the same DRL framework. The agent can learn and apply policies to individual assets while also coordinating across the network, enabling scalable deployment to large infrastructure systems without requiring complete re-computation for each additional asset.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from conventional two-dimensional planning (single asset, single time-point) to multi-dimensional optimization by incorporating temporal dynamics, multiple assets, and uncertain future conditions into the reinforcement learning framework. The agent learns policies that optimize across time and space dimensions simultaneously, enabling scalability to long time-horizons and large networks.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If static optimization formulations are used for maintenance evaluation, then the model is simpler and faster to compute, but the formulation cannot adapt to dynamic maintenance objectives and constraints

Engineering Contradiction:
Improvecomputational speedVSAvoidadaptability to dynamic objectives
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic optimization through reinforcement learning, where the agent continuously adapts its maintenance policies based on changing conditions. The policy can respond to dynamic objectives such as varying budget constraints, priority changes, and real-time asset conditions, allowing the system to optimize maintenance strategies in response to evolving requirements rather than being locked into a static formulation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The reinforcement learning framework incorporates feedback loops where the agent observes the consequences of maintenance actions, receives reward signals based on performance metrics, and uses this feedback to refine its policy. This feedback mechanism enables the system to adapt to dynamic objectives and learn from past experiences, continuously improving its maintenance decisions without requiring re-formulation of the optimization problem.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250278611A1Apparatus and method for improved inspection and/or maintenance management
Publication Date: 2025.09.04 THE PENN STATE RES FOUND INC
  • US20250278611A1 patent drawing
  • US20250278611A1 patent drawing
  • US20250278611A1 patent drawing

AI summary

A communication system, a computer device, as well as a computer program encoded on a non-transitory computer storage medium can be configured to facilitate generation of output and/or one or more graphical user interface displays to provide improved asset maintenance management. Embodiments can be configured to generate a maintenance management process as output, including scheduling inspection and maintenance to be performed to provide significant cost and efficiency improvements.