Hierarchical DRL Planning for Manned Unmanned Platforms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep reinforcement learning (DRL) methods struggle to handle accurate and rapid team planning for multi-domain manned/unmanned platforms across complex, dynamic situations towards a common tactical goal.

Innovation Solution

A hierarchical DRL approach is implemented, comprising a global planning layer for determining collective goals, a platform planning layer for determining platform actions, and a platform control layer for executing functions, with information sharing and separate training of layers to enhance coordination and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If DRL-based methods are used for team planning, then team-level strategies and generalization to new environments are improved, but training time and environment interactions increase significantly

Engineering Contradiction:
Improvegeneralization to new environmentsVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent segments the DRL training process into multiple parallel environments, where multiple instances of the same environment are created and trained simultaneously. This allows the system to explore diverse scenarios and improve generalization to new environments while reducing the total training time through parallel processing. The segmented approach enables efficient utilization of computational resources across multiple environment instances.

Inventive Principle:
Principle #1Segmentation

2Extent of automation

If DRL-based methods are used for team planning, then semantic goal definition and team-level strategies are achieved, but computational complexity and resource requirements increase

Engineering Contradiction:
Improveteam-level strategiesVSAvoidcomputational complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces a hierarchical dimension to the DRL architecture, organizing the system into multiple layers: a global planner that handles high-level team-level strategies and semantic goal definition, and local planners that handle specific platform actions. This dimensional organization separates complex strategic reasoning from tactical execution, reducing the computational burden on individual components while maintaining overall system intelligence and automation capability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If multiple DRL environments are trained in parallel, then training efficiency and data diversity are improved, but system resource consumption and management complexity increase

Engineering Contradiction:
Improvetraining efficiencyVSAvoidsystem management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal environment configuration system where a single base environment class can be instantiated multiple times with different parameters and settings. This multi-functional design allows the same environment framework to serve multiple training purposes simultaneously, managing diverse training scenarios through a unified interface. The universal configuration approach reduces management complexity by providing consistent control mechanisms across all parallel environments while maintaining training efficiency through simultaneous execution.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11960994B2Artificial intelligence-based hierarchical planning for manned/unmanned platforms
Publication Date: 2024.04.16 SRI INTERNATIONAL
  • US11960994B2 patent drawing
  • US11960994B2 patent drawing
  • US11960994B2 patent drawing

AI summary

A method, apparatus and system for artificial intelligence-based HDRL planning and control for coordinating a team of platforms includes implementing a global planning layer for determining a collective goal and determining, by applying at least one machine learning process, at least one respective platform goal to be achieved by at least one platform, implementing a platform planning layer for determining, by applying at least one machine learning process, at least one respective action to be performed by the at least one of the platforms to achieve the respective platform goal, and implementing a platform control layer for determining at least one respective function to be performed by the at least one of the platforms. In the method, apparatus and system despite the fact that information is shared between at least two of the layers, the global planning layer, the platform planning layer, and the platform control layer are trained separately.