Hierarchical DRL Planning for Manned Unmanned Platforms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep reinforcement learning (DRL) methods struggle to handle accurate and rapid team planning for multi-domain manned/unmanned platforms across complex, dynamic situations towards a common tactical goal.
Innovation Solution
A hierarchical DRL approach is implemented, comprising a global planning layer for determining collective goals, a platform planning layer for determining platform actions, and a platform control layer for executing functions, with information sharing and separate training of layers to enhance coordination and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If DRL-based methods are used for team planning, then team-level strategies and generalization to new environments are improved, but training time and environment interactions increase significantly
Solution Approach 1:
The patent segments the DRL training process into multiple parallel environments, where multiple instances of the same environment are created and trained simultaneously. This allows the system to explore diverse scenarios and improve generalization to new environments while reducing the total training time through parallel processing. The segmented approach enables efficient utilization of computational resources across multiple environment instances.
2Extent of automation
If DRL-based methods are used for team planning, then semantic goal definition and team-level strategies are achieved, but computational complexity and resource requirements increase
Solution Approach 1:
The patent introduces a hierarchical dimension to the DRL architecture, organizing the system into multiple layers: a global planner that handles high-level team-level strategies and semantic goal definition, and local planners that handle specific platform actions. This dimensional organization separates complex strategic reasoning from tactical execution, reducing the computational burden on individual components while maintaining overall system intelligence and automation capability.
3Productivity
If multiple DRL environments are trained in parallel, then training efficiency and data diversity are improved, but system resource consumption and management complexity increase
Solution Approach 1:
The patent implements a universal environment configuration system where a single base environment class can be instantiated multiple times with different parameters and settings. This multi-functional design allows the same environment framework to serve multiple training purposes simultaneously, managing diverse training scenarios through a unified interface. The universal configuration approach reduces management complexity by providing consistent control mechanisms across all parallel environments while maintaining training efficiency through simultaneous execution.
Data Source
AI summary
A method, apparatus and system for artificial intelligence-based HDRL planning and control for coordinating a team of platforms includes implementing a global planning layer for determining a collective goal and determining, by applying at least one machine learning process, at least one respective platform goal to be achieved by at least one platform, implementing a platform planning layer for determining, by applying at least one machine learning process, at least one respective action to be performed by the at least one of the platforms to achieve the respective platform goal, and implementing a platform control layer for determining at least one respective function to be performed by the at least one of the platforms. In the method, apparatus and system despite the fact that information is shared between at least two of the layers, the global planning layer, the platform planning layer, and the platform control layer are trained separately.


