Autonomous UAV Policy Switching in Dynamic Tactical Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for autonomous UAV control in tactical environments fail to generalize to unpredictable conditions and do not account for dynamic mission plans, broken data links, and the need for operator and machine teaming.
Innovation Solution
A system, method, and program product for controlling autonomous agents, which includes sensors for receiving environmental and non-environmental inputs, actuators for performing actions, and a controller that manages policies based on inputs, evaluates value functions, and selects new operative policies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If deep reinforcement learning solutions are used for autonomous UAV control, then autonomy in stationary objective tasks is improved, but adaptability to dynamic tactical environments deteriorates
Solution Approach 1:
The system implements dynamic policy switching by maintaining multiple candidate policies (different deep reinforcement learning models trained for different objectives) and selecting among them based on current environmental conditions and mission requirements. This allows the UAV to adapt its behavior dynamically rather than relying on a single static policy.
Solution Approach 2:
The control system is designed to handle multiple objectives and task types through a unified framework that can load and execute different candidate policies. The system serves multiple functions including autonomous navigation, target acquisition, obstacle avoidance, and mission planning within a single integrated architecture.
2Productivity
If single operative policy is used for control, then execution efficiency is improved, but ability to handle terminated policies and dynamic objectives deteriorates
Solution Approach 1:
The system prepares multiple candidate policies in advance, each trained for specific objectives or conditions. When the current operative policy becomes terminated or obsolete, the system can immediately switch to a pre-prepared candidate policy without interruption to mission execution.
Solution Approach 2:
The system changes the operative policy parameter dynamically by selecting different candidate policies based on current mission requirements and environmental conditions. This allows the UAV to adapt its control strategy without restarting or retraining the entire system.
3Reliability
If deep learning solutions are used for object avoidance and collision detection, then safety in stationary environments is improved, but generalization to unpredictable tactical environments deteriorates
Solution Approach 1:
The system uses multiple candidate policies that are dynamically selected based on the tactical situation. Each policy is specialized for certain conditions, allowing the UAV to maintain high safety performance across diverse and unpredictable environments by adapting to each specific scenario.
Solution Approach 2:
The overall control system is segmented into multiple specialized candidate policies, each handling specific aspects of autonomous operation. This segmentation allows each policy to be optimized for its specific function while the system as a whole achieves broad adaptability through policy selection.
4Extent of automation
If autonomous control without operator teaming is implemented, then machine autonomy is improved, but effectiveness in tactical missions requiring human judgment deteriorates
Solution Approach 1:
The system introduces an operator interface that acts as an intermediary between the autonomous UAV and the human operator. Operators can provide high-level guidance, adjust mission parameters, and influence policy selection, combining machine autonomy with human judgment for enhanced tactical effectiveness.
Solution Approach 2:
The system implements feedback loops where operator inputs and mission outcomes inform policy selection and system behavior. This allows human operators to guide the autonomous system while maintaining efficiency, creating a collaborative human-machine teaming architecture.
Data Source
AI summary
Embodiments of the disclosure provide a machine learning framework to control an autonomous agent in a dynamic environment. A system according to the disclosure includes a sensor coupled to the autonomous agent to receive a set of inputs. At least one actuator causes the autonomous agent to perform an action. A controller causes the actuator to perform an action based on the inputs and an operative policy. The controller determines whether the operative policy is terminated, based on the set of inputs. Upon terminating the operative policy, the controller evaluates several candidate policies and selects one of the candidate policies as a new operative policy.


