Vehicle Engagement Control Using Reinforcement-Learned Actions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous control systems for vehicle engagements are inflexible and require extensive manual rule generation, making them inefficient in adapting to new scenarios and consuming significant resources.

Innovation Solution

A vehicle engagement control system utilizing machine-learning logic trained with reinforcement learning techniques to determine sequences of actions that minimize the use of resources and effectively remove opposing vehicles from an engagement zone, by receiving observations from both groups of vehicles and communicating optimized actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If rule-based autonomous control systems are used, then pre-programmed behaviors can be executed, but the systems are inflexible and require extensive manual rule generation

Engineering Contradiction:
Improveadaptability to new scenariosVSAvoidmanual rule generation
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs machine-learning logic that automatically learns and adapts engagement strategies through reinforcement learning from simulated engagements, eliminating the need for extensive manual rule generation. The machine-learning logic independently optimizes action sequences based on observations and outcomes, making the system self-improving rather than requiring continuous human programming.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transitions from fixed rule-based parameters to dynamic parameters learned through reinforcement learning. The machine-learning logic adjusts action sequences and engagement strategies based on simulated engagement outcomes, allowing the system to adapt to new scenarios by changing its internal parameters through learning rather than manual reprogramming.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If reinforcement learning training is used, then adaptability improves, but computational resources and training time increase

Engineering Contradiction:
ImproveadaptabilityVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs reinforcement learning training in advance using simulated engagements before actual deployment. The machine-learning logic is pre-trained on diverse engagement scenarios, allowing it to rapidly adapt to new situations during actual operations without requiring extensive real-time training, thus reducing operational training time while maintaining high adaptability.

Inventive Principle:
Principle #10Preliminary action

3Loss of substance

If machine-learning logic determines action sequences, then resource consumption is minimized, but system complexity increases

Engineering Contradiction:
Improveresource consumptionVSAvoidsystem complexity
Core Design Contradiction:
Loss of substanceVSDevice complexity

Solution Approach 1:

The machine-learning logic uses reinforcement learning with feedback from engagement outcomes to optimize action sequences. The system receives feedback on the effectiveness of actions taken (such as whether opposing vehicles were removed from the engagement zone), and uses this feedback to continuously improve its strategy, minimizing resource consumption like weapons and fuel while managing complexity through iterative optimization rather than static complex rules.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11586200B2Method and system for vehicle engagement control
Publication Date: 2023.02.21 HRL LAB
  • US11586200B2 patent drawing
  • US11586200B2 patent drawing
  • US11586200B2 patent drawing

AI summary

A method includes receiving, by machine-learning logic, observations indicative of a states associated with a first and second group of vehicles arranged within an engagement zone during a first interval of an engagement between the first and the second group of vehicles. The machine-learning logic determines actions based on the observations that, when taken simultaneously by the first group of vehicles during the first interval, are predicted by the machine-learning logic to result in removal of one or more vehicles of the second group of vehicles from the engagement zone during the engagement. The machine-learning logic is trained using a reinforcement learning technique and on simulated engagements between the first and second group of vehicles to determine sequences of actions that are predicted to result in one or more vehicles of the second group being removed from the engagement zone. The machine-learning logic communicates the plurality of actions to the first group of vehicles.