Autonomy Decision Controller for Human-Like Vehicle Maneuvers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in communicating instructions effectively with human operators, leading to inconsistent maneuvers that can confuse other vehicles or human operators during training or testing, as existing systems lack a shared mental model of pilot behavior.

Innovation Solution

A system utilizing a machine learning engine trained with a reward function to generate a value function and policy, enabling autonomous vehicles to understand and replicate human-like behaviors by learning from human pilot responses to various conditions, allowing for more sophisticated interactions with human operators.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If autonomous vehicles use control laws to generate maneuvers, then they can operate autonomously without human input, but the maneuvers may be inconsistent with expected pilot behavior and confuse other vehicles or human operators

Engineering Contradiction:
Improveautonomous operationVSAvoidshared mental model of pilot behavior
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent applies the copying principle by training the autonomous vehicle's decision controller to replicate human pilot behavior patterns. The system learns from training data consisting of human pilot responses to various flight conditions, creating a copy of human decision-making processes. This allows the autonomous vehicle to generate maneuvers that are consistent with expected pilot behavior while maintaining full autonomy, resolving the contradiction between autonomous operation and behavioral consistency.

Inventive Principle:
Principle #26Copying

2Extent of automation

If autonomous vehicles generate instructions based on optimizing a particular objective, then they can operate fully autonomously, but they cannot communicate instructions at a conceptual level that is easy for human operators to understand

Engineering Contradiction:
Improvefully autonomous operationVSAvoidcommunication with human operators
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent applies parameter changes by transitioning the decision controller from purely mathematical optimization parameters to behavioral parameters learned from human pilots. The system changes the nature of its decision-making parameters to match human cognitive patterns, making the autonomous vehicle's instructions and maneuvers understandable to human operators while maintaining full autonomous operation. The value function and policy are trained to reflect human-like trade-offs and priorities.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If a machine learning system is trained with a reward function to learn optimal maneuvers, then it can achieve high productivity in learning pilot behavior, but the complexity of training and generating value functions increases

Engineering Contradiction:
Improvelearning speedVSAvoidtraining system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the learning process into distinct components: the reward function that evaluates maneuvers, the value function that estimates future rewards, and the policy that selects actions. This segmentation allows the complex learning task to be broken down into manageable parts that can be trained and optimized separately, improving learning efficiency while managing system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11107001B1Systems and methods for practical autonomy decision controller
Publication Date: 2021.08.31 ROCKWELL COLLINS INC
  • US11107001B1 patent drawing
  • US11107001B1 patent drawing
  • US11107001B1 patent drawing

AI summary

A system includes a machine learning engine configured to receive training data including a plurality of input conditions associated with a state space and a plurality of response maneuvers associated with the state space and train a learning system using the training data and a reward function including a plurality of terms associated with a plurality of end state spaces, each term in the plurality of terms defines an end reward value for each end state space. A value function and policy are generated. The value function comprising a plurality of values, wherein each response maneuvers in the plurality of response maneuvers is associated with a value in the plurality of values related to transitioning from the state space to each end state space, the policy indicative of connections between the state spaces, plurality of values, and the respective end reward value for the plurality of end state spaces.