USV Formation Path Following With Deep RL and Leader-Follower Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for unmanned surface vessel (USV) formation path following lack efficiency in navigating complex water environments and maintaining formation patterns, especially with multi-underactuated USVs, due to limitations in existing control strategies and reinforcement learning technologies.

Innovation Solution

A deep reinforcement learning-based method is introduced, utilizing a decision-making neural network model that extracts environmental information, employs a random braking mechanism, and combines collaborative exploration with a leader-follower formation control strategy to optimize USV formation path following, incorporating a reward function that maximizes speed and minimizes lateral deviation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional control strategies are used for multi-underactuated USV formation path following, then the system structure is simple, but the navigation efficiency in complex water environments deteriorates

Engineering Contradiction:
Improvenavigation efficiencyVSAvoidcontrol system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical control strategies with deep reinforcement learning algorithms. The decision-making neural network model learns optimal control policies through interaction with the environment, substituting conventional control theory-based approaches with data-driven intelligent decision-making, thereby improving navigation efficiency in complex environments.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the control approach from fixed parameter traditional controllers to adaptive neural network parameters that evolve through training. The decision-making neural network adjusts its internal parameters (weights and biases) based on environmental feedback, enabling adaptive optimization of navigation performance for multi-USV formations.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If existing reinforcement learning technologies are applied to USV formation control, then the adaptability to complex environments improves, but the training time and computational resources increase

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent segments the formation control problem into individual USV control tasks, where each USV has its own decision-making neural network. This segmentation allows parallel training and evaluation, reducing overall training time while maintaining formation coordination through shared environmental observations and collaborative exploration mechanisms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action through pre-training the decision-making neural networks before actual formation operations. The networks are trained in simulation environments beforehand, allowing them to learn basic navigation skills and formation patterns in advance, thereby reducing the time needed for real-world deployment and adaptation.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If collaborative exploration is used to train the decision-making neural network, then the path following performance improves, but the complexity of the training process increases

Engineering Contradiction:
Improvepath following precisionVSAvoidtraining process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where the decision-making neural networks receive continuous environmental observations and reward signals during collaborative exploration. The reward function provides feedback on formation maintenance quality and path following accuracy, guiding the networks to learn precise path following behaviors while maintaining formation patterns through iterative refinement.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11914376B2USV formation path-following method based on deep reinforcement learning
Publication Date: 2024.02.27 WUHAN UNIV OF TECH
  • US11914376B2 patent drawing
  • US11914376B2 patent drawing
  • US11914376B2 patent drawing

AI summary

The invention discloses an unmanned surface vessel (USV) formation path-following method based on deep reinforcement learning, which includes USV navigation environment exploration, reward function design, formation pattern keeping, a random braking mechanism and path following, wherein the USV navigation environment exploration is realized adopting simultaneous exploration by multiple underactuated USVs to extract environmental information, the reward function design includes the design of a formation pattern composition and a path following error, the path following controls USVs to move along a preset path by a leader-follower formation control strategy, and path following of all USVs in a formation is realized by constantly updating positions of the USVs. The invention accelerates the training of a USV path point following model through a collaborative exploration strategy, and combines the collaborative exploration strategy with the leader-follower formation control strategy to form the USV formation path following algorithm.