Unmanned system control method and control system based on lyapunov neural network

By fitting the Lyapunov function of an unmanned system with a Lyapunov neural network and integrating it with a reinforcement learning agent, the safety problem of complex nonlinear systems in unmanned vessel control is solved, achieving autonomous learning and stable control. This method is applicable to unmanned vessels, unmanned vehicles, and drones.

CN115933467BActive Publication Date: 2026-04-14SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2022-12-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing unmanned vessel control methods cannot effectively solve the safety problems of complex nonlinear systems. They rely on prior human knowledge, making it difficult to cover a sufficient Lyapunov stability region in high-dimensional dynamic systems, and they are also difficult to integrate with other control algorithms.

Method used

A Lyapunov neural network is used to fit the Lyapunov function of the unmanned system. The safe area is divided through iterative training and fused with a model-based reinforcement learning agent to guide the control of the unmanned system.

Benefits of technology

It achieves safety assurance for complex nonlinear systems, can learn autonomously without relying on prior human knowledge, can be extended to high-dimensional dynamic systems, improves stability and control efficiency, and is easy to integrate with other algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115933467B_ABST
    Figure CN115933467B_ABST
Patent Text Reader

Abstract

The application discloses an unmanned system control method and control system based on a Lyapunov neural network, and comprises the following steps: fitting a Lyapunov function corresponding to an unmanned system through a Lyapunov neural network; guiding the unmanned system to perform iterative training according to a safety region divided by the Lyapunov neural network; and controlling the unmanned system after fusing a model-based reinforcement learning intelligent agent of the Lyapunov neural network and the unmanned system. The Lyapunov function is fitted through the Lyapunov neural network, most of the Lyapunov stable regions can be covered, and sufficient exploration of the safety region is ensured. The method can be extended to a relatively complex nonlinear system, the Lyapunov neural network can be learned in unmanned systems such as unmanned ships, and the method can be effectively migrated to other control algorithms and combined with other algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned system control technology, and specifically relates to an unmanned system control method and control system based on Lyapunov neural network. Background Technology

[0002] In recent years, in order to address the shortage of skilled professionals and operational efficiency issues in the maritime transport industry, the development of unmanned vessels has been rapid, and various unmanned vessel control methods have emerged.

[0003] Ships navigating at sea are subject to environmental factors such as wind and currents, posing certain safety hazards. Safety has always been a core issue in the field of control; however, because the safety of unmanned surface vessel (USV) systems heavily relies on prior human knowledge of the USV and manual selection, safety issues are rarely addressed in existing USV control methods.

[0004] Safety-assured unmanned surface vessel (USV) control technology is of great significance. Ensuring the safety of USV control can reduce the likelihood of unnecessary damage and dangerous accidents such as capsizing. Furthermore, it can help USVs eliminate high-risk control maneuvers, achieving more stable and effective control and freeing them from excessive reliance on prior human knowledge, thus enabling true intelligence. Therefore, ensuring the safety of USVs is an important research direction and a critical issue that urgently needs to be addressed.

[0005] To address the issue of ensuring safe control, researchers have proposed many methods, which can be broadly categorized into three types: methods for calculating Lyapunov functions based on traditional methods; learning Lyapunov neural networks given a simple dynamical model; and learning Lyapunov neural network controllers. Among these, methods for calculating Lyapunov functions based on traditional methods utilize polynomial fitting; learning Lyapunov neural networks given a simple dynamical model uses a neural network to fit the Lyapunov function of a given dynamical system, solving the problem of finding the Lyapunov function easily; and learning Lyapunov neural network controllers can be applied to some simple nonlinear systems, finding a suitable control function while verifying the Lyapunov conditions.

[0006] Unmanned surface vessel (USV) systems are relatively complex nonlinear systems, and the aforementioned methods cannot directly achieve the task of ensuring safety. Traditional methods for calculating Lyapunov functions can yield suitable functions in simple linear systems, but finding suitable functions in USV systems is difficult, and the found functions only cover a small portion of the Lyapunov stability region. Learning Lyapunov neural networks from simple dynamical models, while generally applicable to low-dimensional, discrete-state dynamical systems, cannot be directly applied to high-dimensional, continuous dynamical systems. Related research has largely remained at the level of simple experiments, such as inverted pendulums, without extending to more complex scenarios. Learning Lyapunov neural network controllers, while validating the controller with Lyapunov conditions during controller generation, fixes the Lyapunov function, failing to capture a significant portion of the Lyapunov stability region, resulting in insufficient exploration and difficulty in algorithm transfer. Furthermore, it cannot effectively integrate with systems that already have control algorithms. Summary of the Invention

[0007] The purpose of this invention is to address the shortcomings of the prior art by providing a control method and control system for unmanned systems based on Lyapunov neural networks, thereby solving at least one of the problems of the prior art.

[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0009] A control method for unmanned systems based on Lyapunov neural networks, characterized by including:

[0010] The Lyapunov function corresponding to the unmanned system was fitted by a Lyapunov neural network.

[0011] The unmanned system is guided to perform iterative training based on the safe zones defined by the Lyapunov neural network.

[0012] By integrating Lyapunov neural networks and model-based reinforcement learning agents for unmanned systems, the unmanned systems can be controlled.

[0013] Furthermore, it also includes training a Lyapunov neural network based on the observation state set of the unmanned system, wherein the input of the Lyapunov neural network is the working parameter data and working environment data of the unmanned system corresponding to the state, and the output of the Lyapunov neural network is the Lyapunov value corresponding to the state.

[0014] As a preferred approach, the state is in the decreasing region during the training of the Lyapunov neural network.

[0015] As a preferred approach, during the training of a Lyapunov neural network, if a state satisfies the defined safety set after a set number of time steps within a potential safe region, then that state is added to the safe set.

[0016] As a preferred approach, after each iteration of training, the Gaussian process model and the Lyapunov neural network are updated based on the latest sample set.

[0017] As a preferred approach, the model-based reinforcement learning agent of the unmanned system is obtained based on the filter probabilistic model predictive control algorithm; the model-based reinforcement learning agent that integrates the Lyapunov neural network and the unmanned system includes training the filter probabilistic model predictive control algorithm guided by the Lyapunov neural network to obtain a reward function guided by Lyapunov, and using the reward function to guide the control of the unmanned system.

[0018] As a preferred embodiment, the unmanned system is an unmanned boat, unmanned vehicle, drone, or robot.

[0019] As a preferred approach, when the unmanned system is an unmanned vessel, the training sample set data includes the real-time positioning data of the unmanned vessel, the speed and direction data of the unmanned vessel, and the wind speed and direction data of the environment in which the unmanned vessel is located; controlling the unmanned system includes controlling its engine throttle and / or rudder angle.

[0020] Based on the same inventive concept, this invention also provides a control system for an unmanned system based on a Lyapunov neural network, characterized by including:

[0021] Lyapunov function acquisition module: used to obtain the Lyapunov function corresponding to the unmanned system by fitting a Lyapunov neural network;

[0022] Iterative training module: used to guide the unmanned system to perform iterative training based on the safe area divided by the Lyapunov neural network;

[0023] Control module: Used to control the unmanned system after integrating Lyapunov neural network and model-based reinforcement learning agent.

[0024] As a preferred embodiment, the unmanned system is an unmanned boat, unmanned vehicle, drone, or robot.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] 1) By fitting the Lyapunov function through a Lyapunov neural network, most of the Lyapunov stable region can be covered, ensuring sufficient exploration of the safe region.

[0027] 2) It can be extended to more complex nonlinear systems and can learn Lyapunov neural networks in unmanned systems such as unmanned ships.

[0028] 3) It can be effectively transferred to other control algorithms, making it convenient to integrate with other algorithms. Attached Figure Description

[0029] Figure 1 This is an overall framework diagram of an unmanned system control method based on Lyapunov neural network according to an embodiment of the present invention (taking an unmanned ship as an example).

[0030] Figure 2 This diagram illustrates a control method for an unmanned surface vessel (USV) system based on a Lyapunov neural network, according to an embodiment of the present invention (taking an USV as an example). Detailed Implementation

[0031] To address the problems and shortcomings of existing technologies, and in order to explore safe areas more efficiently and completely, and to make the control process of unmanned systems such as unmanned ships more stable and efficient, so as to enable practical applications, this invention proposes a reinforcement learning unmanned system control method and control system based on Lyapunov neural networks for iterative learning.

[0032] The present invention proposes a reinforcement learning control method for unmanned systems (such as unmanned ships) based on Lyapunov neural networks for iterative learning. By iteratively learning Lyapunov neural networks, the method ensures the safety of the system and enables autonomous learning of unmanned systems (such as unmanned ships) without the need for prior human knowledge. This allows for safer and more effective control of unmanned systems (such as unmanned ships).

[0033] According to a first aspect of the present invention, the present invention provides a control method for unmanned systems based on Lyapunov neural networks, comprising:

[0034] The Lyapunov function corresponding to the unmanned system was fitted by a Lyapunov neural network.

[0035] The unmanned system is guided to perform iterative training based on the safe zones defined by the Lyapunov neural network.

[0036] By integrating Lyapunov neural networks and model-based reinforcement learning agents for unmanned systems, the unmanned systems can be controlled.

[0037] In some preferred embodiments, the Lyapunov neural network is further trained based on the observation state set of the unmanned system, wherein the input of the Lyapunov neural network is the working parameter data and working environment data of the unmanned system corresponding to the state, and the output of the Lyapunov neural network is the Lyapunov value corresponding to the state.

[0038] In some preferred embodiments, the state is in a decreasing region during the training of the Lyapunov neural network.

[0039] In some preferred embodiments, during the training of the Lyapunov neural network, if a state satisfies the definition of a set of safety sets after a set number of time steps within a potential safe region, then that state is added to the safe set.

[0040] In some preferred embodiments, after each iteration of training, the Gaussian process model and the Lyapunov neural network are updated based on the latest sample set.

[0041] In some preferred embodiments, the model-based reinforcement learning agent of the unmanned system is obtained based on a filtered probabilistic model predictive control algorithm; the model-based reinforcement learning agent that integrates the Lyapunov neural network and the unmanned system includes training the filtered probabilistic model predictive control algorithm guided by the Lyapunov neural network to obtain a reward function guided by Lyapunov, and using the reward function to guide the control of the unmanned system.

[0042] The unmanned system refers to unmanned ships, unmanned vehicles, drones, or robots, etc.

[0043] When the unmanned system is an unmanned vessel, the training sample set data includes the real-time positioning data of the unmanned vessel, the speed and direction data of the unmanned vessel, and the wind speed and direction data of the environment in which the unmanned vessel is located; controlling the unmanned system includes controlling its engine throttle and / or rudder angle.

[0044] According to a second aspect of the present invention, the present invention provides an unmanned system control system based on a Lyapunov neural network, wherein the unmanned system is an unmanned vessel, unmanned vehicle, drone, or robot, etc. The unmanned system control system includes:

[0045] Lyapunov function acquisition module: used to obtain the Lyapunov function corresponding to the unmanned system by fitting a Lyapunov neural network;

[0046] Iterative training module: used to guide the unmanned system to perform iterative training based on the safe area divided by the Lyapunov neural network;

[0047] Control module: Used to control the unmanned system after integrating Lyapunov neural network and model-based reinforcement learning agent.

[0048] This invention addresses the lack of safety considerations in unmanned systems such as unmanned ships, and solves the safety guarantee problem that requires prior human knowledge by utilizing Lyapunov neural networks. It extends Lyapunov neural networks to high-dimensional and complex nonlinear systems, and proposes a reinforcement learning-based unmanned system control method and control system based on iterative learning using Lyapunov neural networks.

[0049] This invention uses iterative training of a neural network to fit the Lyapunov function of an unmanned system (such as an unmanned vessel), and then uses the safe area defined by the Lyapunov neural network to guide the unmanned system (such as an unmanned vessel) in reinforcement learning training. By integrating the Lyapunov neural network with the model-based reinforcement learning algorithm of the unmanned system (such as an unmanned vessel), the stability control capability of the unmanned system (such as an unmanned vessel) is improved, resulting in more efficient and safer operation.

[0050] Taking an unmanned vessel as an example, the overall framework diagram of the control method of this invention is shown below. Figure 1 The technical solution and principles of this invention will be described in detail below.

[0051] I. Unmanned Surface Vessel Control System

[0052] (I) Lyapunov Neural Network

[0053] The core idea of ​​this invention is to use Lyapunov's second method to represent the stability of an unmanned surface vessel system from an energy perspective. For a given policy π, the safety set is defined as S. π Given a system and policy, any x∈S π Trajectories originating within the region remain within this region and gradually approach the equilibrium point. For time step t, the Lyapunov function v(x)... t The following conditions must be met:

[0054] v(0) = 0, v(x) t )>0 for x t ≠0, (1)

[0055] Δv(x t )=v(x t+1 )-v(x t )<0,x t ≠0, (2)

[0056] ||f′(x t+1 )-f′(x t )||≦k||x t+1 -x t ||,k∈R >0 (3)

[0057] We set up a special neural network to fit the Lyapunov function, v θ (x)=φ θ (x) · φ θ (x), φ θ It is a feedback-forward neural network. Lyapunov neural network v θ (x) should satisfy the above conditions for the Lyapunov function:

[0058] v θ (0) = 0; v θ (x)>0, x≠0;Δv θ (x)<0 (4)

[0059] To ensure a simple null space, i.e., satisfying condition (1), the activation function and weight matrix of each layer should also have a simple null space. The output dimension of each layer l is defined as d. l Determine the weight matrix W l For a d l ×d l-1 The matrix, the weight matrix should be a full-rank matrix, W l x = 0 has only the zero solution, and the weight matrix satisfies d l ≥d l-1 The condition (1) can be satisfied. The activation function of each layer and the neural network satisfy the Lipschitz continuity condition, which guarantees that the Lyapunov neural network satisfies the Lipschitz continuity condition, i.e., condition (3). Therefore, if a state satisfies condition (2) Δv θ If (x) < 0, it can be determined that the state is a Lyapunov stable state. This condition is handled during the training process.

[0060] (II) Processing Unmanned Surface Vessel Data Using Lyapunov Neural Networks

[0061] Data from an unmanned surface vessel with eight dimensions is used as input to a neural network, and the output of the network is the Lyapunov value for that state. This is done according to the formula i = 1, 2, ..., N. trial Iteratively training a Lyapunov neural network begins with initializing C using a small numerical value. i Used to describe the safe region, the initial approximate safe region S i Provided according to the following formula:

[0062] S i ≈V(x,C i )={x|v θ (x)<C i} (5)

[0063] In the current sample set X, determine the states that meet condition (3) as the decreasing region D.i :

[0064]

[0065] During training, ensuring the state remains within the decreasing region guarantees compliance with the Lyapunov stability condition, thus determining the safe set S. i for:

[0066] S i =V(x,C) i )={x|v θ (x)<C i},x∈D i (7)

[0067] In the reinforcement learning training process, a parameter α∈R is used. >1 Obtain an extended training set to explore potential safe zones:

[0068] G i =V(x,αC) i )-V(x,C i ),x∈D i (8)

[0069] In region G i Within the safe zone, if a state satisfies the definition of the safe zone (7) after h time steps, then the state is added to the safe zone. States located in potential safe zones and known safe zones are used as the training set, x∈V(x,αC). i And set the label for the safe state to y = +1, otherwise set it to y = -1, following the formula:

[0070]

[0071] The loss function consists of two parts: the first part is based on the misclassification and the Lyapunov value and C. i The second part penalizes the distance; the second part targets S. i Penalize states that violate the diminishing return condition (3):

[0072]

[0073] where λ∈R >0 It uses the Lagrange operator, and the model is updated using stochastic gradient descent. After training, C... i Update according to the following formula:

[0074] C i+1 =maxv θ (x), x∈S i (11)

[0075] In other words, the largest Lyapunov value in the safe set can be used as a critical value to divide the safe set into safe and unsafe regions.

[0076] (III) Integrating Lyapunov Neural Networks into Model-Based Reinforcement Learning Frameworks

[0077] The filtered probabilistic model predictive control is integrated with a Lyapunov neural network, and the learned Lyapunov neural network is introduced into the reward function of the filtered probabilistic model predictive control algorithm.

[0078]

[0079]

[0080] S was calculated based on the Lyapunov neural network. i This is then added to the reward function to guide the training of the unmanned surface vessel.

[0081] II. Unmanned Surface Vessel Hardware System

[0082] During the navigation of an unmanned surface vessel (USV) at sea, the engine and rudder play the primary control roles. Therefore, the control signals generated by this invention mainly control the engine throttle and the rudder angle. Controlling the throttle controls the USV's speed, while controlling the rudder angle controls steering. For the entire system to operate smoothly, the USV must collect data to represent its current state and send it to the control system to generate control signals. This requires hardware to perform this data collection. GPS is used to locate the USV's position and obtain its coordinates; a direction sensor is used to obtain the USV's speed and direction; and a wind sensor obtains wind speed and direction information. By integrating this hardware-collected information, the data is sent to the USV's control system to generate control signals specific to the USV, enabling it to navigate smoothly at sea.

[0083] III. Strengthening the Learning Framework

[0084] Because this invention does not rely on any prior human experience, the unmanned surface vessel (USV) needs to be able to learn autonomously through exploration. Reinforcement learning, as a type of machine learning, has been widely used due to its learning method that does not rely on prior knowledge. Therefore, this invention also incorporates reinforcement learning into the USV's control algorithm, enabling the USV to learn autonomously through exploration. This allows the USV to continuously optimize itself during navigation, making the model more accurate and resulting in superior control performance.

[0085] Taking unmanned systems as an example, unmanned ships Figure 2The diagram illustrates a control method for an unmanned surface vessel (USV) system based on a Lyapunov neural network, according to one embodiment of the present invention. The method begins by initializing parameters such as the sample set and loss function. Then, an initial Gaussian process model is trained using the sample set, and the Lyapunov neural network and C-value are initialized. Following this, iterative processing begins. During iteration, the Lyapunov neural network is first trained, and then predictive control using a filter probability model guided by the Lyapunov neural network is trained. After each iteration, the Gaussian process model and the Lyapunov neural network are updated based on the latest sample set. This entire process ensures that the USV can autonomously explore and learn in the ocean, optimizing the model through its continuous navigation "experience," thereby achieving better control performance.

[0086] The reinforcement learning-based unmanned surface vessel control method based on Lyapunov neural networks proposed in this invention has the following advantages compared to other existing technologies:

[0087] 1. Extend Lyapunov neural networks to high-dimensional data within model-based reinforcement learning frameworks, and extend Lyapunov neural networks to high-dimensional, continuous dynamical systems.

[0088] 2. By using Lyapunov neural networks to guide the learning of unmanned ships, the stability, robustness, and safety of unmanned ship control can be improved, and phenomena such as loss of control can be avoided.

[0089] The method of this invention has been verified by computer simulation, and the results are very good, demonstrating its feasibility.

[0090] This invention has good applicability. In addition to unmanned ship control, it can also be extended to unmanned systems such as unmanned vehicles, drones, and robots, and has broad application prospects.

[0091] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not limiting. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the scope of protection of the present invention.

Claims

1. A control method for unmanned systems based on Lyapunov neural networks, characterized in that, include: The Lyapunov function corresponding to the unmanned system was fitted by a Lyapunov neural network. The unmanned system is guided to perform iterative training based on the safe zones defined by the Lyapunov neural network. The unmanned system is controlled by integrating a Lyapunov neural network with a model-based reinforcement learning agent. The process of fitting the Lyapunov function corresponding to the unmanned system using the Lyapunov neural network includes: For a given policy Define the security set as Under a given system and strategy, any from Trajectories originating within the region remain within this region and gradually approach the equilibrium point; regarding time steps Lyapunov function The following conditions must be met: (1) (2) (3) Set up a neural network to fit the Lyapunov function. , It is a feedback-forward neural network; a Lyapunov neural network. The following conditions for the Lyapunov function should be met: ;(4) To ensure a simple null space, i.e., satisfying condition (1), the activation function and weight matrix of each layer should have a simple null space; define each layer The output dimension is Determine the weight matrix For one The weight matrix should be a full-rank matrix. Only the zero solution exists, and the weight matrix satisfies... The condition is satisfied by condition (1); the activation function of each layer and the neural network satisfy the Lipschitz continuity condition, which guarantees that the Lyapunov neural network satisfies the Lipschitz continuity condition, i.e., condition (3); therefore, if a state satisfies condition (2), then... This state is determined to be a Lyapunov stable state.

2. The unmanned system control method based on Lyapunov neural network according to claim 1, characterized in that, It also includes training a Lyapunov neural network based on the observation state set of the unmanned system, wherein the input of the Lyapunov neural network is the working parameter data and working environment data of the unmanned system corresponding to the state, and the output of the Lyapunov neural network is the Lyapunov value corresponding to the state.

3. The unmanned system control method based on Lyapunov neural network according to claim 2, characterized in that, During the training of a Lyapunov neural network, the state is in a decreasing region.

4. The unmanned system control method based on Lyapunov neural network according to claim 2, characterized in that, During the training of a Lyapunov neural network, if a state satisfies the definition of a set of safety sets after a set number of time steps within a potential safe region, then that state is added to the safe set.

5. The unmanned system control method based on Lyapunov neural networks according to any one of claims 1 to 4, characterized in that, After each iteration of training, the Gaussian process model and Lyapunov neural network are updated based on the latest sample set.

6. The unmanned system control method based on Lyapunov neural networks according to any one of claims 1 to 4, characterized in that, The model-based reinforcement learning agent of the unmanned system is obtained based on the filter probability model predictive control algorithm; the model-based reinforcement learning agent that integrates the Lyapunov neural network and the unmanned system includes training the filter probability model predictive control algorithm guided by the Lyapunov neural network to obtain a reward function guided by Lyapunov, and using the reward function to guide the control of the unmanned system.

7. The unmanned system control method based on Lyapunov neural networks according to any one of claims 1 to 4, characterized in that, The unmanned system refers to an unmanned boat, unmanned vehicle, drone, or robot.

8. The unmanned system control method based on Lyapunov neural network according to claim 7, characterized in that, When the unmanned system is an unmanned vessel, the training sample set data includes the real-time positioning data of the unmanned vessel, the speed and direction data of the unmanned vessel, and the wind speed and direction data of the environment in which the unmanned vessel is located; controlling the unmanned system includes controlling its engine throttle and / or rudder angle.

9. A control system for an unmanned system based on a Lyapunov neural network, characterized in that, The unmanned system control method based on Lyapunov neural network as described in any one of claims 1-8 is adopted; The control system of this unmanned system includes: Lyapunov function acquisition module: used to obtain the Lyapunov function corresponding to the unmanned system by fitting a Lyapunov neural network; Iterative training module: used to guide the unmanned system to perform iterative training based on the safe area divided by the Lyapunov neural network; Control module: Used to control the unmanned system after integrating Lyapunov neural network and model-based reinforcement learning agent.

10. The unmanned system control system based on Lyapunov neural network according to claim 9, characterized in that, The unmanned system refers to an unmanned boat, unmanned vehicle, drone, or robot.

Citation Information

Patent Citations

  • Robust control method based on reinforcement learning and Lyapunov function

    CN110928189A