Underactuated ship motion adaptive control method and device based on meta learning

Through the adaptive control method of under-driven ship motion based on meta-learning, the problem of low control accuracy and stability caused by model uncertainty in traditional ship control methods is solved, and higher control accuracy and stability are achieved.

CN120010251AActive Publication Date: 2025-05-16WUHAN UNIV OF TECH
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510091752.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Traditional ship control methods rely on accurate ship models, but due to uncertain interference factors such as water flow, waves, wind direction or wind speed, it is difficult to build an accurate model, resulting in low control accuracy and stability.

Method used

Adaptive control method for under-driven ship motion based on meta-learning is adopted. By obtaining ship trajectory data, a set of ship trajectory models are constructed, adaptive law is calculated, and meta-parameters are calculated through meta-learning to realize adaptive control of under-driven ship motion.

Benefits of technology

It improves the accuracy and stability of ship motion control, can better cope with uncertainties in complex environments, and achieves precise control of ship motion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010251A_ABST
    Figure CN120010251A_ABST
Patent Text Reader

Abstract

The invention discloses an underactuated ship motion adaptive control method and device based on meta-learning, and the method comprises the steps: obtaining ship trajectory data which comprises a ship state and initial control input; constructing a ship trajectory model set according to the ship trajectory data; calculating an adaptive law according to the control output vector of the controlled system and the control output vector of the reference model; calculating element parameters through element learning according to the ship trajectory model set and the adaptive law; and according to the meta parameters, carrying out under-actuated ship motion self-adaptive control. According to the invention, parameter calculation is carried out through meta-learning to realize ship motion control, and the accuracy and stability are improved. The method can be widely applied to the technical field of ship control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ship control technology, and in particular to a method and device for adaptive control of underactuated ship motion based on meta-learning. Background Art

[0002] Ships are complex dynamic systems that are affected by many factors. Traditional ship control methods rely on accurate ship models, but the actual operating environment of ships has uncertain interference factors such as water flow, waves, wind direction or wind speed, which makes it difficult to build accurate models and determine model parameters. At the same time, it is difficult to determine the nonlinear characteristics of the ship system a priori, and the parameter adjustment is complex, which can easily cause overfitting or underfitting problems, resulting in low accuracy and stability of ship motion control.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention

[0004] The embodiments of the present invention provide a method and device for adaptive control of underactuated ship motion based on meta-learning, which effectively improves accuracy and stability.

[0005] On the one hand, an embodiment of the present invention provides an underactuated ship motion adaptive control method based on meta-learning, comprising the following steps:

[0006] Acquiring ship trajectory data, wherein the ship trajectory data includes a ship state and an initial control input;

[0007] Constructing a set of ship trajectory models according to the ship trajectory data;

[0008] Calculating an adaptive law according to a controlled system control output vector and a reference model control output vector;

[0009] Calculating meta-parameters through meta-learning according to the ship trajectory model set and the adaptive law;

[0010] The motion adaptive control of the underactuated ship is performed according to the meta-parameters.

[0011] In some embodiments, constructing a set of ship trajectory models according to the ship trajectory data includes:

[0012] According to the ship state and the initial control input, a preset neural network is used to perform dynamic characteristic fitting to obtain an initial ship trajectory model;

[0013] Calculating a predicted value of a ship state according to the ship state, the initial control input, the model parameters and the dynamic relationship function;

[0014] Constructing an objective function according to the predicted value of the ship state, the regularization coefficient, the meta-parameters and the number of sampling moments;

[0015] According to the objective function, the initial ship trajectory model is trained by using stochastic gradient descent to obtain a target ship trajectory model;

[0016] A plurality of the target ship trajectory models are combined to obtain the ship trajectory model set.

[0017] In some embodiments, the process of constructing the preset neural network includes:

[0018] Constructing an input layer, wherein the input layer is used to receive the current state of the ship and the initial control input;

[0019] After the input layer, a hidden layer is constructed, wherein the hidden layer is used to calculate the next state of the ship according to the current state and the initial control input using a preset nonlinear activation function, and the number of the hidden layers is 2;

[0020] After the hidden layer, the output layer is constructed.

[0021] In some embodiments, the step of calculating the adaptive law according to the controlled system control output vector and the reference model control output vector comprises:

[0022] Calculating a state vector of a controlled system according to the controlled system control output vector;

[0023] Calculating a control law according to the controlled system state vector, a reference model control input, an initial feedforward controller and an initial state feedback controller;

[0024] Using the control law as a control input of the controlled system;

[0025] Constructing a state space equation group according to the controlled system state vector, the controlled system control input, the controlled system state matrix and the controlled system input matrix;

[0026] Calculate a reference model state vector according to the reference model control output vector;

[0027] constructing a reference model according to the reference model state vector, the reference model control input, the reference model state matrix and the reference model input matrix;

[0028] Calculating the error between the reference output and the actual output according to the state vector of the controlled system and the state vector of the reference model;

[0029] Calculating a system state error based on an error between the reference output and the actual output, the state space equations and the reference model;

[0030] updating the system state error according to the system state convergence equation;

[0031] The adaptive law is calculated according to the system state error.

[0032] In some embodiments, updating the system state error according to the system state convergence equation includes:

[0033] Constructing a target feedforward controller according to the initial feedforward controller, the system state convergence equation and the error between the reference output and the actual output;

[0034] Constructing a target state feedback controller according to the initial state feedback controller, the system state convergence equation and the error between the reference output and the actual output;

[0035] According to the target feedforward controller and the target state feedback controller, a matching equation group is constructed so that the reference model and the controlled system can be fully matched;

[0036] Calculating a feedforward controller difference according to the initial feedforward controller and the target feedforward controller;

[0037] Calculating a state feedback controller difference according to the initial state feedback controller and the target state feedback controller;

[0038] The system state error is updated according to the matching equation group, the feedforward controller difference and the state feedback controller difference.

[0039] In some embodiments, calculating the adaptive law according to the system state error includes:

[0040] Constructing a Lyapunov function according to the reference model state matrix and the positive definite matrix;

[0041] Solving the Lyapunov function to obtain a Lyapunov function value;

[0042] Constructing a real function according to the error between the reference output and the actual output, the Lyapunov function value, the feedforward controller difference and the state feedback controller difference;

[0043] Derivative the real function to obtain a derivative equation of the real function;

[0044] Performing an overall negative definite transformation on the real function derivative equation to obtain a negative definite equation system;

[0045] The adaptive law is constructed according to the negative definite equation group and the gain matrix.

[0046] In some embodiments, calculating meta-parameters by meta-learning according to the ship trajectory model set and the adaptive law includes:

[0047] According to the set of ship trajectory models, a plurality of ship tasks are set;

[0048] Setting an interference signal corresponding to each of the ship tasks;

[0049] Constructing the meta-parameters according to the adaptive law, the parameters related to the pi control input function and the parameters related to the density control input function;

[0050] Calculate the target ship state information according to the initial ship state information and the adaptive input function;

[0051] Calculating a target adaptive mechanism-related quantity according to the meta-parameter, the initial adaptive mechanism-related quantity, the density control input function and the interference signal;

[0052] Calculating control input information according to the meta-parameters, the reference trajectory and the pi control input function;

[0053] Calculating a key information set according to the target ship state information, the target adaptive mechanism related quantity and the control input information;

[0054] Calculating an average tracking error based on the reference trajectory and the interference signal;

[0055] Constructing a loss function according to the key information set, the average tracking error and the control gain adjustment parameter;

[0056] Constructing a meta-problem according to the meta-parameters, the actual state of the ship, the reference trajectory, the control gain adjustment parameter, the initial control input information, the regularization parameter, the number of reference trajectories and the number of interference signals;

[0057] According to the loss function, the meta-problem is solved by using adaptive gradient descent to update the meta-parameters.

[0058] In some embodiments, solving the meta-problem by adaptive gradient descent according to the loss function and updating the meta-parameters includes:

[0059] Initialize the diagonal matrix and the values ​​of the element parameters;

[0060] According to the ship mission, calculating the gradient of the objective function with respect to the meta-parameters;

[0061] updating the diagonal matrix according to the gradient of the objective function with respect to the element parameters;

[0062] Calculate the adaptive learning rate according to the auxiliary positive number, the global learning rate and the updated diagonal matrix;

[0063] The meta-parameters are updated according to the adaptive learning rate and the loss function until a change in the meta-parameters is less than a preset change threshold.

[0064] On the other hand, an embodiment of the present invention provides an underactuated ship motion adaptive control device based on meta-learning, comprising:

[0065] A data acquisition module, used to obtain ship trajectory data, wherein the ship trajectory data includes ship status and initial control input;

[0066] A system dynamics feature learning module, used for constructing a set of ship trajectory models based on the ship trajectory data;

[0067] An adaptive controller module, used for calculating an adaptive law according to a controlled system control output vector and a reference model control output vector;

[0068] A two-layer meta-learning optimization module, used for calculating meta-parameters through meta-learning according to the set of ship trajectory models and the adaptive law;

[0069] A control module is used for performing adaptive control of the motion of the underactuated ship according to the meta-parameters.

[0070] In another aspect, an embodiment of the present invention provides a computer device, comprising:

[0071] at least one processor;

[0072] at least one memory for storing at least one program;

[0073] When the at least one program is executed by the at least one processor, the at least one processor implements the described method.

[0074] The beneficial effects of the present invention are as follows:

[0075] The embodiment of the present invention first obtains ship trajectory data, constructs a ship trajectory model set according to the ship trajectory data, then calculates the adaptive law according to the controlled system control output vector and the reference model control output vector, and then calculates meta-parameters through meta-learning according to the ship trajectory model set and the adaptive law, and finally performs adaptive control of the under-actuated ship motion according to the meta-parameters, so that ship motion control can be realized by parameter calculation through meta-learning, thereby improving accuracy and stability.

[0076] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0078] Figure 1 This is a flow chart of a method for adaptive motion control of an underactuated ship based on meta-learning according to an embodiment of the present invention;

[0079] Figure 2 This is a meta-learning training flow chart of an embodiment of the present invention;

[0080] Figure 3 A schematic diagram of a ship motion controller structure based on model reference adaptive control according to an embodiment of the present invention;

[0081] Figure 4 A schematic diagram of an architecture for controlling ship motion according to an embodiment of the present invention;

[0082] Figure 5 This is a schematic diagram of the structure of an adaptive control device for underactuated ship motion based on meta-learning according to an embodiment of the present invention;

[0083] Figure 6 The figure is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.

[0085] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".

[0086] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0087] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0088] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0089] Meta-learning refers to the learning of individuals' acquisition of learning mechanisms, which involves the question of how individuals acquire the functions they rely on for learning.

[0090] In the related technologies, the classical control algorithm relies on an accurate ship model, but as a complex dynamic system, the ship is affected by many factors, and it is difficult to build an accurate model. At the same time, the actual operating environment is complex, and there are uncertain interference factors such as water flow, waves, wind direction, and wind speed, which lead to uncertainty in model parameters and affect the effect of the classical control algorithm. Although adaptive control can solve some model uncertainty problems, for under-actuated nonlinear systems such as ships, some nonlinear characteristics are difficult to determine a priori, which limits its application effect. Although the data-driven control method does not require an accurate model, the parameter adjustment in the algorithm design is complex, and it is prone to overfitting or underfitting problems, and it is difficult to fully consider the nonlinear and under-actuated characteristics of the ship system. Meta-learning uses tasks as samples and realizes the generalization of new tasks by learning feature representations. Only a small number of samples and iterations are needed to train a model with generalization ability, and the model can fine-tune parameters to adapt to new tasks. In the field of ship control, meta-learning can learn the initial parameters of the neural network, so that the ship has the ability to "learn to learn", simple modeling and flexible control in the new environment, and can better deal with the problems faced by ship motion control. Although the existing data-driven control methods have the advantage of not requiring precise models, there are many parameters that need to be adjusted during the algorithm design process, which increases the complexity of the control process to a certain extent. Excessive parameter adjustments not only increase the computational burden, but may also lead to problems such as overfitting or underfitting, which in turn affects the performance and generalization ability of the model in practical applications. Moreover, when faced with complex ship motion control scenarios, these methods are difficult to fully consider the nonlinear and under-driven characteristics of the ship system, cannot effectively deal with the model uncertainty caused by environmental interference, and are difficult to achieve precise control of ship motion. The accuracy and stability of ship motion control are low.

[0091] In view of this, this embodiment is guided by adaptive control, and by training the model on various learning tasks, it can solve new learning tasks with only a small number of training samples. After training on various tasks, the parameters of the model can better fit the nonlinear relationship of the ship system. When faced with different control tasks, the model can quickly adjust its own parameters to adapt to new task requirements. In addition, this method can effectively deal with the model uncertainty caused by environmental interference, ensuring that the ship can move according to the expected trajectory in a complex environment, thereby improving the accuracy and stability of ship motion control. It mainly includes a system dynamics feature learning module, an adaptive controller module, and a two-layer meta-learning optimization module.

[0092] In the system dynamics feature learning module, the dynamics feature information of the ship in different environments is obtained by analyzing and processing the ship trajectory data. This module first pre-processes the collected trajectory data containing the ship state and control input information, fits a model with parameters to each trajectory under reasonable interference assumptions, and trains the model parameters through methods such as gradient descent. Finally, multiple trained models are combined into a model integration to construct a meta-problem and learn the meta-parameters related to the system dynamics.

[0093] In the adaptive controller module, a suitable adaptive controller structure is designed based on the dynamic characteristics of the ship and the control objectives. This module considers the ship as a dynamic system and determines the specific form of the adaptive controller according to its nonlinear dynamic equations. When some information such as external forces is unknown, neural networks and other methods can be used for approximate processing.

[0094] In the two-layer meta-learning optimization module, it plays a role of bridge and optimization between the system dynamics feature learning module and the adaptive controller module. It is guided by adaptive control, and by establishing multiple loss functions, training data sets and evaluation data sets corresponding to multiple tasks, and defining an adaptation mechanism to map meta-parameters and task-specific training data to task-specific parameters, in order to solve the two-layer problem and find the optimal meta-parameters. In this process, meta-learning is performed from feedback and adaptation, and the static parameters of the adaptive controller are learned using reference trajectories and interference signals as training data. The task loss is calculated through the forward simulation of the closed-loop system, and the meta-problem is constructed and solved.

[0095] The meta-learning-based adaptive control method for underactuated ship motion provided in the embodiment of the present application relates to the field of ship control technology. The meta-learning-based adaptive control method for underactuated ship motion provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms and other basic cloud computing services. The server can also be a node server in a blockchain network; the software can be an application that implements the meta-learning-based adaptive control method for underactuated ship motion, etc., but is not limited to the above forms.

[0096] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0097] The following is a detailed explanation of the embodiments of the present application in conjunction with the accompanying drawings:

[0098] Figure 1 is an optional flow chart of the underactuated ship motion adaptive control method based on meta-learning provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S105.

[0099] Step S101, obtaining ship trajectory data, the ship trajectory data including ship status and initial control input;

[0100] Step S102: construct a ship trajectory model set according to the ship trajectory data;

[0101] Step S103, calculating the adaptive law according to the controlled system control output vector and the reference model control output vector;

[0102] Step S104, calculating meta-parameters through meta-learning according to the ship trajectory model set and the adaptive law;

[0103] Step S105: performing adaptive control of the underactuated ship motion according to the meta-parameters.

[0104] Steps S101 to S105 shown in the embodiment of the present application realize ship motion control and improve accuracy and stability.

[0105] In step S101 of some embodiments, the ship trajectory data can be obtained through the simulator test simulation platform. The ship trajectory data can also be obtained by other means, not limited thereto. Among them, the ship trajectory data includes the ship state and the initial control input. Exemplarily, trajectory control can be performed in the open waters on the simulator test simulation platform under undisturbed conditions, level 3 sea conditions, and level 5 sea conditions to collect multiple sets of ship trajectory data. Among them, under undisturbed conditions, 100 sets of ship trajectory data are collected, each set of ship trajectory data contains 500 sampling points; under level 3 sea conditions, 200 sets of ship trajectory data are collected, each set of ship trajectory data contains 1000 sampling points; under level 5 sea conditions, 300 sets of ship trajectory data are collected, each set of ship trajectory data contains 1500 sampling points. It can be understood that the number of samples can be set according to different working conditions. Under undisturbed conditions, the ship motion law is stable, and this number can meet the modeling requirements. Under level 3 sea conditions, the environment is relatively mild and the ship motion is relatively stable, but more data is still needed to reflect its motion characteristics. Under level 5 sea conditions, the sea conditions are relatively severe and the ship's motion changes are more complex. More data and longer sampling are required to more comprehensively reflect the ship's motion characteristics under this sea condition. Furthermore, during the movement of the ship, the ship's state and control input under the above three conditions are sampled at certain time intervals. Among them, the information on the ship's state may include the ship's position (longitude, latitude), speed or heading, etc., and the control input information may include the rudder angle or propeller speed. By sorting and preprocessing a large amount of original sampled ship trajectory data, the ship's trajectory data in different time periods is obtained. Each track Where M represents the total number of groups of ship trajectory data, and the value range of j is 1 to M. represents the time value corresponding to the kth sampling moment in the jth time period, represents the status information of the ship at the kth sampling moment in the jth time period, represents the control input at the kth sampling moment in the jth time period, and N represents the number of samples in each group.

[0106] In some embodiments, in step S102, constructing a set of ship trajectory models according to the ship trajectory data may include but is not limited to the following steps:

[0107] According to the ship state and initial control input, the preset neural network is used to fit the dynamic characteristics to obtain the initial ship trajectory model;

[0108] Calculate the predicted value of the ship state according to the ship state, initial control input, model parameters and dynamic relationship function;

[0109] Construct the objective function according to the predicted value of ship state, regularization coefficient, meta-parameters and number of sampling moments;

[0110] According to the objective function, the model parameters of the initial ship trajectory model are trained using stochastic gradient descent to obtain the target ship trajectory model;

[0111] Multiple target ship trajectory models are combined to obtain a ship trajectory model set.

[0112] In some embodiments, each trajectory τ may be first calculated based on the ship state x and the initial control input u. j , the preset neural network is used to fit the dynamic characteristics to obtain the initial ship trajectory model. It can be understood that by fitting the data of each trajectory, the model It can describe the ship dynamic characteristics corresponding to the trajectory as accurately as possible, where is the model parameter. The model parameters of different trajectories are different, reflecting the differences in ship dynamics under different tasks and environments. Then, according to the ship state, initial control input, model parameters and dynamic relationship function, the ship state prediction value is calculated, where the calculation formula of the ship state prediction value is: In the formula, is the ship status prediction value of the model for the ship status at the k+1th sampling moment in the jth time period, is the actual status information of the ship at the kth sampling moment in the jth time period, x(t) is the ship status, is the initial control input, is the model parameter corresponding to the jth trajectory, is the dynamic relationship function, that is, the function expression related to the fitting model corresponding to the jth trajectory, which describes the dynamic relationship of the ship state changing with time. Then, according to the ship state prediction value, regularization coefficient, meta-parameters and number of sampling moments, the objective function is constructed, where the expression of the objective function is: Where N j represents the number of sampling moments in the jth trajectory, represents the actual state information of the ship at the k+1th sampling moment in the jth time period, μ is the regularization coefficient, and θ is the meta-parameter.

[0113] Finally, according to the objective function, the model parameters of the initial ship trajectory model are trained using stochastic gradient descent to obtain the target ship trajectory model, and multiple target ship trajectory models are combined to obtain a ship trajectory model set. Exemplarily, during the training process, Adagrad, a variant of the stochastic gradient descent algorithm, can be used for optimization. All fitted target ship trajectory models are combined into a ship trajectory model set. It can be understood that the ship trajectory model set refers to a set of multiple models for subsequent meta-learning and control decision-making processes. In order to improve the accuracy and generalization ability of the model integration, different weights are assigned according to the performance of each target ship trajectory model on the validation data set. Models with better performance are given higher weights, and models with poorer performance are given lower weights, so that the ship trajectory model set can better reflect the dynamic characteristics of the ship as a whole and improve its adaptability to different tasks and environments.

[0114] In some embodiments, the process of constructing a preset neural network includes:

[0115] Construct an input layer, which is used to receive the current state of the ship and the initial control input;

[0116] After the input layer, a hidden layer is constructed. The hidden layer is used to calculate the next state of the ship according to the current state and the initial control input using a preset nonlinear activation function. The number of hidden layers is 2;

[0117] After the hidden layers, the output layer is constructed.

[0118] In some embodiments, one input layer, two hidden layers and one output layer may be constructed in sequence, wherein the input layer is used to receive the current state of the ship and the initial control input, the hidden layer is used to calculate the next state of the ship according to the current state and the initial control input using a preset nonlinear activation function, and the output layer outputs the predicted next state of the ship.

[0119] In some embodiments, in step S103, the adaptive law is calculated according to the controlled system control output vector and the reference model control output vector, which may include but is not limited to the following steps:

[0120] Calculate the state vector of the controlled system according to the control output vector of the controlled system;

[0121] Calculating a control law based on a state vector of a controlled system, a reference model control input, an initial feedforward controller, and an initial state feedback controller;

[0122] Taking the control law as the control input of the controlled system;

[0123] Constructing a state space equation group according to a controlled system state vector, a controlled system control input, a controlled system state matrix and a controlled system input matrix;

[0124] Calculate the reference model state vector according to the reference model control output vector;

[0125] constructing a reference model according to a reference model state vector, a reference model control input, a reference model state matrix, and a reference model input matrix;

[0126] Calculate the error between the reference output and the actual output according to the state vector of the controlled system and the state vector of the reference model;

[0127] Calculate the system state error based on the error between the reference output and the actual output, the state space equations and the reference model;

[0128] According to the system state convergence equation, the system state error is updated;

[0129] Based on the system state error, the adaptive law is calculated.

[0130] In some embodiments, the state vector of the controlled system can be calculated according to the control output vector of the controlled system, wherein the calculation formula of the state vector of the controlled system is: s =Y s , where X s is the state vector of the controlled system, Y s is the control output vector of the controlled system. Then, the control law is calculated according to the state vector of the controlled system, the reference model control input, the initial feedforward controller and the initial state feedback controller. The calculation formula of the control law is: U s =FX s +KU m , where U s is the control law, F is the initial state feedback controller, K is the initial feedforward controller, U m is the reference model control input. The control law is used as the controlled system control input. According to the controlled system state vector, the controlled system control input, the controlled system state matrix and the controlled system input matrix, a state space equation group is constructed, where the expression of the state space equation group is: In the formula, is the derivative of the state vector of the controlled system, A s is the state matrix of the controlled system, B s is the input matrix of the controlled system. According to the reference model control output vector, the reference model state vector is calculated, where the calculation formula of the reference model state vector is: m =Y m , where X m is the reference model state vector, Y mis the reference model control output vector. According to the reference model state vector, the reference model control input, the reference model state matrix and the reference model input matrix, the reference model is constructed, where the expression of the reference model is: In the formula, is the derivative of the reference model state vector, A m is the reference model state matrix, B m Input matrix for the reference model, U m is the reference model control input. According to the state vector of the controlled system and the state vector of the reference model, the error between the reference output and the actual output is calculated. The calculation formula of the error between the reference output and the actual output is: x =X m -X s , where e x is the error between the reference output and the actual output. According to the error between the reference output and the actual output, the state space equations and the reference model, the system state error is calculated, where the calculation formula of the system state error is: Then, according to the system state convergence equation, the system state error is updated, and finally, the adaptive law is calculated according to the system state error.

[0131] In some embodiments, updating the system state error according to the system state convergence equation includes:

[0132] Constructing a target feedforward controller according to the initial feedforward controller, the system state convergence equation and the error between the reference output and the actual output;

[0133] According to the initial state feedback controller, the system state convergence equation and the error between the reference output and the actual output, a target state feedback controller is constructed;

[0134] According to the target feedforward controller and the target state feedback controller, a matching equation group is constructed to achieve a complete match between the reference model and the controlled system;

[0135] Calculating a feedforward controller difference according to the initial feedforward controller and the target feedforward controller;

[0136] Calculate the state feedback controller difference according to the initial state feedback controller and the target state feedback controller;

[0137] The system state error is updated based on the matching equations, the feedforward controller difference and the state feedback controller difference.

[0138] In some embodiments, a target feedforward controller can be constructed based on the initial feedforward controller, the system state convergence equation, and the error between the reference output and the actual output, wherein the expression of the target feedforward controller is: K(e x ,t)=K0 , where K is the target feedforward controller, K 0 is the initial feedforward controller, e x is the error between the reference output and the actual output, and t is the time. According to the initial state feedback controller, the system state convergence equation and the error between the reference output and the actual output, the target state feedback controller is constructed, where the expression of the target state feedback controller is: F(e x ,t)=F 0 , where F is the target state feedback controller, F 0 is the initial state feedback controller. According to the target feedforward controller and the target state feedback controller, a matching equation group is constructed to achieve a complete match between the reference model and the controlled system, where the expression of the matching equation group is: In the formula, A s is the state matrix of the controlled system, B s is the input matrix of the controlled system, A m is the reference model state matrix, B m is the reference model input matrix. According to the initial feedforward controller and the target feedforward controller, the feedforward controller difference is calculated, where the calculation formula of the feedforward controller difference is: In the formula, is the feedforward controller difference. According to the initial state feedback controller and the target state feedback controller, the state feedback controller difference is calculated, where the calculation formula of the state feedback controller difference is: In the formula, is the state feedback controller difference. Finally, according to the matching equations, the feedforward controller difference and the state feedback controller difference, the system state error is updated to eliminate A s and B s , the expression of the updated system state error is: It is understandable that in order to achieve system state convergence, Among them, t represents time, e(t) represents the function of error changing with time, and when time tends to infinity, the error approaches 0, indicating that the system is stable.

[0139] In some embodiments, calculating the adaptive law according to the system state error includes:

[0140] According to the reference model state matrix and positive definite matrix, construct the Lyapunov function;

[0141] Solve the Lyapunov function and obtain the value of the Lyapunov function;

[0142] Construct a real function based on the error between the reference output and the actual output, the Lyapunov function value, the feedforward controller difference, and the state feedback controller difference;

[0143] Derivative the real function to obtain the derivative equation of the real function;

[0144] Perform overall negative definite transformation on the derivative equation of real function to obtain negative definite equation system;

[0145] According to the negative definite equations and the gain matrix, the adaptive law is constructed.

[0146] In some embodiments, a Lyapunov function may be constructed based on a reference model state matrix and a positive definite matrix, wherein the expression of the Lyapunov function is: m T +PA m +Q=0, where A m is the reference model state matrix, Q is a positive definite matrix, and P is the Lyapunov function value. Then the Lyapunov function is solved to obtain the Lyapunov function value P. According to the error between the reference output and the actual output, the Lyapunov function value, the feedforward controller difference, and the state feedback controller difference, a real function is constructed, where the expression of the real function is: In the formula, e x is the error between the reference output and the actual output, F is the target state feedback controller, is the feedforward controller difference, is the state feedback controller difference, P F is the Lyapunov function value of the state feedback controller, P K is the Lyapunov function value of the feedforward controller. By taking the derivative of the real function, we get the derivative equation of the real function, where the expression of the derivative equation of the real function is: In the formula, B m Input matrix for the reference model, U m is the reference model control input, tr represents the trace of the matrix, express The derivative of express The derivative of x T Indicates e x The transpose of Indicates e x The derivative of K 0 -1 K 0 The inverse matrix, P K -1 Indicates P K The inverse matrix of express The transpose of .

[0148] Then, the derivative equation of the real function is transformed into a negative definite system, and the negative definite system is obtained. The expression of the negative definite system is: In the formula, represents the derivative of F, represents the derivative of K, K 0 -T K 0 The negative transpose of B m T For B m The transpose of X s T For X s The transpose of U m T For U m Finally, based on the negative definite equations and the gain matrix, the adaptive law is constructed, where the expression of the adaptive law is: In the formula, γ 1 and γ 2 are all gain matrices. It can be understood that the adaptive law followed by parameters K and F provides the basic control logic for subsequent meta-learning optimization. The calculated adaptive law focuses on real-time adjustment of controller parameters according to information such as system state error within a single control cycle to ensure that the system stably tracks the reference trajectory, while the subsequent two-layer meta-learning optimization uses a large number of different tasks to optimize controller parameters (including the optimal settings of parameters such as K and F under different tasks) from the perspective of multiple tasks and the overall meta-learning framework. The two work together.

[0149] In some embodiments, in step S104, the meta-parameters are calculated by meta-learning according to the ship trajectory model set and the adaptive law, which may include but is not limited to the following steps:

[0150] Set multiple ship missions based on a collection of ship trajectory models;

[0151] Set the jamming signal corresponding to each ship mission;

[0152] Constructing meta-parameters according to the adaptive law, the parameters related to the pi control input function and the parameters related to the density control input function;

[0153] Calculate the target ship state information according to the initial ship state information and the adaptive input function;

[0154] Calculating target adaptive mechanism related quantities according to the meta-parameters, the initial adaptive mechanism related quantities, the density control input function and the interference signal;

[0155] Calculate control input information according to meta-parameters, reference trajectory and pi control input function;

[0156] Calculate the key information set according to the target ship state information, target adaptive mechanism related quantities and control input information;

[0157] Calculate the average tracking error based on the reference trajectory and the interference signal;

[0158] Construct a loss function based on key information set, average tracking error and control gain adjustment parameters;

[0159] A meta-problem is constructed according to meta-parameters, actual ship state, reference trajectory, control gain adjustment parameter, initial control input information, regularization parameter, number of reference trajectories and number of interference signals;

[0160] According to the loss function, the meta-problem is solved using adaptive gradient descent and the meta-parameters are updated.

[0161] In some embodiments, a meta-learning framework can be built to define multiple ship tasks, each of which corresponds to a reference trajectory and an interference signal. A meta-problem is constructed by combining the elements of all ship tasks, and the parameters of the adaptive controller are obtained by gradient descent solution of the meta-problem. The elements mainly include reference trajectories, interference signals, ship status information, control input information, and parameters related to the adaptive controller. Multiple ship tasks can be set according to a set of ship trajectory models, and interference signals corresponding to each ship task can be set. For example, after collecting trajectory data T j In the process of learning, considering the complexity of the actual situation and in order to simplify the model learning process, it is assumed that the interference w(t) takes a fixed unknown value w j . More importantly, the reference trajectory r i (t) and interference signal w j (t) is the training set for ship task (i, j) within the time range T Among them, the ship task (i, j) represents the task set consisting of the i-th reference trajectory and the j-th interference signal. Then, according to the adaptive law, the parameters related to the pi control input function and the parameters related to the density control input function, the meta-parameters are constructed so that the reference trajectory can be well tracked under given interference. The expression of the meta-parameters is:

[0162] θ=(θ π ,θ ρ ), where θ is a meta-parameter, θ π is the parameter related to the pi control input function, which is related to the control input function π, θ ρIt is a parameter related to the density control input function and is related to the control input function ρ. It can be understood that the parameter θ is a generalized set of adaptive controller parameters, which is used to learn the parameters that enable the controller to have a good tracking effect on the reference trajectory under a given disturbance under the meta-learning framework. It is an extension and optimization of the feedforward controller K and the state feedback controller F at the meta-learning level. Through the optimization of this embodiment, the parameters in the adaptive law can obtain better initial values ​​and be adaptively adjusted according to the conditions in more mission scenarios, thereby improving the accuracy, generalization ability and adaptability of the entire ship motion control in different environments and different tasks.

[0163] Then, the target ship state information is calculated according to the initial ship state information and the adaptive input function, where the calculation formula of the target ship state information is: In the formula, x ij (t) is the target ship status information, x ij (0) is the initial ship state information, f is the adaptive input function, x ij (t) is the state information of the ship at time t under ship mission (i, j), w j (t) is the interference signal. According to the meta-parameters, the initial adaptive mechanism related quantity, the density control input function and the interference signal, the target adaptive mechanism related quantity is calculated, where the calculation formula of the target adaptive mechanism related quantity is: In the formula, a ij (t) is the target adaptive mechanism related quantity, a ij (0) is the initial adaptive mechanism related quantity, ρ is the density control input function. According to the meta-parameters, reference trajectory and pi control input function, the control input information is calculated, where the calculation formula of the control input information is: ij (t) = π(x ij (t),r i (t),a ij (t);θ π ), where u ij (t) is the control input information, r i (t) is the reference trajectory, and π is the pi control input function. According to the target ship state information, target adaptive mechanism related quantities and control input information, the key information set is calculated, where the calculation formula of the key information set is: In the formula, is the key information set, and T is the total time.

[0164] Then, the average tracking error is calculated based on the reference trajectory and the interference signal. The expression of the average tracking error is: ij = {r i (t),w j (t)}t∈[0,T] , where D ij is the average tracking error. According to the key information set, the average tracking error and the control gain adjustment parameter, a loss function is constructed, where the expression of the loss function is: Where, L ij is the loss function, α is the control gain adjustment parameter, α≥0. It can be understood that the loss function is obtained by comprehensive calculation based on the deviation between the actual state of the ship and the reference trajectory and the control input. The smaller its value, the better the control effect, that is, the better the ship can track the reference trajectory. According to the meta-parameters, the actual state of the ship, the reference trajectory, the control gain adjustment parameter, the initial control input information, the regularization parameter, the number of reference trajectories and the number of interference signals, the meta-problem is constructed, where the expression of the meta-problem is: In the formula, μ meta is the regularization parameter, N is the number of reference trajectories, and M is the number of interference signals. According to the loss function, the meta-problem is solved using adaptive gradient descent to update the meta-parameters. It can be understood that in the meta-problem, n reference trajectories are constructed. And sampled M interference signals

[0165] In some embodiments, the meta-learning training flowchart is as follows: Figure 2 As shown, we can first train the data to learn the parameters of the adaptive controller, then calculate the adaptation mechanism and task loss, build a meta-problem based on model integration, and use Adagrad gradient descent to solve the meta-parameters. Finally, we determine whether it converges. If it converges, the meta-learning training is completed.

[0166] In some embodiments, according to the loss function, the meta-problem is solved by using adaptive gradient descent to update the meta-parameters, including:

[0167] Initialize the values ​​of the diagonal matrix and meta-parameters;

[0168] According to the ship's mission, the gradient of the objective function with respect to the meta-parameters is calculated;

[0169] Update the diagonal matrix according to the gradient of the objective function with respect to the meta-parameters;

[0170] Calculate the adaptive learning rate based on the auxiliary positive number, the global learning rate and the updated diagonal matrix;

[0171] According to the adaptive learning rate and loss function, the meta-parameters are updated until the change in the meta-parameters is less than the preset change threshold.

[0172] In some embodiments, the values ​​of the diagonal matrix and the meta-parameters can be initialized first. For example, the meta-parameter θ can be initialized to a random value or an initial value set according to prior knowledge. At the same time, a diagonal matrix H=0 is initialized, whose dimension is the same as the meta-parameter θ, and is used to record the historical sum of square gradients of each parameter. Then, according to the ship mission, the gradient of the objective function with respect to the meta-parameters is calculated. For example, a ship mission (i, j) can be randomly selected, and the ship mission is determined by the reference trajectory r i (t) and interference signal w j (t) is composed of calculating the gradient h of the objective function of the current task with respect to the meta-parameter θ ij Then, according to the gradient of the objective function with respect to the element parameters, the diagonal matrix is ​​updated. For example, the diagonal matrix H = H + h can be updated. ij (θ) 2 , add the square of the current gradient to the corresponding diagonal element of H to record the historical sum of the squares of the gradients of each parameter. According to the auxiliary positive number, the global learning rate and the updated diagonal matrix, the adaptive learning rate is calculated, where the adaptive learning rate is calculated as: Where η ij is the adaptive learning rate, η is an initially set global learning rate, H is the updated diagonal matrix, and ξ is an auxiliary positive number, i.e., a very small positive number, used to avoid the situation where the denominator is zero. Finally, according to the adaptive learning rate and the loss function, the meta-parameters are updated until the change in the meta-parameters is less than the preset change threshold. Exemplarily, when the objective function value of the meta-problem no longer decreases significantly and the change in the meta-parameter θ is less than the preset change threshold, that is, for all meta-parameters θ i , after T consecutive iterations, there are Then the algorithm is considered to have converged. represents the value of the i-th meta-parameter at the k-th iteration, δ i is the threshold value when it is the i-th meta-parameter. It can be understood that in the process of solving the meta-problem, a series of operations using the Adagrad gradient descent algorithm, including initialization, gradient calculation, updating the relevant matrix, determining the learning rate, and updating the meta-parameters, are all carried out around the objective function of the meta-problem. The purpose is to minimize the value of the objective function through continuous iterative optimization, so as to find the optimal meta-parameters and solve the meta-problem.

[0173] In some embodiments, in step S105, the underactuated ship motion adaptive control may be performed according to the meta-parameters. For example, the structure of the ship motion controller based on MRAC is as follows: Figure 3 As shown in the figure, in the process of ship motion control, using the updated target state feedback controller F and target feedforward controller K parameters, the MRAC (model reference adaptive control) controller is based on the actual state X of the current ship. s, reference model control output U m And the current disturbance situation, calculate the reference model control input U s At the same time, in each control cycle, the MRAC controller adjusts K and F in real time according to the parameter adaptation law based on the system state error. The adjustment process will be constrained and optimized by the meta-parameter θ. The meta-parameter θ guides F and K to adjust in a direction that is more conducive to the ship's trajectory tracking in a complex environment based on the training results of multiple tasks under the entire meta-learning framework. For example, when a ship encounters a new type of interference, the meta-parameter θ will prompt the MRAC controller's F and K to quickly adapt to this change, making the reference model control input U s It can correct the ship's motion heading and speed in time to ensure that the ship's actual trajectory closely follows the reference trajectory, so as to achieve adaptive control of the under-actuated ship motion.

[0174] The beneficial effects of implementing the embodiments of the present invention include: the embodiments of the present invention first obtain ship trajectory data, construct a set of ship trajectory models based on the ship trajectory data, then calculate the adaptive law based on the controlled system control output vector and the reference model control output vector, and then calculate meta-parameters through meta-learning based on the ship trajectory model set and the adaptive law, and finally perform adaptive control of under-actuated ship motion based on the meta-parameters, so that ship motion control can be achieved through parameter calculation through meta-learning, thereby improving accuracy and stability.

[0175] In some embodiments, the architecture for performing ship motion control is as follows Figure 4 As shown, data can be collected first, and then data processing, model fitting and training and model integration are performed through the system dynamics feature learning module, and the reference model, controller and adaptive mechanism are constructed through the adaptive controller module. Task data training, meta-problem construction and solution, meta-learning and parameter update are performed through the double-layer meta-learning optimization module until the parameters converge, and finally the controlled ship is controlled through the control module. In this embodiment, the system dynamics feature learning module is used to analyze and integrate the ship trajectory data, and the model is trained and optimized under the action of the double-layer meta-learning optimization module, so the accuracy of ship motion control is improved. The adaptive controller module of this embodiment is designed based on ship dynamics, and the model parameters are reasonable values ​​obtained by combining the system dynamics feature learning module and the double-layer meta-learning optimization module, which is more interpretable. In this embodiment, the meta-learning method is used, and only a small number of samples and iterations are needed to train a model with generalization ability. In the face of different tasks, the parameters can be quickly adjusted to adapt to new requirements, thereby improving the generalization ability and learning efficiency of the model. In this embodiment, the model parameters can be quickly adjusted to adapt to different control tasks and environmental changes through the collaboration of various modules, and the environmental adaptability is better.

[0176] like Figure 5As shown, the embodiment of the present invention further provides an underactuated ship motion adaptive control device based on meta-learning, comprising:

[0177] The data acquisition module 801 is used to obtain the ship track data, which includes the ship status and initial control input;

[0178] The system dynamics feature learning module 802 is used to construct a set of ship trajectory models based on the ship trajectory data;

[0179] The adaptive controller module 803 is used to calculate the adaptive law according to the controlled system control output vector and the reference model control output vector;

[0180] A two-layer meta-learning optimization module 804 is used to calculate meta-parameters through meta-learning according to the ship trajectory model set and the adaptive law;

[0181] The control module 805 is used to perform adaptive control of the motion of the underactuated ship according to the meta-parameters.

[0182] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0183] like Figure 6 As shown, an embodiment of the present invention further provides a computer device, including:

[0184] at least one processor 901;

[0185] At least one memory 902, used to store at least one program;

[0186] When at least one program is executed by at least one processor, the at least one processor implements Figure 1 The method shown.

[0187] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0188] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. An adaptive control method for underactuated ship motion based on meta-learning, characterized in that: The following steps are involved: Acquiring ship trajectory data, wherein the ship trajectory data includes a ship state and an initial control input; Constructing a set of ship trajectory models according to the ship trajectory data; Calculating an adaptive law according to a controlled system control output vector and a reference model control output vector; Calculating meta-parameters through meta-learning according to the ship trajectory model set and the adaptive law; The motion adaptive control of the underactuated ship is performed according to the meta-parameters.

2. The method according to claim 1, characterized in that The step of constructing a set of ship trajectory models according to the ship trajectory data comprises: According to the ship state and the initial control input, a preset neural network is used to perform dynamic characteristic fitting to obtain an initial ship trajectory model; Calculating a predicted value of a ship state according to the ship state, the initial control input, the model parameters and the dynamic relationship function; Constructing an objective function according to the predicted value of the ship state, the regularization coefficient, the meta-parameters and the number of sampling moments; According to the objective function, the initial ship trajectory model is trained by using stochastic gradient descent to obtain a target ship trajectory model; A plurality of the target ship trajectory models are combined to obtain the ship trajectory model set.

3. The method according to claim 2, characterized in that The construction process of the preset neural network includes: Constructing an input layer, wherein the input layer is used to receive the current state of the ship and the initial control input; After the input layer, a hidden layer is constructed, wherein the hidden layer is used to calculate the next state of the ship according to the current state and the initial control input using a preset nonlinear activation function, and the number of the hidden layers is 2; After the hidden layer, the output layer is constructed.

4. The method according to claim 1, characterized in that: The step of calculating the adaptive law according to the controlled system control output vector and the reference model control output vector comprises: Calculating a state vector of a controlled system according to the controlled system control output vector; Calculating a control law according to the controlled system state vector, a reference model control input, an initial feedforward controller and an initial state feedback controller; Using the control law as a control input of the controlled system; Constructing a state space equation group according to the controlled system state vector, the controlled system control input, the controlled system state matrix and the controlled system input matrix; Calculate a reference model state vector according to the reference model control output vector; constructing a reference model according to the reference model state vector, the reference model control input, the reference model state matrix and the reference model input matrix; Calculating the error between the reference output and the actual output according to the state vector of the controlled system and the state vector of the reference model; Calculating a system state error based on an error between the reference output and the actual output, the state space equations and the reference model; updating the system state error according to the system state convergence equation; The adaptive law is calculated according to the system state error.

5. The method according to claim 4, characterized in that The updating of the system state error according to the system state convergence equation comprises: Constructing a target feedforward controller according to the initial feedforward controller, the system state convergence equation and the error between the reference output and the actual output; Constructing a target state feedback controller according to the initial state feedback controller, the system state convergence equation and the error between the reference output and the actual output; According to the target feedforward controller and the target state feedback controller, a matching equation group is constructed so that the reference model and the controlled system can be fully matched; Calculating a feedforward controller difference according to the initial feedforward controller and the target feedforward controller; Calculating a state feedback controller difference according to the initial state feedback controller and the target state feedback controller; The system state error is updated according to the matching equation group, the feedforward controller difference and the state feedback controller difference.

6. The method according to claim 5, characterized in that The calculating the adaptive law according to the system state error comprises: Constructing a Lyapunov function according to the reference model state matrix and the positive definite matrix; Solving the Lyapunov function to obtain a Lyapunov function value; Constructing a real function according to the error between the reference output and the actual output, the Lyapunov function value, the feedforward controller difference and the state feedback controller difference; Derivative the real function to obtain a derivative equation of the real function; Performing an overall negative definite transformation on the real function derivative equation to obtain a negative definite equation system; The adaptive law is constructed according to the negative definite equation group and the gain matrix.

7. The method according to claim 1, characterized in that The calculating of meta-parameters by meta-learning according to the ship trajectory model set and the adaptive law comprises: According to the set of ship trajectory models, a plurality of ship tasks are set; Setting an interference signal corresponding to each of the ship tasks; Constructing the meta-parameters according to the adaptive law, the parameters related to the pi control input function and the parameters related to the density control input function; Calculate the target ship state information according to the initial ship state information and the adaptive input function; Calculating a target adaptive mechanism-related quantity according to the meta-parameter, the initial adaptive mechanism-related quantity, the density control input function and the interference signal; Calculating control input information according to the meta-parameters, the reference trajectory and the pi control input function; Calculating a key information set according to the target ship state information, the target adaptive mechanism related quantity and the control input information; Calculating an average tracking error based on the reference trajectory and the interference signal; Constructing a loss function according to the key information set, the average tracking error and the control gain adjustment parameter; Constructing a meta-problem according to the meta-parameters, the actual state of the ship, the reference trajectory, the control gain adjustment parameter, the initial control input information, the regularization parameter, the number of reference trajectories and the number of interference signals; According to the loss function, the meta-problem is solved by using adaptive gradient descent to update the meta-parameters.

8. The method according to claim 7, characterized in that Solving the meta-problem by using adaptive gradient descent according to the loss function and updating the meta-parameters includes: Initialize the diagonal matrix and the values ​​of the element parameters; According to the ship mission, calculating the gradient of the objective function with respect to the meta-parameters; updating the diagonal matrix according to the gradient of the objective function with respect to the element parameters; Calculate the adaptive learning rate according to the auxiliary positive number, the global learning rate and the updated diagonal matrix; The meta-parameters are updated according to the adaptive learning rate and the loss function until a change in the meta-parameters is less than a preset change threshold.

9. An adaptive control device for underactuated ship motion based on meta-learning, characterized in that: include: A data acquisition module, used to obtain ship trajectory data, wherein the ship trajectory data includes ship status and initial control input; A system dynamics feature learning module, used for constructing a set of ship trajectory models based on the ship trajectory data; An adaptive controller module, used for calculating an adaptive law according to a controlled system control output vector and a reference model control output vector; A two-layer meta-learning optimization module, used for calculating meta-parameters through meta-learning according to the set of ship trajectory models and the adaptive law; A control module is used for performing adaptive control of the motion of the underactuated ship according to the meta-parameters.

10. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Under-actuation unmanned light boat track tracking control method of ICA-CMAC neural network based on RBF identification

    CN107255923A

  • Training method and device for financial risk identification model, computer equipment and medium

    CN111724083A

  • Unmanned aerial vehicle tracking instruction online generation method under fault condition based on meta-learning

    CN116466743A

  • Unmanned ship trajectory tracking controller structure based on meta-learning and control method

    CN117539239A

  • Queue vehicle multi-constraint anti-interference trajectory tracking control method and system

    CN117991776A