Adaptive learning and training method and system for dynamic deep network model of aircraft

By using neural network modeling and online compensation methods, and training a neurodynamic model and an online compensator with time-series state data, the problem of insufficient accuracy in nonlinear modeling of traditional mechanistic models is solved, and more efficient dynamic model accuracy and control effect are achieved.

CN121436041APending Publication Date: 2026-01-30TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511389728.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Traditional mechanistic modeling methods lack accuracy when modeling nonlinear components, leading to control strategy failure and making it difficult to effectively improve the accuracy of dynamic models.

Method used

A neural network modeling and online compensation method is adopted. By acquiring the temporal state data of the controlled object, a neurodynamic model and an online compensator are trained to construct a model predictive controller. The temporal information is processed by a temporal encoder, and the parameters of the neurodynamic model are updated through a meta-learning method.

Benefits of technology

It improves the accuracy of the dynamic model and the control effect, and can better adapt to time-varying control objects, thus achieving more precise control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121436041A_ABST
    Figure CN121436041A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive learning and training method for an aircraft dynamic deep network model, and the method comprises the following steps: firstly, obtaining the time sequence state data of a group of controlled objects; secondly, on the basis of the time sequence state data of the group of controlled objects, a neurodynamic model and an online compensator for the neurodynamic model are trained; and finally, on the basis of the neurodynamic model and the online compensator, taking the neurodynamic model and the online compensator as a prediction model in model prediction control, thereby constructing a model prediction controller. According to the data-driven time-varying system method, more accurate neurodynamics is compensated online through neurodynamics representation and neurodynamics compensation, and then a control instruction is given through a model prediction control method. When historical time sequence data information is obtained, the current dynamical model of the time-varying system can be accurately calculated by encoding the state by using a time sequence encoder and compensating the neural dynamical model by using meta-learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic control, specifically to an adaptive learning and training method and system for a dynamic deep network model of an aircraft, belonging to the interdisciplinary field of high-end equipment manufacturing and artificial intelligence. Background Technology

[0002] Model-based control methods typically consist of a mechanistic model, nonlinear component modeling, and error disturbances. They are commonly used in scenarios where, for a given controlled object, a dynamic model is established through mechanistic modeling. Then, based on this dynamic model, a control strategy is derived recursively using a model-based control method. This control method has a wide range of applications, such as drone control and unmanned vehicle control.

[0003] Traditional mechanistic modeling methods analyze the mechanisms of each environment within the controlled object, such as state-space equations and transfer functions. However, many nonlinear components cannot be modeled mechanistically. If the model representation is inaccurate, the control strategy may fail. With the development of artificial intelligence, data-driven methods have become one of the important approaches for modeling dynamics. Using data-driven methods to learn a dynamic model can effectively improve the model's accuracy and the tracking performance of the control.

[0004] Therefore, how to combine data-driven methods to characterize dynamic model design has become a challenging problem of concern to those skilled in the art. Summary of the Invention

[0005] To address the aforementioned challenges, this invention proposes a method for neural network modeling and compensation of time-varying dynamic models, which at least partially improves the problems mentioned above.

[0006] This invention provides an adaptive learning and training method for a dynamic deep network model of an aircraft, characterized by the following steps: Step 1, acquiring time-series state data of a family of controlled objects; Step 2, training a neurodynamic model and an online compensator for the neurodynamic model based on the time-series state data of the family of controlled objects; Step 3, using the neurodynamic model and the online compensator as the predictive model in model predictive control, thereby constructing a model predictive controller. In Step 1, the time-series state data consists of state-control trajectories collected at multiple moments from a set of dynamic models constructed by introducing parameter deviations based on the nominal dynamic model of the controlled objects. In Step 2, the neurodynamic model uses short-series historical data and the current control quantity as input to infer the system state at the next moment; the online compensator uses short-series data as input to output partial parameters used to update the neurodynamic model. In Step 3, the state of the model predictive control is planned using the encoded state in neurodynamics, and the online compensator performs online compensation on the neurodynamic model at a predetermined frequency.

[0007] The adaptive learning and training method for the dynamic deep network model of an aircraft provided by this invention may also have the following features: The neurodynamic model, as the predictive model in model predictive control, is modeled and trained using neurodynamic methods, and a time-series encoder is added to process time-series information. The online compensator uses time-series state data to perform online compensation of the dynamic model, and then performs control solving based on the above model framework to determine the control commands for the time-varying controlled object. The specific process includes an offline training stage and an online control and compensation stage: The offline training stage includes: Step A, establishing a family of controlled object dynamic models based on the nominal dynamic model of the controlled object. The family of controlled objects includes controlled objects with similar basic dynamic models but with some deviations in their internal dynamic model parameters; Step B, collecting trajectory data during the motion process based on the family of controlled object dynamic models. The trajectory data includes state data and control data at each moment from the start to the end of the motion; Step C, dividing the collected trajectory data into time-series data groups. The time-series data groups include data from a certain moment k forward N... h After time k, N h Step D involves training a neurodynamic network with a time-series encoder and an online dynamic network compensator based on the time-series data set. The online control and compensation stage includes: Step E, using the offline-trained neurodynamic model as the predictive model for model predictive control based on the current time-series state data and the control tracking trajectory, performing rolling optimization to solve for the current optimal control quantity, and recording the current state and control information to update the time-series state data sequence; Step F, every N interval... fEach control step inputs the accumulated time-series state data to the online compensator, which then outputs updated neurodynamic model parameters. Based on these updated parameters, the predictive model in model predictive control is updated.

[0008] The adaptive learning and training method for the dynamic deep network model of aircraft provided by this invention may also have the following feature: In step B, for a specific controlled object i in a family of controlled objects, the collected state-action trajectory data at T time points are... in It is an m-dimensional state variable. It is a control variable with dimension n.

[0009] The adaptive learning and training method for the aircraft dynamic deep network model provided in this invention may also have the following feature: the neurodynamic model is a linear model with a temporal encoder, and its dynamic equation is expressed as:

[0010]

[0011] In the formula, x t u represents the state quantity at time t; t This represents the control quantity at time t. This represents the time from the current time t, and the time N forward. h Historical trajectory time-series data at each moment. Based on the historical trajectory time-series data, using the formula... Infer the hidden state r at the current time. t In the formula h ψ It is a network architecture based on LSTM design, r t It is to use historical time series data The last set of data in the hidden state sequence output by the subsequent LSTM network. Temporal hidden state r t and the current state x t Together as KOOPMAN timing encoder g ψ The input is x' t ={x t ,r t In the formula, the timing encoder g ψ Composed of a multilayer perceptron (MLP), based on x' t Together with the timing encoder, we obtain the state variable z' in the current Koopman state space. t , where z' t ∈R p Then, based on the Koopman state space, the dynamic model utilizes the linear model A. θ z t +B θ ut This indicates that the hidden state z t ={z' t x t} is formed by concatenating the state after the timing encoder and the current state, z t ∈R p+n A θ ∈R p+n×p+n The system matrix B represents the current linear model. θ ∈R p+n×m The control matrix represents the current linear model, while the decoder x t+1 =C z t+1 Used to extract a state from a hidden state z in the Koopman state space. t+1 Decoded into state space x t+1 , C∈R n×p+n The characterization decoding matrix is ​​used, while the neurodynamic model uses f. θ (x' t ,u t ) representation.

[0012] The adaptive learning and training method for the aircraft dynamic deep network model provided in this invention may also have the following characteristics: The online dynamic network compensator utilizes meta-learning representations and implements an online compensation neural dynamics model based on time-series data. The meta-learning mapping method is: μ Φ =τ i →θ i , In the formula, τ i This represents the historical time-series state data collected under the current time-varying controlled object, and the mapped parameters. This represents the system matrix and control matrix in the neurodynamic model. Meta-learning mapping is achieved through one of the following methods: (a) gradient-based updates:

[0013]

[0014] In the formula, Φ={θ init ,α} represents meta-learning μ Φ (τ i The parameters of θ init For parameters The initial value, where α represents the learning rate. This indicates the current parameter θ init Below, for θ init Calculate the loss value and its gradient, then update the parameters using several steps of gradient descent;

[0015] (b) Neural network-based updates:

[0016] θ i =μ Φ (τ i )=H Φ L i (θ init ;τ i );

[0017] In the formula, Φ={θ init ,φ} represents meta-learning μ Φ (τ i The parameters of θ init For parameters The initial values ​​are given by φ, which represents the network parameters. The meta-learning method takes the data and initial parameters as input and uses the neural network to output new neurodynamic parameters.

[0018] The adaptive learning and training method for the aircraft dynamic deep network model provided by this invention may also have the following feature: In step D, the loss function of the neural dynamics model is established using a multi-step iterative prediction error method, and its expression is:

[0019]

[0020] In the formula, x t+1 Trajectory data from a group of controlled objects {τ i},and The predicted trajectory is derived through iterative calculations using neurodynamic equations; similarly, The Koopman space state variables are derived through iteration using a neurodynamic model, while It is through x t+1 The state variable N obtained through the timing encoder h It represents the number of forward iterations, consistent with the length of the planning process for model predictive control.

[0021] The adaptive learning and training method for the dynamic deep network model of an aircraft provided by this invention may also have the following feature: wherein, in step E, the optimization problem of model predictive control is expressed as:

[0022]

[0023] In the formula, u i z i The target variable for optimization is the hidden state z. i With the planned control quantity u i N p Let z be the length of the sequence programming problem, and in the constraints, z t+1 =A θ z i +Bθ u i For neurodynamic models, Represents the hidden state z i The linear constraint matrix, Indicates the control quantity u t The linear constraint matrix, b i The constraints represent the state and control parameters, and z0 represents the initial hidden state of the program, which is generated from the current initial state {x0, r0} through the encoder z0 = g. ψ ({x0,r0}) is encoded to obtain the objective function. for:

[0024]

[0025] In the formula, Q i ∈R (p+n)×(p+n) This indicates that for the latent variable z i From the current time N p -1 is the cost matrix at each time step, R i ∈R m×m This indicates that for the control variable u t From the current time N p -1 is the cost matrix at each time step. Indicates the last moment of the planning process, N. p Latent variable z i The cost matrix for the control variable u t The planning length is greater than the latent variable z i The planned length is short by 1. This indicates that for the latent variable z i Track the trajectory state at time i, that is, the desired state at time i.

[0026] The adaptive learning and training method for the dynamic deep network model of aircraft provided by this invention may also have the following feature: wherein the latent variable z i At the current time, k is less than N. h At that time, the timing data input to the timing encoder All are filled by the first time step {x0}; at one time step All are filled with {x0}, and each time control time passes, a value is extracted. The last data in the sequence is selected, and the current data is added to the first position of the time series data, while the queue length remains unchanged. The tracking trajectory status is... Tracking states using trajectory in state space It is formed by merging a p-dimensional vector with a content of 0. Cost matrix Q i ∈R (p +n)×(p+n) In the middle, qi ∈R n×n It is Q i The first n×n matrix is ​​for the trajectory tracking state. The cost matrix, and Q i The rest of the text is filled with 0.

[0027] The adaptive learning and training method for the dynamic deep network model of aircraft provided by this invention may also have the following feature: In step F, the online dynamic model compensation process is as follows: the online compensator, based on a meta-learning method, operates at a preset frequency and updates the parameters of the neural dynamic model using historical time-series data as input. The compensation method is based on real-time time-series data control sequence d. i The time series data is from the current time k to the past Nx+N. p Data at any given time, i.e. The compensator is μ Φ =τ i →θ i Based on online time series data τ i Mapping the current compensation variable, using Time series data, after an error iterative prediction process, are used to obtain... Then, the neurodynamic parameters are updated using a meta-learning mapping method. If the current time k is less than N h +N p Then the neurodynamic model uses initial parameters; if k is greater than N h +N p And k is the update frequency N hz If the value is an integer multiple of the value, then the above update process is executed.

[0028] This invention provides an adaptive learning and training system for a dynamic deep network model of an aircraft, characterized by the following features: a data preparation and acquisition module for acquiring time-series state data of a family of controlled objects; a model training and construction module for training a neurodynamic model and an online compensator for the neurodynamic model based on the time-series state data of the family of controlled objects; and a control and online compensation module for constructing a model predictive controller based on the neurodynamic model and the online compensator, using them as the predictive model in model predictive control. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of a data-driven time-varying system control method provided in an embodiment of the present invention;

[0030] Figure 2 This is a design diagram of the neurodynamic model and the neurodynamic compensation method in an embodiment of the present invention;

[0031] Figure 3 This is a diagram illustrating the tracking effect of the control trajectory for each state in an embodiment of the present invention;

[0032] Figure 4 This describes the overall effect of the control trajectory in the embodiments of the present invention. Detailed Implementation

[0033] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0034] This embodiment provides an adaptive learning and training method for a dynamic deep network model of an aircraft. Specifically, it provides a data-driven time-varying system control system, including a data-driven neurodynamic model representation, an online compensation scheme based on meta-learning, and a model control scheme. Both the neurodynamic model representation and the meta-learning compensator are input in a time-series manner. During the control phase, model predictive control is used for control, while the neurodynamic model is compensated online. Using this method, a time-varying control system can be controlled.

[0035] Figure 1 This is a schematic diagram of a data-driven time-varying system control method provided in an embodiment of the present invention.

[0036] like Figure 1 As shown, the adaptive learning and training method for the aircraft dynamic deep network model provided in this embodiment includes the following steps:

[0037] Step S101: Obtain the time-series state data of a family of controlled objects. The time-series state data consists of state-control trajectories collected at multiple moments from a set of dynamic models constructed based on parameter deviations introduced from the nominal dynamic model of the controlled objects.

[0038] In this embodiment, it is applied to the tracking control of a planar quadcopter. Time-varying deviations are included in this simulation to further test the robustness of the framework. The dynamic modeling of the planar quadcopter is as follows:

[0039]

[0040] In the formula, the state variables are y, z is the aircraft position, and θ is the pitch angle of the quadcopter. Its state variable x t For the dynamic model, where parameter g is the acceleration due to gravity, m is the mass of the spacecraft, and I... xxLet l be the moment of inertia of the aircraft, l be the length of the aircraft arm, and ΔT be the sampling time.

[0041] Here, T1 and T2 are the thrust generated by the motor, which are the action variables u. t The dynamic disturbance distribution ρ will be in I xx ,m,l. Where ρ=[-20%,20%] is the deviation added to the nominal dynamic model.

[0042] The system utilizes various stochastic dynamic disturbances ρ for control, where the data collected under each deviation is used as a flight trajectory data τ. i The data length is T time steps. Then, n data points are collected to form a family of control trajectory data.

[0043] Step S102: Based on the time-series state data of a family of controlled objects, a neurodynamic model and an online compensator for the neurodynamic model are trained. The neurodynamic model takes short-series historical data and the current control input as input to infer the system state at the next moment; the online compensator takes short-series data as input and outputs some parameters used to update the neurodynamic model. The specific implementation process is as follows:

[0044] S102-1, Segmentation of the temporal state data of a family of controlled objects.

[0045] First, based on the nominal dynamic model of the controlled object, a family of controlled objects dynamic models are established. The family of controlled objects includes controlled objects with similar basic dynamic models, but with some deviations in their internal dynamic model parameters.

[0046] Secondly, based on the dynamic model of a family of controlled objects, trajectory data during their motion process is collected. The trajectory data includes state data and control data at each moment from the start to the end of the motion.

[0047] Finally, based on the trajectory data of a group of controlled objects, the data is divided into time-series data groups. The time-series data sequence includes the first N data points from a certain time k. h After time k, N h Status data and control data.

[0048] S102-2, Construct the neurodynamic equations with a time-series encoder.

[0049] The established neural dynamics equations are linear dynamics equations with a time-sequence encoder;

[0050] The linear neurodynamic equations are described as follows:

[0051]

[0052] Where x t Let u represent the state quantity at time t. t This represents the control quantity at time t. This represents the time from the current time t, and the time N forward. h Historical trajectory time-series data at each moment. Based on the above historical trajectory time-series data, using... Infer the hidden state r at the current time. t , where h ψ It is a network architecture based on LSTM design, r t It is to use historical time series data The last set of data in the hidden state sequence output by the LSTM network. And the time-series hidden state r... t and the current state x t As a KOOPMAN timing encoder g ψ Input x' t ={x t ,r t}, where the timing encoder g ψ It is composed of a multilayer perceptron (MLP). Based on x' t Together with the timing encoder, we obtain the state variable z' in the current Koopman state space. t , where z' t ∈R p Then, based on the Koopman state space, the dynamic model utilizes the linear model A. θ z t +B θ u t This indicates that the hidden state z t ={z' t x t} is formed by concatenating the state after the timing encoder and the current state, z t ∈R p+n Among them, A θ ∈R p+n×p+n The system matrix B represents the current linear model. θ ∈R p+n×m The control matrix represents the current linear model. The decoder x... t+1 =Cz t+1 Used to extract the state from the hidden state z in the Koopman state space. t+1 Decoded into state space x t+1 , where C∈R n×p+n The characterization decoding matrix. The neurodynamic model uses f... θ (x' t ,u t ) representation.

[0053] S102-3, Establishment of a neurodynamic model compensation method. The specific process is as follows:

[0054] The online dynamics network compensator utilizes meta-learning representations and implements an online compensatory neural dynamics model based on time-series data. The mapping method of meta-learning is as follows:

[0055] μ Φ =τ i →θ i ,

[0056] Where, τ i This represents the historical time-series state data collected under the current time-varying controlled object, and the mapped parameters. This represents the system matrix and control matrix in a neurodynamic model.

[0057] In this embodiment, two meta-learning mapping methods are designed. The first mapping method is based on gradients:

[0058]

[0059] Where, Φ={θ init ,α} represents meta-learning μ Φ (τ i The parameters of ) are θ. init For parameters The initial value of , where α represents the learning rate. This indicates the current parameter θ init Below, for θ init Calculate the loss value and its gradient, then update the parameters using several steps of gradient descent.

[0060] The second meta-learning method is specifically expressed as follows: This method uses a network architecture to represent the gradient descent process:

[0061] θ i =μ Φ (τ i )=H Φ L i (θ init ;τ i )

[0062] Where, Φ={θ init ,φ} represents meta-learning μ Φ (τ i The parameters of ) are θ. init For parameters The initial values ​​are denoted by φ, which represents the network parameters. The meta-learning method takes the data and initial parameters as input and uses the neural network to output new neurodynamic parameters.

[0063] The loss function of the neurodynamic model is established using a multi-step iterative prediction error method, and its expression is as follows:

[0064]

[0065] Where, x t+1 Trajectory data from a group of controlled objects {τ i},and The predicted trajectory is derived through iterative calculations using neurodynamic equations; similarly, The Koopman space state variables are derived through iteration using a neurodynamic model, while It is through x t+1 The state variables are obtained through a timing encoder. Where N... h It represents the number of forward iterations, consistent with the length of the planning process for model predictive control.

[0066] Figure 2 This is a design diagram of the neurodynamic model and the neurodynamic compensation method in an embodiment of the present invention.

[0067] Step S103: Based on the neurodynamic model and the online compensator, it is used as the predictive model in model predictive control, thereby constructing a model predictive controller. The state of the model predictive control is planned using the encoded state in neurodynamics, and the neurodynamic model is compensated online at a predetermined frequency using the online compensator. The specific process is as follows:

[0068] S103-1, Design Model Predictive Control Method.

[0069] The optimization problem of model predictive control is formulated as follows:

[0070]

[0071] stz t+1 =A θ z i +B θ u i i = 0, ..., N p -1;

[0072] E i z i +F i u i ≤b i ,

[0073]

[0074] z0 = g(x0),

[0075] In the formula, ui z i The target variable for optimization is the hidden state z of the aircraft. i The control quantity u of the planned aircraft i N p Let z be the length of the sequence programming problem, and in the constraints, z t+1 =A θ z i +B θ u i For neurodynamic models, Represents the hidden state z i The linear constraint matrix, Indicates the control quantity u t The linear constraint matrix, b i The constraints represent the state and control parameters, and z0 represents the initial hidden state of the program, which is generated from the current initial state {x0, r0} through the encoder z0 = g. ψ ({x0,r0}) is encoded to obtain the objective function. for:

[0076]

[0077] In the formula, Q i ∈R (p+n)×(p+n) This indicates that for the latent variable z i From the current time N p -1 is the cost matrix at each time step, R i ∈R m×m This indicates that for the control variable u t From the current time N p -1 is the cost matrix at each time step. Indicates the last moment of the planning process, N. p Latent variable z i The cost matrix for the control variable u t The planning length is greater than the latent variable z i The planned length is short by 1. This indicates that for the latent variable z i Track the trajectory state at time i, that is, the desired state at time i.

[0078] Latent variable z i At the current time, k is less than N. h At that time, the timing data input to the timing encoder All are filled by the first time step {x0}; at one time step All are filled with {x0}, and each time control time passes, a value is extracted. The last data in the sequence is selected, and the current data is added to the first position of the time series data, while the queue length remains unchanged. The tracking trajectory status is... Tracking states using trajectory in state space It is formed by merging a p-dimensional vector with a content of 0.

[0079] Cost matrix Q i ∈R (p+n)×(p+n) In the middle, q i ∈R n×n It is Q i The first n×n matrix is ​​for the trajectory tracking state. The cost matrix, and Q i The rest of the text is filled with 0.

[0080] S103-2, every interval N f Each control step inputs the accumulated time-series state data to the online compensator, which then outputs updated neurodynamic model parameters. Based on these updated parameters, the predictive model in model predictive control is updated.

[0081] In this embodiment, the online dynamic model compensation process is as follows:

[0082] The online compensator is based on a meta-learning method, operates at a preset frequency, and updates the parameters of the neurodynamic model using historical time-series data as input.

[0083] The compensation method is based on real-time time-series data control sequence d. i Time series data is from the current time k to the past N. h +N p Data at any given time, i.e. The compensator is μ Φ =τ i →θ i Based on online time series data τ i Mapping the current compensation variable, using Time series data, after an error iterative prediction process, are used to obtain... Then, the neurodynamic parameters are updated using a meta-learning mapping method. If the current time k is less than N h +N p Then the neurodynamic model uses initial parameters; if k is greater than N h +N p And k is the update frequency N hz If the value is an integer multiple of the value, then the above update process is executed.

[0084] Figure 3 This is a diagram showing the tracking effect of the control trajectory for each state in an embodiment of the present invention.

[0085] Figure 4This is a diagram illustrating the overall effect of the control trajectory in an embodiment of the present invention.

[0086] Specifically, the data-driven time-varying system method in this embodiment uses neurodynamic representation and compensation to obtain more accurate neurodynamics online, and then provides control commands through model predictive control. When historical time-series data is obtained, the current dynamic model of the time-varying system can be accurately calculated by using a time-series encoder to encode the state and by using meta-learning to compensate the neurodynamic model.

[0087] The above embodiments are preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention.

Claims

1. An adaptive learning and training method of an aircraft dynamic deep network model, characterized in that, The method comprises the following steps: Step 101, obtaining time sequence state data of a group of controlled objects; Step 102, training a neural dynamics model and an online compensator for the neural dynamics model based on the time sequence state data of the group of controlled objects; Step 103, taking the neural dynamics model and the online compensator as a prediction model in model predictive control, thereby constructing a model predictive controller, In step 101, the time sequence state data is state-control trajectories collected at multiple time points from a group of dynamic models constructed by introducing parameter deviations based on a nominal dynamic model of a controlled object, In step 102, the neural dynamics model takes short time sequence historical data and current control quantity as input to infer the system state at the next time point; and the online compensator takes short time sequence data as input to output part of parameters of the neural dynamics model for updating, In step 103, the state of model predictive control is planned by using the state encoding in neural dynamics, and the online compensator is used to perform online compensation on the neural dynamics model at a predetermined frequency.

2. The method of claim 1, wherein the aircraft dynamic depth network model is trained using a plurality of training data sets, each training data set comprising a plurality of flight parameters and a plurality of flight conditions. The method is characterized in that: The neural dynamics model is taken as a prediction model in model predictive control, is modeled and trained by using a neural dynamics method, and is provided with a time sequence encoder for processing time sequence information, The online compensator performs online compensation on the dynamic model by using time sequence state data, and then performs control solving based on the above model architecture to determine the control instruction of a time-varying controlled object, and the specific process comprises an offline training phase and an online control and compensation phase: The offline training phase comprises: Step A, establishing a group of dynamic models of controlled objects based on a nominal dynamic model of the controlled object, wherein the group of controlled objects comprises controlled objects similar to a basic dynamic model but with deviations in internal dynamic model parameters; Step B, collecting trajectory data in the motion process of the group of dynamic models based on the group of dynamic models of the controlled objects, wherein the trajectory data comprises state quantity data and control quantity data at each time point from the start to the end of the motion; Step C: the collected trajectory data is divided into time series data sets, which include state quantity data and control quantity data from time k to time k+N h forward and time k-N h backward. Step D, training a neural dynamics network with a time sequence encoder and an online dynamic network compensator based on the time sequence data set, The online control and compensation phase comprises: Step E, taking the neural dynamics model trained offline as a prediction model of model predictive control to perform rolling optimization based on the time sequence state data and the trajectory of control tracking at the current time, solving the current optimal control quantity, and recording the state information and control information at the current time to update the time sequence state data sequence; Step F, every N f control steps, the accumulated timing state data is input to the online compensator, which outputs updated neuromuscular model parameters, and based on the updated parameters, the prediction model in model predictive control is updated.

3. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 2 is characterized in that: wherein, In step B, for one of the group of controlled objects i, the collected state-action trajectory data at T time instants is where is an m-dimensional state quantity, is an n-dimensional control quantity.

4. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 2 is characterized in that: wherein, The neural dynamics model is a linear model with a time sequence encoder, and its dynamic equation is represented as: In the formula, x t represents the state quantity at the tth moment; u t represents the control quantity at the tth moment; k = t - N h , t - N h + 1, ……, t, represents the historical trajectory time series data from the current moment t, and N h previous moments. Based on the historical trajectory time series data, a formula is used Infer the hidden state r at the current moment t , wherein h ψ is a network architecture designed based on LSTM, r t is the last set of data of the hidden state sequence output by the LSTM network after the historical time series data ​ The timing hidden state r t And the current state x t Together as the input of the KOOPMAN timing encoder g ψ , get x' t ={x t ,r t}, in the formula, the timing encoder g ψ It is composed of a multi-layer perception mechanism, based on x' t And the timing encoder, get the state variable z' t Under the current KOOPMAN state space, wherein z' t ∈R p , Then, based on the KOOPMAN state space, the dynamic model utilizes a linear model A θ z t +B θ u t represents, wherein the hidden state z t ={z' t , x t} is spliced by the state after the time series encoder and the current state, z t ∈R p+n , A θ ∈R p+n×p+n characterizes the system matrix under the current linear model, B θ ∈R p+n×m characterizes the control matrix under the current linear model, and the decoder x t+1 =Cz t+1 is used to decode the state from the hidden state z t+1 under the KOOPMAN state space to x t+1 under the state space, C∈R n×p+n characterizes the decoding matrix, and the neural dynamic model is represented by f θ (x' t , u t ).

5. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 2 or 4 is characterized in that: wherein The online dynamics network compensator utilizes meta-learning representation, and implements an online compensating neural dynamics model based on time series data, and the meta-learning mapping mode is: In the formula, τ i represent the historical time sequence state data collected under the current time-varying controlled object, the mapped parameters represent the system matrix and the control matrix in the neural dynamics model, The meta-learning mapping mode is implemented by one of the following modes: (a) Gradient-based update: In the formula, Φ = {θ init , α} represents the parameters of meta-learning μ Φ (τ i ), θ init is the initial value of the parameter θ i = θ , α represents a learning rate, represents that the loss value is calculated for θ init under the current parameter θ init , and the gradient is obtained, and then the parameter is updated by using several steps of gradient descent; (b) Neural network-based update: θ i = μ Φ (τ i ) = H Φ L i (θ init ; τ i ) ; In the formula, Φ = {θ init , φ} represents parameters of meta-learning μ Φ (τ i ), θ init is an initial value of the parameter , and φ is a parameter of the network. The meta-learning method takes data and an initial parameter as input, and outputs new neural dynamics parameters by using a neural network.

6. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 2, characterized in that: In step D, the loss function of the neural dynamics model is established in a multi-step iterative prediction error mode, and the expression is: In the formula, x t+1 Trajectory data from a group of controlled objects {τ i },and The predicted trajectory is derived through iterative calculations using neurodynamic equations; similarly, The Koopman space state variables are derived through iteration using a neurodynamic model, while It is through x t+1 The state variable N obtained through the timing encoder h It represents the number of forward iterations, consistent with the length of the planning process for model predictive control.

7. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 2, characterized in that: wherein, In step E, the optimization problem of the model predictive control is expressed as: In the formula, u i is the control variable, z i is the target variable of optimization, that is, the hidden state z t and the planned control variable u t , N p is the length of sequence planning, in the constraint condition, z t+1 =A θ z i +B θ u i is the neural dynamics model, represents the linear constraint matrix of the hidden state z i , represents the linear constraint matrix of the control variable u t , b i represents the constraint quantity of the state and control, z0 represents the initial hidden state of planning, which is obtained by encoding the current initial state {x0, r0} through the encoder z0=g ψ ({x0, r0}), and the objective function is: In the formula, Q i ∈R (p+n)×(p+n) This indicates that for the latent variable z i From the current time N p -1 is the cost matrix at each time step, R i ∈R m×m This indicates that for the control variable u t From the current time N p -1 is the cost matrix at each time step. Indicates the last moment of the planning process, N. p Latent variable z i The cost matrix for the control variable u t The planning length is greater than the latent variable z i The planned length is short by 1. This indicates that for the latent variable z i Track the trajectory state at time i, that is, the desired state at time i.

8. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 7, characterized in that: wherein The hidden variable z i At the current time k less than N h When its input to the timing encoder timing data All by the first time {x0} fill; in a time All by {x0} fill, every control time, the last data in Put forward, and fill the current data to the first timing data, the queue length remains unchanged, The tracking trajectory state is The tracking trajectory state is and a vector with p-dimensional content of 0, The cost matrix Q i ∈R (p+n)×(p+n) where q i ∈R n×n is the upper n x n matrix in Q i is the cost matrix for the trajectory tracking state while the rest of Q i is padded with zeros.

9. The adaptive learning and training method of the aircraft dynamic deep network model according to claim 2 or 5, characterized in that: wherein In step F, the online dynamics model compensation process is: The online compensator is based on the meta-learning method, operates at a preset frequency, and updates the parameters of the neural dynamics model with historical time series data as input, wherein the compensation method controls the sequence d based on real-time timing data i Timing data is data from the current time k to the past N h +N p time, i.e. The compensator is μ Φ = τ i → θ i , based on online timing data τ i , maps out the current compensation variable, using timing data, through an error iterative prediction process, to obtain Then the neural dynamics parameters are updated using the meta-learning mapping method If the current time instant k is less than N h +N p , the neural dynamics model uses the initial parameters; if k is greater than N h +N p and k is an integer multiple of the update frequency N hz , the above update procedure is performed.

10. An adaptive learning and training system for an aircraft dynamic depth network model, the system comprising: Comprise: a data preparation and acquisition module for acquiring time series state data of a family of controlled objects; a model training and construction module for training a neural dynamics model and an online compensator for the neural dynamics model based on the time series state data of the family of controlled objects; a control and online compensation module for using the neural dynamics model and the online compensator as a prediction model in model predictive control, thereby constructing a model predictive controller.