Full hovering hovercraft hovering pressure control method and device based on generative adversarial imitation learning and storage medium

Through the generative adversarial imitation learning method, combined with multivariable autoencoder and LSTM network, the dependence of the full-pad hoverboat hover lift system on driver experience is solved, intelligent pad lift pressure control is realized, and operating efficiency and system stability are improved.

CN120335512APending Publication Date: 2025-07-18HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419446.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The control of the fully padded hoverboat padded hoist system depends on the driver's experience, making it difficult to achieve automatic adaptation and global optimization in complex environments, increasing operational difficulty and risk.

Method used

Generative adversarial imitation learning method is adopted, and the generative adversarial model is constructed, combined with multivariate autoencoder and expert data discriminator of LSTM network, and simulation data and driving simulator data are used for progressive training to assist drivers in pad lift pressure control.

Benefits of technology

It reduces the operation difficulty of the pad lift system, improves the robustness and stability of the control system, can automatically adapt to different environments and working conditions, and reduces the dependence on driver experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335512A_ABST
    Figure CN120335512A_ABST
Patent Text Reader

Abstract

The invention discloses a full-hovering hovercraft hovering pressure control method and device based on generative adversarial imitation learning and a storage medium, and belongs to the field of hovercraft hovering control. The control method based on generative adversarial imitation learning is adopted, a progressive training mode is used, and intelligent control over hovering pressure of the full hovering hovercraft is achieved. The generative adversarial model adopted by the invention comprises an expert data discriminator combined by a multi-variation auto-encoder and an LSTM network, and an SAC network is used as a hover state control variable generator, and through adversarial training of the two models, the performance of the two network models is improved at the same time. According to the progressive training mode provided by the invention, the existing non-physical simulation data, hovercraft driving simulator data and limited expert data are fully utilized to improve the network performance, and the expert data are simulated. The hover control method solves the problem that hover control of the full hover is excessively dependent on experience of a driver and the like, and the hover control difficulty is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of hovercraft lift control, and particularly relates to a lift pressure control method, device and storage medium for a fully-lifted hovercraft based on generative adversarial imitation learning. Background Art

[0002] A fully-lifted hovercraft is an amphibious special ship that hovers over water, land, swamps and other complex environments through a lift system. Due to its good mobility and speed, it is widely used in naval equipment and civilian rescue, etc. The existence of the lift system of the hovercraft reduces the driving resistance while increasing the instability of the hovercraft. Therefore, it is of great significance to study the pressure control technology of the hovercraft lift system. Lift pressure control is to adjust the pressure control mechanism according to the state of the hovercraft and in combination with the actual task requirements, and control the air cushion pressure to keep the hovercraft stable. In previous studies on lift pressure control, it was mainly the lift fan flow-pressure control based on traditional control methods, the air release valve control combined with pressure sensors, and the passive air cushion pressure control relying on skirt design. The above methods rely on manual adjustment and traditional control methods, and the lift effect depends on the driver's operation experience. With the improvement of computer computing power and the birth of various precision sensors, the driver's operation process and environmental information can be accurately recorded. Generative adversarial imitation learning is constructed based on generative adversarial networks and can obtain the same data distribution as expert data under the condition of sufficient samples. By using this method to learn the driver's operation data, it can assist the driver's operation and reduce the difficulty of controlling the lift pressure system. Summary of the Invention

[0003] The present invention solves the problem of the control difficulty faced by the driver during the lift process of a fully-lifted hovercraft. By adopting a control method based on generative adversarial imitation learning, it realizes the intelligent control of the lift pressure of the fully-lifted hovercraft, thereby helping the driver reduce the operation difficulty and improve the operation efficiency. The present invention adopts the generative adversarial imitation learning method, and realizes the lift pressure control of the fully-lifted hovercraft by learning the existing effective data, and obtains a lift control system with better robustness and stability. The progressive training method proposed by the present invention makes full use of the existing non-physical simulation data, hovercraft driving simulator data and limited expert data to improve the network performance and achieve the imitation of expert data.

[0004] The present invention provides a lift pressure control method for a fully-lifted hovercraft based on generative adversarial imitation learning, including:

[0005] Step 1: Establish a mathematical model of the lift system of the fully-lifted hovercraft, and design a controller according to the six-degree-of-freedom model of the fully-lifted hovercraft to initially realize the lift control of the simulation environment, and obtain simulation training data;

[0006] Step 2: Construct a generative adversarial model; the structure of the generative adversarial model includes: an expert data discriminator and a lift state control variable generator; the expert data discriminator uses a combination of a variational autoencoder and an LSTM structure to discriminate between the generated data and the real data of the input samples; the lift state control variable generator uses the SAC algorithm, and the SAC network model includes: an Actor network, a Critic network, and a Critic-target target network;

[0007] Step 3: Use the simulation training data and the simulation data of the hovercraft driving simulation platform to train the constructed generative adversarial model in a progressive training manner; first, use the simulation training data to perform the first-stage training on the generative adversarial model; then, use the simulation data obtained from the hovercraft driving simulation platform to perform the second-stage training to further optimize the network parameters of the expert data discriminator; use the simulation data to input the variational autoencoder for encoding to extract features, splice the feature encodings, input them into the LSTM network, and output the discrimination results of the actions executed in the continuous state;

[0008] Step 4: The driver inputs the current working state of the hovercraft into the generative adversarial model trained in Step 4, and the lift flow rate of the hovercraft is output; use the driver operation data sample as the result to further optimize the generative adversarial model.

[0009] Furthermore, the specific steps of Step 1 include:

[0010] Step 1.1: The six-degree-of-freedom maneuvering equation of the six-degree-of-freedom motion model of the fully-lifted hovercraft:

[0011]

[0012] where I x 、I y 、I z are the moments of inertia, u, v, w are the longitudinal, lateral, and vertical velocities, p, q, r are the roll angular velocity, pitch angular velocity, and yaw angular velocity, F x 、F y 、F z are the resultant forces, M x 、M y 、M z are the torques;

[0013] According to the air chamber layout of the fully-lifted hovercraft, the apron response model, and the structural parameters of the lift fan, construct a mathematical model of the lift system of the fully-lifted hovercraft to obtain the four-air chamber force balance equation of the hovercraft:

[0014]

[0015] Among them, W is the gravity of the hovercraft; F0 is all longitudinal environmental forces acting on the hovercraft; X x0 , Y y0 , M x0 , M y0 , M z0 are the resultant force and resultant moment generated by environmental forces; X x1 , Y y1 , M x1 , M y1 , M z1 are the resultant force and resultant moment of the hovercraft's propellers, side air dampers, and air rudders;

[0016] Step 1.2: Set the step size and the initial cushion lift pressure P c and the airbag pressure P n , and perform one round of iteration at each step size to obtain the attitude information of the hovercraft;

[0017] Step 1.3: According to the attitude information of the hovercraft, use neural network PID to control the state of the hovercraft, and obtain the time series composed of [p, q, r, u, v, w, P c , P n and the cushion lift fan flow rate [Q] as training data.

[0018] Furthermore, the said Step 3 includes the following steps:

[0019] Step 3.1: Screen and preprocess the simulation training data to construct a training set;

[0020] Step 3.2: Pretrain the cushion lift state control variable generator and the expert data discriminator; first, use the simulation training data to pretrain the cushion lift state control variable generator; then use the pretrained cushion lift state control variable generator to generate fake samples and real samples, and cross-train the expert data discriminator using loss functions such as entropy.

[0021] Step 3.3: Perform adversarial training using the cushion lift state control variable generator and the expert data discriminator, and adjust the parameters of the cushion lift state control variable generator and the expert data discriminator to make the cushion lift state control variable generator generate more realistic synthetic data; during the adversarial training process, first fix the parameters of the expert data discriminator, use the cushion lift state control variable generator to generate fake samples and input them into the expert data discriminator, calculate the loss function and gradient of the cushion lift state control variable generator based on the output of the expert data discriminator to update its parameters; subsequently, fix the parameters of the cushion lift state control variable generator, mix real samples and generated samples, calculate the gradient of the expert data discriminator through cross-entropy loss and update its parameters; this process realizes the adversarial learning of the generative adversarial model through the alternating training of the cushion lift state control variable generator and the expert data discriminator;

[0022] Step 3.4: Evaluate the performance of the generative adversarial model using expert data; determine whether the generative adversarial model is optimized and meets the standards based on the flow rate [Q] of the lift fan, the average return, and the average method; if the generative adversarial model does not meet the standards, return to Step 3.3.

[0023] Further, in Step 3.2, the pre-trained lift state control variable generator is the input S of the Actor network j , and outputs A j , and then use an unsupervised method to pre-train the Actor network; first randomly initialize the parameters of the lift state control variable generator, and then use an autoencoder to train the lift state control variable generator; the input of the autoencoder is random noise, the output is the output of the Actor network, and the middle layer is the hidden layer of the lift state control variable generator;

[0024] The loss function of the Actor network of the lift state control variable generator is:

[0025]

[0026] where n is the number of samples, x i is the input sample, z i is the random noise, and G is the lift state control variable generator;

[0027] The gradient of the Actor network of the lift state control variable generator is:

[0028]

[0029] where θ G is the parameter of the lift state control variable generator, is the partial derivative of the parameter of the lift state control variable generator.

[0030] Further, the adversarial training in Step 3.3 specifically includes the following steps:

[0031] Step 3.3.1: Obtain the expert data sequence τ' and generate the lift control sequence τ;

[0032] Step 3.3.2: Train the expert data discriminator; input the expert data sequence τ' and the generated lift control sequence τ into the expert data discriminator to obtain the discrimination probabilities dτ′ i and dτ j of the expert data discriminator for the expert data sequence τ' and the generated lift control sequence τ; calculate the loss function and update the network parameters of the expert data discriminator using the gradient descent loss function;

[0033] Step 3.3.3: Use the discrimination result of the expert data discriminator on the data generated for cushioning state control as the reward, and update the parameters of the cushioning state control variable generator using the policy gradient method.

[0034] Further, in Step 3.3.1, the acquisition of the expert data sequence is as follows: Input the expert feature state S i ' into the Actor network in the cushioning state control generator to obtain the action A i ' in this state. Use the six-degree-of-freedom motion model of the fully cushioned hovercraft and the mathematical model of the cushioning system of the fully cushioned hovercraft in Step 1 to obtain the next cushioning state S i ' +1 ; Continuously substitute S i ' +1 into the cushioning state control generator to obtain the state-control sequence {S′1,A′1,S′2,A′2,...,S′ m ,A′ m}; Concatenate S′ i and A′ i to obtain the tensor τ′ i . Represent the complete sequence {τ1',τ2'...τ n '} with τ', and the sequence length is n;

[0035] The acquisition of the generated cushioning control sequence is as follows: Input the hovercraft feature state S j into the Actor network in the cushioning state control generator to obtain the action A j in this state. Use the six-degree-of-freedom motion model of the fully cushioned hovercraft and the mathematical model of the cushioning system of the fully cushioned hovercraft in Step 1 to obtain the next cushioning state S j+1 ; Continuously substitute S j+1 into the cushioning state control generator to obtain the state-control sequence {S1,A1,S2,A2...S m ,A m}; Concatenate S j and A j to obtain the tensor τ j . Represent the complete sequence {τ1,τ2...τ m} with τ, and the sequence length is m.

[0036] Further, in Step 3.3.2, the gradient descent loss function of the expert data discriminator is:

[0037]

[0038] where D(τ) is the discrimination output result of the expert data discriminator on the input trajectory sample τ; D(τ') is the discrimination output result of the expert data discriminator on the input trajectory sample τ'.

[0039] Further, in step 3.3.3, the lift state control variable generator is updated to first calculate the difference Δ between the Critic network and the target network Critic-target:

[0040] Δ = Critic_target(τ) + reward - Critic(τ).

[0041] Wherein, reward is the discrimination result D(τ) of the expert data discriminator for the lift state control generated data;

[0042] Use Δ to calculate the advantage function adv for each τ j The gradient of the lower entropy is weighted to obtain the policy gradient Update the Actor network parameter π(θ) by the policy gradient method:

[0043]

[0044] Wherein, P(τ j |π) is the probability of executing the action τ under the policy π j ; H(π(S i ; θ)) is the entropy value of the current policy,

[0045] At the end of each episode, the Critic-target target network is updated in a soft-update manner, that is, a very small proportion of new network parameters are updated each time:

[0046] θ target ← λθ target + (1 - λ)θ

[0047] Wherein, θ target is the parameter of the target network; λ is the update law designed by the soft-update method; θ is the network parameter.

[0048] The present invention also provides a computer device / equipment / system, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the steps of the above-mentioned full-lift hovercraft lift pressure control method based on generative adversarial imitation learning are implemented.

[0049] The present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the above-mentioned full-lift hovercraft lift pressure control method based on generative adversarial imitation learning are implemented.

[0050] The present invention also provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the cushion lift pressure control method for a fully lifted hovercraft based on generative adversarial imitation learning described in any one of the above.

[0051] The beneficial effects of the present invention are as follows:

[0052] (1) The present invention solves the problem that the control of the cushion lift system of a fully lifted hovercraft in the past relied on the driver's experience. The control of the cushion lift system is an important part of hovercraft driving. The present invention makes full use of the simulation environment, the hovercraft driving simulation platform and the actual operation data of experts, and trains the constructed generative adversarial imitation learning network by sampling progressive training to assist the driver in controlling the cushion lift system.

[0053] (2) A discriminator network structure for expert data combining a multi-variational autoencoder and an LSTM network is proposed. The multi-variational autoencoder encodes the state of the cushion lift system to obtain features, and the LSTM network ensures the correlation between state-action sequences. These two operations greatly improve the discriminative ability of the expert data discriminator for expert data and generated data, and improve the performance of the entire generative adversarial imitation learning network.

[0054] (3) The present invention uses the SoftActor Critic network as the generator of cushion lift state control variables. Sampling softupdate promotes the convergence of the Critic network. For the Actor network, on the one hand, the advantage function is set in combination with the context before and after the sequence, and on the other hand, the policy gradient is calculated to increase terms, improve the dispersion of the search strategy of the Actor, and increase the policy generalization ability. Description of the Drawings

[0055] Figure 1 It is a structural diagram of the air cushion chamber of a fully lifted hovercraft;

[0056] Figure 2 It is a simulation design diagram of the motion state of a fully lifted hovercraft;

[0057] Figure 3 It is a structural diagram of the generative adversarial imitation learning network for expert data of cushion lift state control of a fully lifted hovercraft;

[0058] Figure 4 It is a progressive network training flow chart of the generative adversarial model for expert data of cushion lift state control;

[0059] Figure 5 It is a schematic diagram of the encoder extracting cushion lift data features. Detailed Embodiments

[0060] The present invention will be described in detail below with reference to the accompanying drawings. The following provides a detailed design process and specific solution steps, but the protection scope of the present invention is not limited to the following examples.

[0061] One of the disadvantages of the traditional cushion lift state control of hovercraft is the reliance on the professional knowledge and experience accumulated by operators. In complex sea conditions, it is difficult for traditional control methods to fully cover all situations. The limitations of traditional control methods, such as the inability to automatically adapt to changing environments and parameters and the difficulty in achieving global optimization, also increase the difficulty of operator adjustment and decision-making, as well as the operation difficulty and risk.

[0062] The present invention discloses a cushion lift pressure control method for a fully skirted hovercraft based on generative adversarial imitation learning, including:

[0063] Step 1: Establish a mathematical model of the hovercraft, design a controller to initially achieve cushion lift control in the simulation environment, and obtain the first-stage simulation training data. In the present invention, the bottom of the fully skirted hovercraft is divided into air chambers by the cross-division method, the skirt adopts a single-capsule finger structure, and the cushion lift fans are distributed on both sides. First, determine the air chamber layout of the fully skirted hovercraft, the skirt response model, and the structural parameters of the cushion lift fans, establish a mathematical model of the cushion lift system of the fully skirted hovercraft, and obtain the four-air-chamber force balance equation of the hovercraft. Set the step size and the initial cushion lift pressure P c and the airbag pressure P n , and perform one round of iteration at each step size to obtain the attitude information of the hovercraft. Based on the model, design a neural network PID to control the state of the hovercraft, and obtain a time series composed of [p, q, r, u, v, w, P c , P n and the cushion lift fan flow rate [Q]. This time series will be used as the training data for the subsequent steps.

[0064] The following is the four-air-chamber force balance equation:

[0065]

[0066] Among them, W represents the gravity of the hovercraft, F0 represents all longitudinal environmental forces acting on the hovercraft, including wave-making resistance, wind-wave-current force. X x0 , Y y0 , M x0 , M y0 , M z0 are the resultant force and resultant moment generated by the environmental forces, and X x1 , Y y1 , M x1 , M y1 , M z1 are the resultant force and resultant moment of the hovercraft's propeller, side air doors, and air rudders.

[0067] Step 2: Design the structure of the generative adversarial model, which mainly includes a discriminator for driver expert data and a generator for cushioning state control variables.

[0068] The discriminator for expert data adopts an LSTM structure, and the output layer is a softmax layer, with the output between 0 and 1. The features of the hovercraft state S i and the cushioning fan control variable A i are concatenated into a tensor {S i , A i}, which is denoted by τ i . For complex states S i , a variational autoencoder can be used to complete the extraction of state features and then perform concatenation. The concatenated tensor is input into the LSTM network, and the output represents the true probability that the concatenated tensor is expert data.

[0069] The generator for cushioning state control variables adopts Soft Actor Critic (SAC). The network consists of an Actor network, a Critic network, and a Critic-target target network. The Actor network is composed of fully connected layers and Relu layers, and the number of network layers is selected according to the data characteristics. The Actor network takes S j as input and outputs A j . The input of the Critic network is the state S j+1 at the next sampling point, and the output is a score. The Critic-target target network has the same structure as the Critic network and is used to assist in the calculation of the TD error.

[0070] Step 3: Use the data collected in Step 1 to perform the first-stage training on the generative adversarial structure designed in Step 2:

[0071] Step 3.1: Screen and preprocess the data, and use methods such as adding perturbations to enhance the data. Collect training data from different scenarios, balance the sample distribution, and pre-train the actor part of the generator and the discriminator.

[0072] Step 3.2: Based on the pre-training, use the simulation training data obtained in Step 1 for adversarial training. The discriminator loss function and the policy gradient for updating the generator in adversarial training are:

[0073]

[0074]

[0075] Step 3.3: After obtaining the model, use expert data or test data to evaluate the performance of the model and compare the differences with the expert data; if the differences are large, return to Step 3.2 for adversarial training to adjust the structure and parameters.

[0076] Step 4: Use the operation data obtained from the hovercraft driving simulation platform to conduct the second phase of training to further optimize the network parameters and model structure.

[0077] Since the boost system of the driving simulation platform involves more complex processes, many variables are coupled and related, the amount of data is huge, and the variables between systems affect each other, which increases the difficulty of training. Therefore, it is necessary to strengthen the screening and balance of data. On the one hand, the input-output relationship is expanded, the structure is adjusted, and the features of the newly added variables are extracted and coupled to the network. On the other hand, the evaluation indicators are refined and specific evaluation indicators are set for different subjects.

[0078] Step 5: Use the network to assist the driver in manipulating the hovercraft lift system during actual driving. Record the lift system status and environmental information, and learn and optimize in real time. At the same time, use the driver's real-time operation, feedback learning, and update the network.

[0079] Example 1

[0080] A method for controlling the cushion pressure of a full-cushion hovercraft based on generative adversarial imitation learning, comprising:

[0081] Step 1: Determine the full-lift hovercraft cushion system model, and design the control law based on the obtained mathematical model to realize the control of the cushion pressure system of the hovercraft in the simulation environment. Analyze and screen the data generated by the control process to establish the first stage training data set.

[0082] Step 1-1: Establish the six-degree-of-freedom model of the full-lift hovercraft, the pressure distribution model of the full-lift hovercraft air cushion chamber, and the mathematical model of the lift fan. Analyze the characteristics of the air cushion skirt, and obtain the kinematic and dynamic equations in combination with other external forces acting on the hovercraft.

[0083] Six-degree-of-freedom control equations of the six-degree-of-freedom model of the hovercraft:

[0084]

[0085] Among them, I x ,I y ,I z is the moment of inertia, u, v, w are the longitudinal, lateral and vertical velocities, p, q, r are the roll angular velocity, pitch angular velocity and bow angular velocity, F x 、F y 、F z is the resultant force, M x 、M y 、M z is the torque.

[0086] The cushion lift fan is the air supply source for the apron system and the air cushion system. The flow pressure formula of a single cushion lift fan is:

[0087] P f = P t - P fin - P fout

[0088] P fin = C fin ρ a (Q f / S fin ) 2 / 2

[0089] P fout = C fout ρ a (Q f / S fout ) 2 / 2

[0090] Among them, P t is the total flow pressure formula of a single lift fan, P fin and P fout are the inlet pressure loss and the outlet pressure loss of the lift fan respectively, S is the area of the fan inlet and outlet, and C is the loss coefficient of the fan inlet and outlet.

[0091] Combined with the flow pressure equation, the dynamic force expressions of the lift fans on both sides can be obtained:

[0092]

[0093] Among them, Q fi represents the flow rate of the i-th lift fan, (x fi , y fi , z fi ) is the position of the fan.

[0094] The dynamic force of the apron drainage is the reaction force generated by the outward drainage of the gap between the bottom of the apron and the wave surface. Integrating the air cushion pressure P c (x, y, t) at the air cushion edge can obtain the dynamic force and moment, denoted by F I and M I as follows:

[0095] F I = -∫ I P c (x, y, t)n t dI

[0096] M I = -∫ I P c (x, y, t)(r I × n I )dI

[0097] The air chamber of the hovercraft is divided as Figure 1 shown. The air cushion is divided into four air chambers by the cross - division method. The high - pressure gas in the skirt bag flows into the four air chambers, and the gas flow direction is divided into three parts: one part flows from the air cushion through the skirt finger bottom around the air cushion to the atmosphere, one part flows along the bottom edge of the skirt to the adjacent air chamber, and one part fills the air chamber to lift the hull. Assuming that the gas is non - viscous but has potential and is incompressible, the flow - continuity balance relationship of each air chamber is established with the help of the theoretical model of the pressurizing chamber, and then the dynamic model of the four - air - chamber air - cushion system is established.

[0098]

[0099] Among them, Q ni represents the flow rate of the skirt airbag discharging into the i - th air chamber, Q i represents the flow rate of the i - th air chamber discharging to the atmosphere, Q bi is the air - pumping flow rate, and its expression is:

[0100]

[0101] Among them, K ni is the orifice discharge coefficient, A ni is the orifice area, P ci 、P ni are the air - cushion lift pressure and the skirt - bag pressure.

[0102] The Newton - Simpson iteration can be used to solve the air - cushion system flow equation to obtain the pressure and flow rate of each air chamber, which act on the hull as vertical force, rolling moment and pitching moment:

[0103]

[0104] Among them, S i is the horizontal projection area of the four air chambers, x gi 、y gi are the transverse and longitudinal positions of the pressure action point from the center of gravity.

[0105] Combining the above lift fan, skirt force and moment, and air - chamber model, the force - balance equation of the four - air - chamber hovercraft can be summarized as:

[0106]

[0107] Among them, W represents the gravity of the hovercraft, F0 represents all longitudinal environmental forces acting on the hovercraft, including wave - making resistance, wind - wave - current force. X x0 、Y y0 、M x0 、M y0 、M z0 are the resultant force and resultant moment generated by the environmental forces, X x1 、Y y1 、Mx1 , M y1 , M z1 is the resultant force and resultant moment of the hovercraft propeller, side air door, and air rudder. Controlling the lift fan can control the air cushion flow rate, thereby affecting the hull attitude. Combining with the hovercraft kinematic model under different sea conditions, the real-time attitude of the hovercraft can be obtained.

[0108] Step 1-2: Based on Step 1-1, design a controller by combining the six-degree-of-freedom maneuvering equation of the hovercraft, and realize the control of the hovercraft's lift attitude in the simulation environment. Set the sampling step size, record the hull attitude information, lift pressure and other information related to the lift system, and the output of the controlled lift fan, obtain a time series, and screen and process the series as the training data for the first stage of the subsequent model. The controller uses a neural network PID. This controller effectively utilizes the adaptive ability of the neural network. The controller design is as Figure 2 shown, using incremental PID, and its expression is:

[0109] Δu(k) = O K Δe(k) + O I e(k) + O D [Δe(k) - Δe(k - 1)]

[0110]

[0111] where, is the output of each layer of the neural network, k is the layer number of the network, O k , O I , O D correspond to the output layer results.

[0112] Using the controller to control the hovercraft state, the input sequence obtained is: [p, q, r, u, v, w, h, p c , p n , and the output sequence obtained is: [Q f .

[0113] Step 2: Build a hovercraft expert data generation adversarial imitation learning network according to the network shown in Figure 3 . Design the structure of the generative adversarial model, which mainly includes a driver expert data discriminator D(x) and a lift state control variable generator G(z).

[0114] The driver expert data discriminator adopts an LSTM structure. This structure can judge whether the actions taken in each state have the same distribution as the expert data, and can also explore the connection between actions in continuous states, thereby improving the ability to distinguish expert data from generated data. The input of the discriminator is the hovercraft state S i and the lift fan control variable A iThe spliced tensor {S i , A i}, and the tensor is represented by τ i . A i represents the flow rate of the lift fan, and S i includes states such as lift height, lift speed and acceleration, roll angle and angular acceleration, pitch angle and angular acceleration, etc. As shown in Figure 5 , the variational autoencoder is used to encode and abstract low-dimensional data, learn the latent distribution of lift control data, and input the low-dimensional features into the discriminator LSTM network. The size and number of layers of the network hidden layer are set according to the actual data effect. The output layer of the discriminator is a softmax layer, and the output result is the discrimination probability, and the numerical range is between 0 and 1.

[0115] The lift state control variable generator adopts SoftActor Critic (SAC). The generator contains three networks, namely the Actor network, the Critic network and the Critic-target target network. The Actor network consists of a fully connected layer and a Relu layer, and the number of network layers is selected according to the data characteristics. The Actor network inputs S j and outputs A j . The input of the Critic network is the state S of the next sampling point j+1 . The purpose of the Critic network is to evaluate the state S to obtain a score, and the output is the score. The Critic-target target network corresponds to the target value of the Critic network and is used to assist in the calculation of the TD error. After convergence, it has the same parameters as the Critic network. The convergence of the two Critic networks corresponds to the physical meaning of the quality of the lift state obtained after control.

[0116] Step 3: Use the data obtained in Step 1 for the first-stage training, adjust the network structure, optimize the parameters, and initially realize the basic imitation of the expert data in the simulation environment by the generative adversarial network. The process is as shown in the training process of the generative adversarial network in Figure 4 . The specific process is as follows:

[0117] Step 3-1: Screen and preprocess the data. In different application scenarios of hovercraft lift control, add training data to improve the generalization ability of the model. Data augmentation techniques, such as random perturbation, can be used to generate more training data. The lift state and lift fan flow rate of the hovercraft can also be standardized or normalized for easy model learning.

[0118] Step 3-2: Pre-train the lift state control generator model and the expert data discriminator model, restrict the output distribution of the model to improve the initial performance of the generative adversarial model. Screen data sets with different characteristics, such as the lift conditions in the absence of sea waves, the lift conditions under regular waves and irregular waves, and the lift conditions under different wind directions, etc. Divide the lift state data for different operating environments such as beaching, obstacle crossing, docking, etc., and use classical imitation learning to fit the input and output data. Combine self-supervised learning methods, such as autoencoders, etc., to initially realize the functions of generating data and distinguishing data in the network, and the specific process.

[0119] Step 3-2-1: First, use the training data in the first stage to fit the input S of the Actor of the generator j , and output A j . Then, use an unsupervised method to pre-train the Actor of the generator. Specifically, first randomly initialize the parameters of the generator, and then use an autoencoder to train the generator. The input of the autoencoder is random noise, the output is the output of the Actor of the generator, and the middle layer is the hidden layer of the generator. The training objective is to minimize the reconstruction error, that is, the distance between the input and the output.

[0120] The loss function of the pre-trained Actor of the generator is:

[0121]

[0122] where n is the number of samples, x i is the input sample, z i is the random noise, and G is the generator;

[0123] The gradient of the pre-trained Actor of the generator is:

[0124]

[0125] where θ G are the parameters of the generator.

[0126] Step 3-2-2: Use the pre-trained generator to generate "fake" samples and real samples to train the expert data discriminator. If the input sample is a real sample, set its label to 1, otherwise set it to 0, and then use loss functions such as cross-entropy to train the discriminator.

[0127] Step 3-3: The training objective of the generator is to minimize the difference between its output and the expert data, and the training objective of the discriminator is to maximize the probability of correctly distinguishing the generated data from the expert data. The two models are trained adversarially to improve the performance of both network models simultaneously. During the adversarial training process, first fix the discriminator parameters, use the generator to generate pseudo-samples and input them into the discriminator, and calculate the loss function and gradient of the generator based on the discriminator output to update its parameters; then fix the generator parameters, mix the real samples and the generated samples, calculate the gradient of the discriminator through the cross-entropy loss, and update its parameters. This process realizes the adversarial learning of the model through the alternating training of the generator and the discriminator. The specific process of adversarial training is as follows:

[0128] Step 3-3-1: Obtain the expert data sequence and the generated lift control sequence. Input the hovercraft feature state S into the Actor in the lift state control generator j , and obtain the action A in this state j . Use the six-degree-of-freedom motion model of the hovercraft and the mathematical model of the lift system in Step 1 to obtain the next lift state S j+1 . Substitute S j+1 into the generation model continuously to obtain the state-control sequence {S1, A1, S2, A2... S m , A m}. Concatenate S j and A j to obtain the tensor τ j . Use τ to represent the complete sequence {τ1, τ2... τ m}, and the sequence length is m. Similarly, the expert data sequence {τ1', τ2'... τ n '} with the same initial hovercraft feature state can be obtained. Use τ' to represent it, and the sequence length is n.

[0129] Step 3-3-2: Train the expert data discriminator. Input the sequence into the discriminator to obtain the discrimination probabilities of the discriminator for τ and τ', denoted by dτ i '(dτ j ). Calculate the loss function and update the discriminator network parameters by gradient descent:

[0130]

[0131] where D(τ) is the discrimination output result of the expert data discriminator for the input trajectory sample τ; D(τ') is the discrimination output result of the expert data discriminator for the input trajectory sample τ'.

[0132] Step 3-3-3: Retain the discrimination result of the discriminator for the data generated by the cushion lifting state control. Use it as another form of reward to update the generator. The generator uses the Soft Actor Critic algorithm (SAC). The algorithm process is as follows:

[0133] First, calculate the difference Δ between the Critic network and the target network Critic-target, and use the least squares method to perform gradient descent on the Critic network:

[0134] Δ = Critic_target(τ) + reward - Critic(τ).

[0135] Among them, reward is the discrimination result D(τ) of the expert data discriminator for the data generated by the cushion lifting state control;

[0136] After that, according to the principle that the action state only affects the subsequent benefits, use Δ to calculate the advantage function adv for each τ j weight the gradient of the following entropy to obtain the policy gradient, and update the parameters π(θ) of the Actor network in SAC:

[0137]

[0138] Among them, P(τ j |π) is the probability of executing the action τ under the policy π j ; when calculating the policy gradient, it includes which is the entropy value of the current policy. The larger the entropy value, the more dispersed the policy distribution means. A hyperparameter β is set in the policy gradient to balance the dispersion degree of the learning policy. The greater the proportion it occupies means the higher the randomness of the policy. A suitable β is beneficial to fully explore the environment and increase the generalization ability.

[0139] At the end of each episode, the target network Critic-target is updated in a soft-update manner, that is, a very small proportion of new network parameters are updated each time:

[0140] θ target ← λθ target + (1 - λ)θ

[0141] Among them, θ target is the parameter of the target network (used to stabilize training); λ is the update law designed by the soft-update method; θ is the network parameter.

[0142] Step 4: Use the simulated operation data obtained from the hovercraft driving simulation platform for the second-phase training to further optimize the network parameters and model structure. The lift system of the driving simulation platform involves a more complex process, with many variables coupled and correlated, a huge amount of data, and the variables between systems affecting each other, which increases the training difficulty. The specific process is as follows:

[0143] Step 4-1: Screen the data and use data augmentation techniques to expand the dataset, such as adding random wave wind direction disturbances to expand the coverage of the dataset. To improve the generalization ability, assign greater weights to low-probability events. For different hovercraft driving training subjects, such as beach landing and climbing, low-speed straight navigation, high-speed turning, etc., set specific sampling weights according to the different requirements of the lift system in different scenarios to solve the problem of sample imbalance.

[0144] Step 4-2: Expand the input and output data and adjust the network structure and quantity. The fan speed-flow control unit of the hovercraft driving simulator is included in the hovercraft engine control system. The engine operation process is complex, including self-check, start-up, warm-up, shutdown, etc. The engine working mode will change according to the hovercraft driving mode, including automatic control mode, manual operation mode, etc. There are many and complex engine state variables, such as turbine speed, internal pressure, working temperature, etc. Use an autoencoder to encode and extract features from various simulation data of the engine, then splice the feature codes and input them into the LSTM network to output the discrimination results of actions executed in a continuous state. The expanded network structure is as Figure 5 shown in the combined structure, where A, B, C…S are the states of different mechanisms in the system.

[0145] Step 4-3: Evaluate the trained model. In addition to setting network training indicators, combine the expected requirements of the lift pressure control system in actual work, add model indicators, obtain model evaluation feedback, and adjust and optimize the network structure parameters. The specific indicators include boiler pressure, engine efficiency, main reducer working threshold, fuel supply pressure, etc. Set specific constraints and evaluation criteria for special subjects, such as spin operation training, large roll operation training, entering and leaving the mooring area operation training, apron damage operation training, etc. For example, when designing the apron damage operation training, design the hovercraft operation stability time and roll angle constraints as evaluation indicators, and when designing the entering and leaving the mooring area training, design the lift height, entering and leaving time, parking attitude, etc. as evaluation indicators.

[0146] Step 5: During the actual driving process, apply the trained network to assist the driver in operating the hovercraft lift system. During the process of using network assistance, record the state of the lift system and the surrounding environment information, and screen and process this data for feedback and optimization.

[0147] During the optimization process, a feedback learning method is used to feed the actual driving data back to the network and optimize according to the results output by the network. By continuously iteratively optimizing the model, the network can be made to more accurately predict the state and response of the lift system.

[0148] Substitute the working conditions of the hovercraft lift system and consider the working environment and characteristics of the hovercraft. Hovercraft usually operate on water or other smooth surfaces and need to control the height and direction through the lift system. Therefore, when training the network, we need to consider environmental factors such as the center of gravity of the hull, wind speed, water flow, etc., as well as factors such as the response speed and stability of the lift system. By appropriately adjusting the training data and optimization algorithm, we can make the network more accurately predict the state and response of the hovercraft lift system, thereby improving the safety and efficiency of driving. The overall process of progressive training is as Figure 5 shown.

[0149] In particular, in some preferred embodiments of the present invention, a computer device is further provided, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the steps of the full-lift hovercraft lift pressure control method based on generative adversarial imitation learning described in any of the above embodiments are implemented.

[0150] In some other preferred embodiments of the present invention, a computer-readable storage medium is further provided, on which a computer program / instructions are stored. When the computer program is executed by a processor, the steps of the full-lift hovercraft lift pressure control method based on generative adversarial imitation learning described in any of the above embodiments are implemented.

[0151] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the full-lift hovercraft lift pressure control method embodiment based on generative adversarial imitation learning as described above, which will not be repeated here.

[0152] In summary, the lift system control method based on adversarial imitation learning proposed by the present invention largely solves problems such as the over-reliance on the driver's experience in the lift control of a fully-lifted hovercraft, and reduces the difficulty of lift operation. By adopting a control method based on generative adversarial imitation learning and using a progressive training method, the intelligent control of the lift pressure of a fully-lifted hovercraft is realized, thereby helping the driver reduce the operation difficulty and improve the operation efficiency. The present invention uses a combination of a variational autoencoder and an LSTM network to construct an expert data discriminator network. This structure improves the discrimination ability of the expert data discriminator for expert data and generated data, and promotes the performance of the entire generative adversarial imitation learning network. The present invention uses a SoftActor Critic network as the generator of the lift state control variable. For the Actor network therein, on the one hand, an advantage function is set in combination with the context before and after the sequence, and on the other hand, an item is added when calculating the policy gradient to improve the diversity of the search strategy of the Actor and increase the generalization ability, realizing the adaptive control of the lift pressure of a fully-lifted hovercraft, being able to automatically adapt to different environments and working conditions, and improving the robustness and stability of the system.

[0153] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the method of identifying and tracking aquatic biological communities using a bionic robotic fish as described above.

[0154] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0155] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0156] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0157] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate and then storing it in a computer memory.

[0158] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following technologies known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0159] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium, and when the program is executed, it includes one or a combination of the steps of the method embodiments.

[0160] In addition, in each of the embodiments of the present invention, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0161] The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A lift pressure control method for a fully air-cushioned hovercraft based on generative adversarial imitation learning, characterized in that, Including: Step 1: Establish a mathematical model of the lift system of an air-cushion vehicle with full lift, and design a controller according to the six-degree-of-freedom model of the air-cushion vehicle with full lift to initially achieve lift control in the simulation environment and obtain simulation training data; Step 2: Construct a generative adversarial model; The structure of the generative adversarial model includes: an expert data discriminator and a lift state control variable generator; the expert data discriminator uses a combination of a multi-variational autoencoder and an LSTM structure to discriminate between the generated data and the real data of the input samples; the lift state control variable generator uses the SAC algorithm, and the SAC network model includes: an Actor network, a Critic network, and a Critic-target network; Step 3: Use the simulation training data and the simulation data of the air-cushion vehicle driving simulation platform to train the constructed generative adversarial model in a progressive training manner; first, use the simulation training data to conduct the first-stage training of the generative adversarial model; then use the simulation data obtained from the air-cushion vehicle driving simulation platform to conduct the second-stage training to further optimize the network parameters of the expert data discriminator; use the simulation data to input the multi-variational autoencoder for encoding to extract features, splice the feature encodings, input them into the LSTM network, and output the discrimination results of the actions executed in the continuous state; Step 4: The driver inputs the current working state of the air-cushion vehicle into the generative adversarial model trained in Step 4, and outputs the lift flow rate of the air-cushion vehicle; use the driver operation data samples as the results to further optimize the generative adversarial model.

2. The cushion lift pressure control method for an all-foil air-cushion vehicle based on generative adversarial imitation learning according to claim 1, wherein, The specific content of Step 1 includes: Step 1.1: The six-degree-of-freedom maneuvering equation of the six-degree-of-freedom motion model of the air-cushion vehicle with full lift: Among them, I x , I y , I z are the moments of inertia, u, v, w are the longitudinal, lateral, and vertical velocities, p, q, r are the roll angular velocity, pitch angular velocity, and yaw angular velocity, F x , F y , F z are the resultant forces, M x , M y , M z are the torques; According to the air chamber layout of the air-cushion vehicle with full lift, the skirt response model, and the structural parameters of the lift fan, construct a mathematical model of the lift system of the air-cushion vehicle with full lift, and obtain the four-air-chamber force balance equation of the air-cushion vehicle; Among them, W is the gravity of the hovercraft; F0 is all longitudinal environmental forces acting on the hovercraft; X x0 , Y y0 , M x0 , M y0 , M z0 are the resultant force and resultant moment generated by environmental forces; X x1 , Y y1 , M x1 , M y1 , M z1 are the resultant force and resultant moment of the hovercraft's propellers, side air doors, and air rudders; Step 1.2: Set the step size and the initial air-cushion lift pressure P c and the airbag pressure P n , and perform one round of iteration at each step size to obtain the hovercraft attitude information; Step 1.3: According to the hovercraft attitude information, use neural network PID to control the hovercraft state, and obtain a time series composed of [p, q, r, u, v, w, P c , P n and the lift fan flow rate [Q] as training data.

3. The cushion lift pressure control method for an all-cushion lift hovercraft based on generative adversarial imitation learning according to claim 1, wherein Step 3 includes the following steps: Step 3.1: Screen and preprocess the simulation training data to construct a training set; Step 3.2: Pretrain the lift state control variable generator and the expert data discriminator; first, use the simulation training data to pretrain the lift state control variable generator; then use the pretrained lift state control variable generator to generate fake samples and real samples, and cross-train the expert data discriminator using loss functions such as entropy. Step 3.3: Conduct adversarial training using the lift state control variable generator and the expert data discriminator, and adjust the parameters of the lift state control variable generator and the expert data discriminator to make the lift state control variable generator generate more realistic synthetic data; during the adversarial training process, first fix the parameters of the expert data discriminator, use the lift state control variable generator to generate fake samples and input them into the expert data discriminator, calculate the loss function and gradient of the lift state control variable generator based on the output of the expert data discriminator to update its parameters; then fix the parameters of the lift state control variable generator, mix the real samples and the generated samples, calculate the gradient of the expert data discriminator through the cross-entropy loss and update its parameters; this process realizes the adversarial learning of the generative adversarial model through the alternating training of the lift state control variable generator and the expert data discriminator; Step 3.4: Use expert data to evaluate the performance of the generative adversarial model; judge whether the generative adversarial model is optimized and meets the standard according to the lift fan flow rate [Q], average return, and average method; if the generative adversarial model does not meet the standard, return to Step 3.

3.

4. The lift pressure control method for a fully-elevated hovercraft based on generative adversarial imitation learning according to claim 3, wherein In step 3.2, the pre-trained lift state control variable generator inputs S to the Actor network j and outputs A j . Then, an unsupervised method is used to pre-train the Actor network. First, the parameters of the lift state control variable generator are randomly initialized, and then an autoencoder is used to train the lift state control variable generator. The input of the autoencoder is random noise, the output is the output of the Actor network, and the middle layer is the hidden layer of the lift state control variable generator. The loss function of the Actor network of the lift state control variable generator is: where n is the number of samples, x i is the input sample, z i is the random noise, and G is the generator of cushioning state control variables; The gradient of the Actor network of the lift state control variable generator is: where θ G is the parameter of the lift state control variable generator, is the partial derivative of the parameter of the lift state control variable generator.

5. The lift pressure control method for a fully air-cushioned hovercraft based on generative adversarial imitation learning according to claim 3, characterized in that The specific steps of the adversarial training in Step 3.3 are as follows: Step 3.3.1: Obtain the expert data sequence τ' and the generated lift control sequence τ; Step 3.3.2: Train the expert data discriminator; input the expert data sequence τ' and the generated lift control sequence τ into the expert data discriminator to obtain the discrimination probabilities dτ′ i and dτ j ; calculate the loss function and update the network parameters of the expert data discriminator using the gradient descent loss function; Step 3.3.3: Use the discrimination result of the expert data discriminator on the lift state control generated data as the reward, and update the parameters of the lift state control variable generator using the policy gradient method.

6. The cushion lift pressure control method for an all-cushion lift hovercraft based on generative adversarial imitation learning according to claim 5, characterized in that, In step 3.3.1, the expert data sequence is obtained as follows: Input the expert feature state S i ' into the Actor network in the lift state control generator to obtain the action A i ' in this state. Use the six-degree-of-freedom motion model of the fully-lifted hovercraft and the mathematical model of the fully-lifted hovercraft lift system in step 1 to obtain the next lift state S i '. +1 ; Continuously substitute S i ' +1 into the lift state control generator to obtain the state-control sequence {S′1,A′1,S′2,A'2,...,S' m ,A' m}; Concatenate S′ i and A′ i to obtain the tensor τ′ i . Represent the complete sequence {τ1′,τ2′...τ n ′} with τ', and the sequence length is n. The generation of the lift control sequence is obtained as follows: Input the hovercraft feature state S into the Actor network in the lift state control generator j , and obtain the action A in this state j . Using the six-degree-of-freedom motion model of the fully-lifted hovercraft and the mathematical model of the lift system of the fully-lifted hovercraft in step 1, obtain the next lift state S j+1 ; Substitute S j+1 into the lift state control generator continuously to obtain the state-control sequence {S1, A1, S2, A2... S m , A m}; Concatenate S j and A j to obtain the tensor τ j , and use τ to represent the complete sequence {τ1, τ2... τ m}, and the length of the sequence is m 7. The method for controlling the lift pressure of an air-cushion vehicle with full lift based on generative adversarial imitation learning according to claim 6, wherein In Step 3.3.2, the gradient descent loss function of the expert data discriminator is: where D(τ) is the discrimination output result of the expert data discriminator on the input trajectory sample τ; D(τ') is the discrimination output result of the expert data discriminator on the input trajectory sample τ'.

8. The full-lift hovercraft lift pressure control method based on generative adversarial imitation learning according to claim 7, characterized in that, In Step 3.3.3, the lift state control variable generator is updated by first calculating the difference Δ between the Critic network and the target network Critic-target: Δ = Critic_target(τ) + reward - Critic(τ). where reward is the discrimination result D(τ) of the expert data discriminator on the lift state control generated data; Using Δ to calculate the advantage function adv for each τ j Weight the gradient of the lower entropy to obtain the policy gradient Update the Actor network parameter π(θ) through the policy gradient method: where P(τ j |π) is the probability of taking action τ j under policy π; H(π(S i ; θ)) is the entropy value of the current policy, At the end of each episode, the Critic-target target network is updated in a soft-update manner, that is, a very small proportion of new network parameters are updated each time: θ target ← λθ target +(1 - λ)θ where, θ target is the parameter of the target network; λ is the update law designed by the soft-update method; θ is the network parameter.

9. A computer device / apparatus / system, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 8 are implemented.