An air-cushion vehicle lift control method, device, equipment and storage medium
By constructing the sample data of the hovercraft and conducting neural network training, calculating and applying control parameters, the problem of instability of the hovercraft lifting system under harsh sea conditions is solved, and its intelligence and autonomy are improved.
Patent Information
- Application Number
- CN202510368727.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The hovercraft's pad lifting system is affected by the terrain environment under harsh sea conditions, causing air cushion pressure fluctuations, causing violent rise and sinking movements, increasing the instability of the ship.
By obtaining the current actual attribute data of the hovercraft, setting the expected attribute data, using the first strategic neural network and the first value network to construct sample data, conduct training and update, and calculate control parameters to control the operation of the hovercraft.
The intelligence and autonomy of the hovercraft are improved, so that it can better adapt to various environmental and task requirements, and reduce the instability of the ascending and sinking movement.
Smart Images

Figure CN119882413B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data control, and particularly to a hovercraft lift control method, device, equipment and storage medium. Background Art
[0002] The fully air-cushioned hovercraft is an amphibious ship that can achieve autonomous lift on various terrains such as land, shoals and water surfaces. The lift system of the hovercraft is a key element to realize the unique amphibious characteristics of the hovercraft, and it is also the premise and foundation for the hovercraft to achieve autonomous navigation. However, the air cushion generated by the lift system is extremely vulnerable to the terrain environment. Under severe sea conditions, the wave pumping motion caused by waves and the undulating water surface and the change of the skirt discharge flow will significantly affect the air cushion pressure. This effect causes the hovercraft to suffer severe heaving motion, increasing the instability of the ship. Summary of the Invention
[0003] The present disclosure provides a hovercraft lift control method, device, equipment and storage medium to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, there is provided a hovercraft lift control method, including:
[0005] Obtain the current actual attribute data of the hovercraft, and set the desired attribute data of the hovercraft based on the actual attribute data;
[0006] Determine at least one sample data of the hovercraft based on a first policy neural network, the desired attribute data and the actual attribute data;
[0007] Input the at least one sample data into a first value network for training to obtain a first value function, and update the first policy neural network and the first value network based on the first value function to obtain a second value network and a second policy neural network;
[0008] Calculate a first control parameter based on the second value network and the second policy neural network, and control the operation of the hovercraft based on the first control parameter.
[0009] In an implementable manner, the determining at least one sample data of the hovercraft based on the first policy neural network, the desired attribute data and the actual attribute data includes:
[0010] Calculate the error data of the hovercraft at the current moment based on the desired attribute data and the actual attribute data, and determine the initial state data of the hovercraft at the current moment based on the error data;
[0011] Determine the first action data of the hovercraft based on the first policy neural network, input the first action data into the first control algorithm, and determine the state data of the hovercraft at the next moment;
[0012] Construct the reward function of the hovercraft based on the error data, and determine the first sample data based on the reward function, the initialization state data, the state data at the next moment, and the first action data.
[0013] In an implementable manner, the calculating the error data of the hovercraft at the current moment based on the expected attribute data and the actual attribute data, and determining the initialization state data of the hovercraft at the current moment based on the error data includes:
[0014] Obtain the actual attribute data at the previous moment, and calculate the error change rate data of the hovercraft based on the actual attribute data at the previous moment and the actual attribute data at the current moment;
[0015] Calculate the error vector data of the hovercraft based on the expected attribute data and the actual attribute data;
[0016] Determine the initialization state data of the hovercraft at the current moment based on the error vector data and the error change rate data.
[0017] In an implementable manner, the updating the first policy neural network and the first value network based on the first value function includes:
[0018] Obtain the initialization network parameters of the first policy neural network and the first value network respectively;
[0019] Update the initialization network parameters based on the first value function to obtain the current network parameters corresponding to the first policy neural network and the first value network;
[0020] Determine the value network parameters based on the initialization network parameters and the current network parameters, and update the first policy neural network and the first value network based on the value network parameters.
[0021] In an implementable manner, the constructing the reward function of the hovercraft based on the error data includes:
[0022] Determine the error reward function based on the error vector data, and determine the error change rate reward function based on the error change rate data;
[0023] Set the first weight coefficient and the second weight coefficient, and construct the reward function of the hovercraft based on the first weight coefficient, the second weight coefficient, the error reward function, and the error change rate reward function.
[0024] According to a second aspect of the present disclosure, a cushion lift control device for an air-cushion vehicle is provided. The device includes:
[0025] A data acquisition unit configured to acquire actual attribute data of the air-cushion vehicle at present, and set desired attribute data of the air-cushion vehicle based on the actual attribute data;
[0026] A sample data calculation unit configured to determine at least one sample data of the air-cushion vehicle based on a first policy neural network, the desired attribute data, and the actual attribute data;
[0027] A neural network update unit configured to input the at least one sample data into a first value network for training to obtain a first value function, and update the first policy neural network and the first value network based on the first value function to obtain a second value network and a second policy neural network;
[0028] A control parameter calculation unit configured to calculate a first control parameter based on the second value network and the second policy neural network, and control the operation of the air-cushion vehicle based on the first control parameter.
[0029] In an implementable embodiment, the sample data calculation unit is further configured to calculate error data of the air-cushion vehicle at the current moment based on the desired attribute data and the actual attribute data, and determine initialization state data of the air-cushion vehicle at the current moment based on the error data; determine first action data of the air-cushion vehicle based on the first policy neural network, input the first action data into a first control algorithm to determine state data of the air-cushion vehicle at the next moment; construct a reward function of the air-cushion vehicle based on the error data, and determine first sample data based on the reward function, the initialization state data, the state data at the next moment, and the first action data;
[0030] The neural network update unit is further configured to respectively obtain initialization network parameters of the first policy neural network and the first value network; update the initialization network parameters based on the first value function to obtain current network parameters corresponding to the first policy neural network and the first value network; determine value network parameters based on the initialization network parameters and the current network parameters, and update the first policy neural network and the first value network based on the value network parameters.
[0031] In one implementable manner, the sample data calculation unit is further configured to obtain the actual attribute data at the previous moment, calculate the error change rate data of the hovercraft based on the actual attribute data at the previous moment and the current actual attribute data; calculate the error vector data of the hovercraft based on the desired attribute data and the actual attribute data; determine the current initialization state data of the hovercraft based on the error vector data and the error change rate data;
[0032] Determine an error reward function based on the error vector data, and determine an error change rate reward function based on the error change rate data; set a first weight coefficient and a second weight coefficient, and construct a reward function of the hovercraft based on the first weight coefficient, the second weight coefficient, the error reward function, and the error change rate reward function.
[0033] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0034] At least one processor; and
[0035] A memory communicatively connected to the at least one processor; wherein,
[0036] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the present disclosure.
[0037] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method described in the present disclosure.
[0038] The hovercraft lift control method, device, equipment and storage medium of the present disclosure construct sample data corresponding to the hovercraft through the actual attribute data and the desired attribute data of the hovercraft, and use the sample data as the training samples of the first policy neural network and the first value network to update the network parameters. Calculating the control parameters of the hovercraft based on the neural network with updated network parameters can improve the intelligence and autonomy of the hovercraft, enabling the hovercraft to better adapt to various environments and task requirements.
[0039] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0040] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understandable. In the drawings, several embodiments of the present disclosure are shown in an exemplary but not restrictive manner, where:
[0041] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0042] Figure 1 The implementation process schematic of a hovercraft lift control method according to an embodiment of the present disclosure is shown. Figure 1 ;
[0043] Figure 2 The implementation process schematic of a hovercraft lift control method according to an embodiment of the present disclosure is shown. Figure 2 ;
[0044] Figure 3 The schematic diagram of a hovercraft lift control device according to an embodiment of the present disclosure is shown;
[0045] Figure 4 The schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure is shown. Detailed implementation manners
[0046] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present disclosure.
[0047] Figure 1 The implementation process schematic of a hovercraft lift control method according to an embodiment of the present disclosure is shown. Figure 1 , as Figure 1 shown, a hovercraft lift control method according to an embodiment of the present disclosure includes the following steps:
[0048] Step 101, obtain the current actual attribute data of the hovercraft, and set the expected attribute data of the hovercraft based on the actual attribute data.
[0049] In the embodiments of the present disclosure, the attribute data of the hovercraft includes: heave height, heave speed, and air chamber pressure. The corresponding actual attribute data is the attribute data of the hovercraft at the current moment. According to the current actual attribute data and the target lift position, the expected attribute data of the hovercraft is set.
[0050] Step 102: Determine at least one sample data of the hovercraft based on the first policy neural network, the expected attribute data, and the actual attribute data.
[0051] In the embodiments of the present disclosure, calculate the error data of the hovercraft at the current moment based on the expected attribute data and the actual attribute data, and determine the initial state data of the hovercraft at the current moment based on the error data; wherein, by obtaining the actual attribute data at the previous moment, calculate the error change rate data of the hovercraft based on the actual attribute data at the previous moment and the actual attribute data at the current moment; calculate the error vector data of the hovercraft based on the expected attribute data and the actual attribute data; determine the initial state data of the hovercraft at the current moment based on the error vector data and the error change rate data, and the initial state data consists of the heave height error, heave speed error, and air chamber pressure error of the hovercraft at the current moment. Input the initial state data into the first policy neural network, and determine the first action data of the hovercraft based on the first policy neural network, wherein the first policy neural network is the policy neural network (Actor) of Deep Deterministic Policy Gradient (DDPG). Input the first action data into the first control algorithm to determine the state data of the hovercraft at the next moment, wherein the first control algorithm is the current control algorithm of the DDPG neural network.
[0052] In the embodiments of the present disclosure, construct the reward function of the hovercraft based on the error data, wherein determine the error reward function based on the error vector data, and determine the error change rate reward function based on the error change rate data; set the first weight coefficient and the second weight coefficient, and construct the reward function of the hovercraft based on the first weight coefficient, the second weight coefficient, the error reward function, and the error change rate reward function. Determine the first sample data based on the reward function, the initial state data, the state data at the next moment, and the first action data. Store the first sample data in the experience pool as the training data set. Meanwhile, the experience pool is also used to store and replay historical experience data.
[0053] Step 103: Input the at least one sample data into the first value network for training to obtain the first value function, and update the first policy neural network and the first value network based on the first value function to obtain the second value network and the second policy neural network.
[0054] In the embodiments of the present disclosure, at least one sample data is selected from the experience pool and input into the first value network for model training. The first value network is the value network (Critic) in the DDPG neural network, and the generated first value function is the value function of the state-action pair, including the state and action of the hovercraft at the current moment.
[0055] In the embodiments of the present disclosure, the initial network parameters of the first policy neural network and the first value network are obtained respectively; the initial network parameters are updated based on the first value function. Specifically, the current target value is calculated according to the first value function, and the initial network parameters of the Critic and the initial network parameters of the Actor are updated to obtain the current network parameters corresponding to the first policy neural network and the first value network; a learning rate parameter is set, and the value network parameters are determined based on the learning rate parameter, the initial network parameters and the current network parameters. The first policy neural network and the first value network are updated based on the value network parameters to obtain the updated second value network and the second policy neural network.
[0056] In the embodiments of the present disclosure, the initial round of iteration enters a loop, and the maximum number of rounds is set to M. The maximum number of steps per round is set to N. When the maximum number of steps N is reached, the iterative training of this round ends. At this time, the number of iterations is incremented by 1, and the operations in the above steps 101-103 are repeated, and so on, until the preset maximum number of iterations M is reached. After reaching the maximum number of iterations, the updated second value network and the second policy neural network are obtained.
[0057] Step 104, calculate the first control parameter based on the second value network and the second policy neural network, and control the operation of the hovercraft based on the first control parameter.
[0058] In the embodiments of the present disclosure, the process of determining the first action data of the hovercraft to the updated second value network and the second policy neural network is set as a round, and at the same time, the maximum number of rounds M is set. The maximum number of steps per round is N. When the maximum number of steps N is reached, the iterative training of this round ends. At this time, the number of iterations is incremented by 1, and the above steps are repeated, and so on, until the preset maximum number of iterations M is reached.
[0059] In the embodiments of the present disclosure, a linear sliding mode switching function is designed according to the error change rate data, the lift motion model of the hovercraft is substituted into the linear sliding mode switching function, and iterative training is carried out in combination with the second value network and the second policy neural network. Specifically, the control rate parameter is calculated based on the maximum number of steps N per round as described above, and the set of control rate parameters obtained per round is iteratively trained based on the maximum number of rounds M to obtain the first control rate parameter, that is, the iteratively optimized control rate parameter.
[0060] Figure 2 The implementation process of a hovercraft lift control method according to an embodiment of the present disclosure is shown schematically Figure 2 , as Figure 2 shown, a hovercraft lift control method according to an embodiment of the present disclosure includes the following steps:
[0061] Step 201, obtain the attribute data of the hovercraft.
[0062] In the embodiment of the present disclosure, an Actor network and a Critic value network are first constructed, and their parameters and are initialized. At the same time, an experience pool is initialized to store and replay experience data. At the same time, the initial round iteration (Episode = 1) enters a loop, and the maximum number of rounds is set to M. The maximum number of steps per round is set to N. The attribute data of the hovercraft includes heave height, heave speed, and chamber pressure, which correspondingly include actual values and expected values. According to the obtained attribute data, the expected attribute data is input, that is, the expected heave height , heave speed , and chamber pressure .
[0063] Step 202, construct the sample data of the hovercraft.
[0064] In the embodiment of the present disclosure, assuming that the pitch angle has only a small change range, the heave motion model of the hovercraft is simplified to:
[0065]
[0066] where represents the first derivative of the heave position, represents the heave speed, represents the mass of the hovercraft, represents the first derivative of the heave speed, represents the longitudinal speed, represents the lateral speed, represents the roll angular velocity, represents the pitch angular velocity, represents the chamber chamber pressure, represents the chamber chamber area, represents the acceleration due to gravity, represents the air density, represents the relative wind speed, represents the aerodynamic (moment) coefficient, represents the horizontal projected area of the hull.
[0067] In the embodiment of the present disclosure, the error at this moment is calculated respectively With the error change rate , the initial state is obtained , specifically: The error vector is expressed as:
[0068]
[0069] where represents the heave position error represents the heave velocity error represents the air chamber pressure error represents the actual heave position represents the desired heave position, with the unit of m represents the actual heave velocity represents the desired heave velocity, with the unit of m / s. P represents the air chamber pressure of the hovercraft represents the desired air chamber pressure
[0070] The state vector of the state space is:
[0071] where represents the heave position error at time t represents the heave velocity error at time t represents the air chamber pressure error at time t
[0072] In the embodiments of the present disclosure, based on the Actor policy network, the state vector is input into the Actor policy network, and an action is selected according to the current policy , and the action space expression is:
[0073]
[0074] where represents the parameter of the heave position error represents the parameter of the heave velocity error
[0075] In the embodiments of the present disclosure, the action is input into the neural network DDPG to obtain the state at the next moment , and the immediate reward function is calculated. Among them, in the hovercraft lift height control system, the design of the reward function should consider the error between the expected value and the actual value, and the influence of the change rate of the error on the control performance of the system, and reflect the stability and response speed of the system according to the change rate of the error. A smaller error and a stable error change rate should obtain a higher reward to encourage the system to maintain the expected height and correct the error in a timely manner. The reward function is used to qualitatively measure the evaluation of the environment on the performance of the control system, so as to guide the decision-making of the intelligent agent to take actions in different environmental states. Therefore, the change rate of its error can be defined as:
[0076]
[0077] Among them, represents the heave position error, represents the heave speed error, represents the air chamber pressure error. Design the reward function The expression is:
[0078]
[0079]
[0080]
[0081] Among them, and are weight coefficients. , are the reward functions of the error and the reward function of the error change rate respectively. and are the minimum error tolerance and the maximum error tolerance respectively. When and , the hovercraft is at the desired cushion height.
[0082] In the embodiments of the present disclosure, based on the above calculations, the corresponding sample data is obtained and stored in the experience pool as the training data set.
[0083] Step 203, neural network update.
[0084] In the embodiments of the present disclosure, N groups of sample data are extracted from the experience pool to start training. Using the current state and the current state as the input of the Critic current network, the value function of the state-action pair is obtained. Calculate the current target value according to the formula and update the Critic network parameters and the Actor network parameters . The specific expressions for updating the current network parameters and are respectively:
[0085]
[0086]
[0087] Among them, N is the sampling amount, is the sample weight, is the target action value, is the discount factor. Update the target parameters of the Critic network according to the formula and the target parameters of the Actor network . The specifically updated value network parameters and the policy network parameters The expression is:
[0088]
[0089] where is the learning rate.
[0090] Step 204, calculate the control parameter, and control the cushion lift of the hovercraft based on the control parameter.
[0091] In the embodiments of the present disclosure, since the air chamber pressure error is determined by the rotational speed of the cushion lift fan, the controller is mainly designed to control the heave position error and the heave speed error. According to the error change rate in step 202, we can obtain:
[0092]
[0093] In the formula, the actual heave position and the heave speed are defined as the errors between the expected values respectively. Take the first derivative of the above formula, and we can obtain:
[0094]
[0095] Design the linear sliding mode switching function as:
[0096]
[0097] where , , are the sliding mode parameters, , , . Take the first derivative of the above formula, and we can obtain:
[0098]
[0099] Substitute the formula of the hovercraft heave motion model into the above formula, and we can obtain:
[0100]
[0101] According to the stability condition, let , then the designed control parameter is derived as:
[0102]
[0103] where is the sliding mode parameter. is the sign function.
[0104] In the embodiment of the present disclosure, when the maximum number of steps N is reached, this round of iterative training ends, and the sliding mode control parameter at this moment can be obtained according to the formula. Otherwise, return to step 201, and at the same time, the iteration number Episode is incremented by 1, and the operations are continued according to steps 201 to 203, and so on, until the preset maximum number of iterations M is reached, and the sliding mode control parameter after iterative optimization is obtained. Since the DDPG algorithm based on reinforcement learning needs to continuously learn and train to obtain a good control strategy. This learning process requires a large amount of sample data and iteration times, so the time cost is relatively high. Therefore, when the hovercraft autonomously lifts on land, the traditional control algorithm has a better control effect, does not require a large amount of training data and iteration process, and can quickly achieve a stable control effect. For other environments, the DDPG algorithm is superior to other algorithms in terms of steady-state convergence time and steady-state error, achieving the best control effect, and its performance is also significantly better than the fuzzy adaptive control method. In contrast, the traditional PID control method has the worst effect.
[0105] In the embodiment of the present disclosure, the designed controller combines the learning ability of reinforcement learning and the stability of the sliding mode controller, and adjusts the parameters of the sliding mode controller in real time according to different terrain environments. This controller makes up for the shortcoming of the insufficient adaptive ability of the traditional controller, and does not require any manual tuning and parameter adjustment. Through the DDPG algorithm, the controller can autonomously learn and optimize parameters according to real-time feedback information and environmental conditions to achieve a better control effect. Compared with the traditional control method, this algorithm has higher stability, enables the controller to better adapt to complex and changing environments, and achieves a more accurate control effect.
[0106] Figure 3 shows a schematic diagram of a hovercraft lift control device according to an embodiment of the present disclosure, as Figure 3 shown, a hovercraft control device according to an embodiment of the present disclosure includes:
[0107] A data acquisition unit 301, configured to acquire the current actual attribute data of the hovercraft, and set the desired attribute data of the hovercraft based on the actual attribute data;
[0108] A sample data calculation unit 302, configured to determine at least one sample data of the hovercraft based on the first policy neural network, the desired attribute data, and the actual attribute data.
[0109] The sample data calculation unit 302 is further configured to calculate the error data of the hovercraft at the current moment based on the expected attribute data and the actual attribute data, and determine the initialization state data of the hovercraft at the current moment based on the error data; determine the first action data of the hovercraft based on the first policy neural network, input the first action data into the first control algorithm to determine the state data of the hovercraft at the next moment; construct the reward function of the hovercraft based on the error data, and determine the first sample data based on the reward function, the initialization state data, the state data at the next moment, and the first action data.
[0110] The sample data calculation unit 302 is further configured to obtain the actual attribute data at the previous moment, and calculate the error change rate data of the hovercraft based on the actual attribute data at the previous moment and the actual attribute data at the current moment; calculate the error vector data of the hovercraft based on the expected attribute data and the actual attribute data; determine the initialization state data of the hovercraft at the current moment based on the error vector data and the error change rate data; determine the error reward function based on the error vector data, and determine the error change rate reward function based on the error change rate data; set the first weight coefficient and the second weight coefficient, and construct the reward function of the hovercraft based on the first weight coefficient, the second weight coefficient, the error reward function, and the error change rate reward function.
[0111] The neural network update unit 303 is configured to input the at least one sample data into the first value network for training to obtain the first value function, and update the first policy neural network and the first value network based on the first value function to obtain the second value network and the second policy neural network.
[0112] The neural network update unit 303 is further configured to respectively obtain the initialization network parameters of the first policy neural network and the first value network; update the initialization network parameters based on the first value function to obtain the current network parameters corresponding to the first policy neural network and the first value network; determine the value network parameters based on the initialization network parameters and the current network parameters, and update the first policy neural network and the first value network based on the value network parameters.
[0113] The control parameter calculation unit 304 is configured to calculate the first control parameter based on the second value network and the second policy neural network, and control the operation of the hovercraft based on the first control parameter.
[0114] In an exemplary embodiment, the data acquisition unit 301, the sample data calculation unit 302, the neural network update unit 303, the control parameter calculation unit 304, etc. can be implemented by one or more central processing units (CPUs, Central Processing Unit), graphics processing units (GPUs, Graphics Processing Unit), application specific integrated circuits (ASICs, Application Specific Integrated Circuit), DSPs, programmable logic devices (PLDs, ProgrammableLogic Device), complex programmable logic devices (CPLDs, Complex Programmable Logic Device), field programmable gate arrays (FPGAs, Field-Programmable Gate Array), general purpose processors, controllers, microcontroller units (MCUs, Micro Controller Unit), microprocessors (Microprocessor), or other electronic components.
[0115] Regarding the device in the above embodiment, the specific manners in which each module and unit perform operations have been described in detail in the embodiment related to the method, and will not be elaborated herein.
[0116] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0117] Figure 4 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0118] As Figure 4As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 can also be stored. The computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0119] Multiple components in device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0120] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as a hovercraft lift control method. For example, in some embodiments, a hovercraft lift control method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of a hovercraft lift control method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute a hovercraft lift control method in any other appropriate way (e.g., by means of firmware).
[0121] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0122] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0123] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).
[0125] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0126] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.
[0127] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0128] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of this disclosure, "a plurality" means two or more unless otherwise specifically defined.
[0129] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claimed rights.
Claims
1. A method for controlling the lift of a hovercraft, characterized in that: The method comprises: Acquiring current actual attribute data of the hovercraft, and setting expected attribute data of the hovercraft based on the actual attribute data, wherein the actual attribute data includes heave height, heave speed, and air chamber pressure; Determine at least one sample data of the hovercraft based on the first strategy neural network, the expected attribute data and the actual attribute data; Inputting the at least one sample data into a first value network for training to obtain a first value function, and updating the first policy neural network and the first value network based on the first value function to obtain a second value network and a second policy neural network, wherein the first value network is a value network in a deep deterministic policy gradient neural network, and the first value function is a value function of a state-action pair, which includes the state and action of the hovercraft at the current moment; Calculate a first control parameter based on the second value network and the second strategy neural network, and control the operation of the hovercraft based on the first control parameter; Wherein, the determining of at least one sample data of the hovercraft based on the first strategy neural network, the expected attribute data and the actual attribute data comprises: Calculate error data of the hovercraft at a current moment based on the expected attribute data and the actual attribute data, and determine current initialization state data of the hovercraft based on the error data; Determining first motion data of the hovercraft based on the first strategy neural network, inputting the first motion data into a first control algorithm, and determining state data of the hovercraft at a next moment; A reward function of the hovercraft is constructed based on the error data, and first sample data is determined based on the reward function, the initialization state data, the next moment state data and the first action data.
2. The method according to claim 1, characterized in that The step of calculating error data of the hovercraft at a current moment based on the expected attribute data and the actual attribute data, and determining current initialization state data of the hovercraft based on the error data, comprises: Acquire actual attribute data at a previous moment, and calculate error change rate data of the hovercraft based on the actual attribute data at the previous moment and the current actual attribute data; calculating error vector data of the hovercraft based on the expected attribute data and the actual attribute data; The current initialization state data of the hovercraft is determined based on the error vector data and the error change rate data.
3. The method according to claim 1, characterized in that The updating of the first policy neural network and the first value network based on the first value function includes: Respectively obtaining initialization network parameters of the first strategy neural network and the first value network; Based on the first value function, the initialized network parameters are updated to obtain current network parameters corresponding to the first strategy neural network and the first value network; Determine value network parameters based on the initialization network parameters and the current network parameters, and update the first policy neural network and the first value network based on the value network parameters.
4. The method according to claim 2, characterized in that: The step of constructing a reward function of the hovercraft based on the error data comprises: determining an error reward function based on the error vector data, and determining an error change rate reward function based on the error change rate data; A first weight coefficient and a second weight coefficient are set, and a reward function of the hovercraft is constructed based on the first weight coefficient, the second weight coefficient, the error reward function, and the error change rate reward function.
5. A hovercraft lift control device, characterized in that: The device comprises: a data acquisition unit, configured to acquire current actual attribute data of the hovercraft, and set expected attribute data of the hovercraft based on the actual attribute data, wherein the actual attribute data includes heave height, heave speed and air chamber pressure; a sample data calculation unit, configured to determine at least one sample data of the hovercraft based on the first strategy neural network, the expected attribute data and the actual attribute data; a neural network updating unit, configured to input the at least one sample data into a first value network for training to obtain a first value function, and update the first policy neural network and the first value network based on the first value function to obtain a second value network and a second policy neural network, wherein the first value network is a value network in a deep deterministic policy gradient neural network, and the first value function is a value function of a state-action pair, which includes the state and action of the hovercraft at a current moment; a control parameter calculation unit, configured to calculate a first control parameter based on the second value network and the second strategy neural network, and control the operation of the hovercraft based on the first control parameter; The sample data calculation unit is also used to calculate the error data of the hovercraft at a current moment based on the expected attribute data and the actual attribute data, and determine the current initialization state data of the hovercraft based on the error data; determine the first action data of the hovercraft based on the first strategy neural network, input the first action data into a first control algorithm, and determine the state data of the hovercraft at a next moment; construct a reward function of the hovercraft based on the error data, and determine the first sample data based on the reward function, the initialization state data, the state data at the next moment and the first action data.
6. The device according to claim 5, characterized in that The neural network updating unit is also used to respectively obtain the initialization network parameters of the first strategy neural network and the first value network; update the initialization network parameters based on the first value function to obtain the current network parameters corresponding to the first strategy neural network and the first value network; determine the value network parameters based on the initialization network parameters and the current network parameters, and update the first strategy neural network and the first value network based on the value network parameters.
7. The device according to claim 6, characterized in that The sample data calculation unit is further used to obtain the actual attribute data at the previous moment, and calculate the error change rate data of the hovercraft based on the actual attribute data at the previous moment and the current actual attribute data; calculating error vector data of the hovercraft based on the expected attribute data and the actual attribute data; Determine current initialization state data of the hovercraft based on the error vector data and the error change rate data; determining an error reward function based on the error vector data, and determining an error change rate reward function based on the error change rate data; A first weight coefficient and a second weight coefficient are set, and a reward function of the hovercraft is constructed based on the first weight coefficient, the second weight coefficient, the error reward function, and the error change rate reward function.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Autonomous underwater robot model-free control method based on double-neural-network reinforcement learning technology
CN111240344A
Hovercraft path tracking method, device and equipment based on reinforcement learning and storage medium
CN119668271A