An underwater vehicle multi-energy hybrid energy storage system and coordinated energy management strategy

By introducing a multi-energy hybrid energy storage system and a coordinated energy management strategy into underwater vehicles, and utilizing a combination of fuel cells, lithium-ion batteries, and supercapacitors, the problem of slow dynamic response of fuel cells has been solved, achieving efficient dynamic response and improved endurance.

CN120165483BActive Publication Date: 2026-06-26NORTHWESTERN POLYTECHNICAL UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2025-03-05
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing underwater vehicles mainly rely on battery power. Fuel cells have relatively soft power characteristics and slow dynamic response, which leads to a decline in endurance and performance, making them unable to meet the requirements of complex missions.

Method used

A multi-energy hybrid energy storage system is adopted, including proton exchange membrane fuel cells, lithium-ion batteries and supercapacitors. By predicting the load through a BP neural network, the output of different energy sources is decoupled and coordinated energy management is achieved, and a composite topology is constructed.

Benefits of technology

It improves the dynamic response and endurance of underwater vehicles, and enables efficient and smooth switching of operating modes and optimal power distribution of hybrid energy storage systems, saving costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165483B_ABST
    Figure CN120165483B_ABST
Patent Text Reader

Abstract

The application particularly relates to a multi-energy hybrid energy storage system and a coordinated energy management strategy for an underwater vehicle. The system comprises a lithium ion battery, a super capacitor and a fuel cell using a proton exchange membrane, which are connected to a DC bus respectively. The method of the coordinated energy management strategy comprises: obtaining current motion parameters and current energy parameters of the underwater vehicle; inputting the current motion parameters, the current energy parameters and a desired speed into a load prediction model based on a BP neuron network to estimate a desired total power at a next time; determining a corresponding action strategy according to a current operation mode of the underwater vehicle, and performing power distribution on the desired total power based on the action strategy. The scheme can enable the hybrid energy system to perform smooth operation mode switching and optimal power distribution under different working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater vehicle control technology, specifically to a multi-energy hybrid energy storage system for underwater vehicles, and a coordinated energy management strategy applied to the multi-energy hybrid energy storage system for underwater vehicles. Background Technology

[0002] In related technologies, underwater vehicles are intelligent equipment capable of autonomous underwater navigation for extended periods and retrievable. Underwater vehicles can carry various sensors and specialized equipment to perform specific tasks. The power system is the "heart" of an underwater vehicle, typically composed of components such as a power source, an electric motor / engine, and a propulsion system. Currently, underwater vehicles are primarily powered by electricity, using batteries to drive the electric motor for propulsion. However, fuel cells, due to their relatively soft power characteristics and slow dynamic response, experience performance degradation, making them unsuitable for the application scenarios of underwater vehicles.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] This invention provides a multi-energy hybrid energy storage system for underwater vehicles and a coordinated energy management strategy, as well as a computer program product that can realize a composite topology structure with decoupled outputs of different energy sources, thereby overcoming the defects of the prior art to a certain extent.

[0005] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0006] According to a first aspect of the present invention, a multi-energy hybrid energy storage system for underwater vehicles is provided, the system comprising: a lithium-ion battery, a supercapacitor, and a fuel cell using a proton exchange membrane, all connected to a DC bus;

[0007] Fuel cells are used to provide average power for underwater vehicles;

[0008] Lithium-ion batteries are used to supplement power or absorb additional power.

[0009] Supercapacitors are used to smooth out power fluctuations at DC buses.

[0010] The DC bus is connected to the power components of the underwater vehicle.

[0011] In some exemplary embodiments, the fuel cell is connected to the DC bus via a unidirectional DC-DC power converter; the lithium-ion battery and the supercapacitor are respectively connected to the DC bus via bidirectional DC-DC power converters.

[0012] According to a second aspect of the present invention, a coordinated energy management strategy is provided, applied to the aforementioned multi-energy hybrid energy storage system of an underwater vehicle, the method for coordinating the energy management strategy comprising:

[0013] Acquire the current motion parameters and current energy parameters of the underwater vehicle;

[0014] The current motion parameters, current energy parameters, and desired speed are input into a load prediction model based on a BP neural network to estimate the desired total power at the next moment.

[0015] The corresponding action strategy is determined based on the current operating mode of the underwater vehicle, and the expected total power is allocated based on the action strategy.

[0016] In some exemplary embodiments, the current motion parameters include the current speed of the underwater vehicle and the current motor speed;

[0017] Current energy parameters include: current current.

[0018] In some exemplary embodiments, the current motion parameters, current energy parameters, and desired speed are input into a load prediction model based on a BP neural network to estimate the desired total power at the next moment, including:

[0019] Input the parameters into the input layer of the load prediction model;

[0020] In the load prediction model, each neuron in the intermediate hidden layer performs fully connected operations on the input parameters to obtain feature data.

[0021] The output layer of the load prediction model processes the feature data, weight matrix, and bias vector to output the expected total power at the next time step.

[0022] In some exemplary embodiments, the method further includes: pre-training a load prediction model based on a BP neural network, including:

[0023] Randomly initialize the model's weight matrix and bias vector;

[0024] Input samples, weight matrix, and bias vector are input into the input layer of the model, and the model's intermediate hidden layers are used to perform fully connected operations on the input samples to obtain multiple feature values.

[0025] The feature values ​​are nonlinearized by an activation function, and then fully connected with the weight matrix. The input values ​​are then output by the model's output layer.

[0026] The model loss is calculated based on the output value and the labels of the input samples, and the weight matrix and bias vector are updated using the model loss to complete the backpropagation training of the model.

[0027] In some exemplary embodiments, the method further includes:

[0028] A composite topology model corresponding to a multi-energy hybrid energy storage system is constructed; the composite topology model includes: fuel cell, lithium-ion battery, and supercapacitor;

[0029] Considering the activation loss, ohmic loss, and concentration loss of the fuel cell, a fuel cell model is defined based on the equivalent circuit form;

[0030] A lithium-ion battery model is defined based on a dynamic first-order RC equivalent circuit.

[0031] The supercapacitor model is defined based on the fractional-order model.

[0032] In some exemplary embodiments, the operating modes of the underwater vehicle include: acceleration mode, deceleration mode, and cruise mode.

[0033] In some exemplary embodiments, the method further includes:

[0034] Obtain the system state of the underwater vehicle in the underwater environment after executing the current action strategy, and calculate the reward score for the interaction between the current action strategy and the environment;

[0035] If the current reward score exceeds the preset constraints, an update action strategy will be triggered.

[0036] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, a method for implementing the above-described coordinated energy management strategy is provided.

[0037] The embodiments of this invention provide an underwater vehicle multi-energy hybrid energy storage system and coordinated energy management strategy. This system establishes a multi-energy hybrid energy storage system composed of a proton exchange membrane fuel cell, a lithium-ion battery, and a supercapacitor. It enables the use of lithium-ion batteries and supercapacitors to compensate for insufficient output power from the fuel cell, improving the efficiency and power quality of the energy storage system while saving costs. During the control of the hybrid energy storage system, by predicting the required power at the next moment, the operating mode and power distribution are switched and adjusted in advance, thereby improving the dynamic response capability of the hybrid energy storage system. This allows the hybrid energy system to smoothly switch operating modes and achieve optimal power distribution under different operating conditions.

[0038] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0040] Figure 1 This schematic diagram illustrates a multi-energy hybrid energy storage system structure for an underwater vehicle, as an exemplary embodiment of the present invention.

[0041] Figure 2 This schematic diagram illustrates an exemplary embodiment of the present invention of a method for a coordinated energy management strategy applied to a multi-energy hybrid energy storage system for underwater vehicles;

[0042] Figure 3 The diagram illustrates an equivalent circuit model of a fuel cell according to an exemplary embodiment of the present invention.

[0043] Figure 4 The diagram illustrates an equivalent circuit model of a lithium battery according to an exemplary embodiment of the present invention.

[0044] Figure 5 The diagram illustrates an equivalent circuit model of a supercapacitor according to an exemplary embodiment of the present invention.

[0045] Figure 6 The diagram illustrates a load prediction model architecture according to an exemplary embodiment of the present invention.

[0046] Figure 7 This diagram illustrates the flow of an iterative optimization method for an action policy network model, as exemplified by an embodiment of the present invention.

[0047] Figure 8 This diagram illustrates the training prediction results of a load prediction model according to an exemplary embodiment of the present invention. Detailed Implementation

[0048] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the invention will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0049] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0050] In related technologies, fuel cells directly convert the chemical energy stored in fuel and oxidant into electrical energy through electrode reactions. Because they do not involve a heat engine process and are not limited by the Carnot cycle, their energy conversion efficiency reaches up to 60%. A fuel cell is like a factory; as long as there is a continuous supply of fuel, it can continuously generate electricity. With increasing environmental awareness, users' demand for clean energy is growing stronger, and countries have begun to develop fuel cells that are easy to store hydrogen, have high durability, and are low-cost. Due to their advantages such as high efficiency, low pollution, and renewability, fuel cells are gradually gaining attention, and the required hydrogen and oxygen are renewable, thus their application prospects in various fields are very broad.

[0051] Underwater vehicles, primarily supported by submarines or surface ships, are intelligent devices capable of long-term autonomous underwater navigation and recovery. They can carry various sensors and specialized equipment to perform specific missions. The propulsion system is the "heart" of an underwater vehicle, typically composed of components such as energy source, motor / engine, and propulsion. Currently, electric propulsion is the main method, using batteries to drive the motor for propulsion. However, fuel cells suffer from relatively soft power characteristics and slow dynamic response, leading to performance degradation and limiting range and future development.

[0052] To address the shortcomings and deficiencies of existing technologies, this example embodiment provides a multi-energy hybrid energy storage system for underwater vehicles and a coordinated energy management strategy. The multi-energy structure can achieve complementary operation or independent operation, realizing a composite topology structure with decoupled output of different energy sources.

[0053] The following will describe in more detail the multi-energy hybrid energy storage system and coordinated energy management strategy for underwater vehicles in this exemplary embodiment, with reference to the accompanying drawings and embodiments.

[0054] This example embodiment provides a multi-energy hybrid energy storage system for underwater vehicles. (Reference) Figure 1 As shown, the system includes a lithium-ion battery 3, a supercapacitor 2, and a fuel cell 1, all connected to a DC bus 10. The fuel cell 1 may be a proton exchange membrane fuel cell (PEMFC) to provide average power to the underwater vehicle; the lithium-ion battery 3 is used for power replenishment or to absorb additional power; the supercapacitor 2 is used to smooth power fluctuations on the DC bus; the DC bus is connected to the underwater vehicle's power assembly. The fuel cell provides continuous and stable power output, while the lithium-ion battery and supercapacitor smooth the power output of the fuel cell. The power assembly includes a motor inverter 6, a thruster 7, and a propeller 8 connected in sequence. The motor inverter 6 is connected to the DC bus 10.

[0055] The fuel cell 1 is connected to the DC bus 10 via a unidirectional DC-DC converter 4; the lithium-ion battery 3 and the supercapacitor 2 are connected to the DC bus 10 via bidirectional DC-DC converters (5 and 9), respectively.

[0056] This example implementation provides a coordinated energy management strategy, which can be applied to, for example... Figure 1 The underwater vehicle shown has a multi-energy hybrid energy storage system. (Reference) Figure 2 As shown, methods for coordinating energy management strategies include:

[0057] Step S11: Obtain the current motion parameters and current energy parameters of the underwater vehicle;

[0058] Step S12: Input the current motion parameters, current energy parameters, and expected speed into the load prediction model based on BP neural network to estimate the expected total power at the next moment.

[0059] Step S13: Determine the corresponding action strategy based on the current operating mode of the underwater vehicle, and allocate the expected total power based on the action strategy.

[0060] The following will describe in more detail the method steps of the coordinated energy management strategy in this exemplary embodiment, with reference to the accompanying drawings and embodiments.

[0061] For example, the method further includes: constructing a composite topology model corresponding to a multi-energy hybrid energy storage system; the composite topology model includes: fuel cells, lithium-ion batteries, and supercapacitors;

[0062] Considering the activation loss, ohmic loss, and concentration loss of the fuel cell, a fuel cell model is defined based on the equivalent circuit form;

[0063] A lithium-ion battery model is defined based on a dynamic first-order RC equivalent circuit.

[0064] The supercapacitor model is defined based on the fractional-order model.

[0065] Specifically, refer to Figure 1 The multi-energy hybrid energy storage system shown can first construct a corresponding fuel cell-lithium-ion battery-supercapacitor composite topology. By creating a virtual mathematical model of the physical entity in a digital way, the behavior of the physical entity in the real environment is simulated with data. Through virtual-real interaction between the physical entity and the mathematical model, relevant parameters at each moment during underwater navigation are fed back.

[0066] For example, refer to Figure 3 As shown, the proton exchange membrane fuel cell model can be modeled using an equivalent circuit. Based on the operating current-voltage output characteristics, the fuel cell can be configured into three operating regions: activation loss, ohmic loss, and concentration loss, with corresponding equivalent resistance capacitors set. Considering the activation voltage loss and ohmic loss of the fuel cell while neglecting the concentration polarization voltage loss, the influence of operating parameters on the fuel cell can be characterized. That is, when parameters such as fuel and air pressure, temperature, composition, and flow rate change, it will affect the output voltage, Tafel slope (A), and exchange current (i). o ) and open-circuit voltage (E oc This affects the output voltage. The corresponding formula can be expressed as:

[0067]

[0068] in,

[0069] E oc =K c *E n

[0070]

[0071] Among them, K c It is the voltage constant under rated conditions; T d E represents the reaction time. n denoted as NN voltage; N is the number of fuel cell cells; T is the operating temperature of the fuel cell; z is the number of transferred electrons (equal to 2); F is the Faraday constant, with a value of 96485°C; R is the molar gas constant, with a value of 8.3145 J / (mol·K); k is the Boltzmann constant, with a value of 1.38 × 10-1. -23J / K; h is Planck's constant, with a value of 6.626 × 10⁻⁶. -34 Js; γ is the charge transfer coefficient; ΔG is the Gibbs free energy; and These are the supply pressures for hydrogen and oxygen, respectively. The partial pressure of water vapor; i fc and r ohm These are the fuel cell current and internal resistance, respectively.

[0072] For example, refer to Figure 4 As shown, a dynamic first-order RC equivalent circuit model of a lithium-ion battery can be defined, and sliding mode control can be used to track the real-time output voltage of the lithium-ion battery. According to Kirchhoff's voltage law, the corresponding formulas can include:

[0073] E O -V O -V C -V out =0

[0074] Among them, the open-circuit voltage of the lithium battery is represented by E. O The voltage across the ohmic internal resistance is represented by V. o It means, V c It is the voltage across the polarized capacitor and resistor, V out It is the real-time output voltage of the lithium battery.

[0075] The change in polarization capacitance can be obtained from the relationship between capacitance, current, and voltage:

[0076]

[0077] Based on the above formula, we can obtain:

[0078]

[0079] The real-time output voltage of a lithium battery can be calculated based on an equivalent circuit model. The operating current and output voltage can be obtained through testing with current and voltage sensors. Therefore, the measured operating current, current change I, and real-time output voltage of the lithium battery can be used as input variables for the model to estimate the internal parameters or state of the lithium battery.

[0080] Furthermore, the SOC of a lithium-ion battery can then be estimated using a Kalman filter algorithm, with the following formula:

[0081] X(k+1)=AX(k)+BI O (k)

[0082]

[0083] Where X(k) represents the estimated state variables of the battery system at time k, and C SOC (k) and U P (k) represent the state of charge and polarization capacitor voltage of the lithium battery, respectively. and P k Let K represent the prior and posterior estimation error covariance matrices at time k, respectively. k is the Kalman filter correction coefficient, Q is the process noise covariance matrix, and R is the measurement noise covariance matrix.

[0084] For example, a supercapacitor model can be constructed using a fractional-order model, and the dynamic behavior of the supercapacitor can be described using fractional-order components. (Reference) Figure 5 The figure shows a representative fractional-order model of a supercapacitor, consisting of a constant-phase CPE element, a Warburg-like element, a parallel resistor Rp, and a series resistor Rs. The Warburg-like element characterizes the main capacitance properties and describes the semi-wireless diffusion impedance characteristics of the planar electrodes caused by the ion diffusion effect inside the supercapacitor. The impedance expression for the Warburg-like element is:

[0085]

[0086] Where W represents the main capacitance coefficient of the supercapacitor, and β represents the end of capacitance distribution. When β = 0, the Warburg-like element is a resistor; when β equals 1, the Warburg-like element is an ideal capacitor.

[0087] The impedance expression for a constant-phase CPE element is:

[0088]

[0089] Where C1 represents the capacitance coefficient of the constant-phase CPE element, and α represents the fractional order. When α = 1, the CPE element represents an ideal capacitor.

[0090] The series resistor Rs is used to measure the resistance of the electrolyte and current collector in an equivalent supercapacitor.

[0091] The impedance expression for a fractional-order supercapacitor model is:

[0092]

[0093] Among them, u sc Let represent the terminal voltage of the supercapacitor, and ... P R s W].

[0094] A fractional-order SOC estimation method is used to track the SOC of supercapacitors in real time, achieving real-time monitoring and estimation of the supercapacitor's SOC. The formula is as follows:

[0095]

[0096] For example, the model of an underwater vehicle is defined as follows:

[0097] N = F 水

[0098]

[0099] Among them, P m Yes, the drive motor requires an input power of N. m For the drive motor torque, θ m For the efficiency of the drive motor, ω m Let F be the speed of the drive motor, N be the thrust generated by the motor, and L be the structural length. 水 For navigation resistance.

[0100] For example, the operating modes of an underwater vehicle include: acceleration mode, deceleration mode, and cruise mode.

[0101] Specifically, the load conditions of the underwater vehicle can be predefined. Considering various situations the underwater vehicle will encounter during navigation, the operating modes can be divided into acceleration mode, deceleration mode, and cruise mode. Factors affecting its operating mode include ocean current speed, current direction, current speed, desired speed, and motor speed. During high-power acceleration, reconnaissance, and special circumstances, the load changes significantly; during low-speed, low-power navigation and routine positioning and navigation, the load changes little, and the power output is relatively stable.

[0102] For example, pre-training a load prediction model based on a BP neural network includes:

[0103] Step S21: Randomly initialize the model's weight matrix and bias vector;

[0104] Step S22: Input the input samples, weight matrix, and bias vector into the input layer of the model, and use the intermediate hidden layer of the model to perform a fully connected operation on the input samples to obtain multiple feature values;

[0105] Step S23: The feature values ​​are nonlinearly processed by the activation function, and a fully connected operation is performed with the weight matrix. The input values ​​are then output by the output layer of the model.

[0106] Step S24: Calculate the model loss based on the output value and the labels of the input samples, and use the model loss to update the weight matrix and bias vector to complete the backpropagation training of the model.

[0107] Specifically, for the aforementioned hybrid energy system topology, a corresponding load prediction model can be pre-trained to predict the load of the underwater vehicle. Specifically, a load prediction model based on a backpropagation (BP) neural network can be constructed. The input data for the load prediction model can include: current current, current rotational speed, current speed, and desired speed; these four influencing factors serve as the input layer of the neural network, with the required power at the next predicted moment serving as the output layer. Factors such as temperature, wind speed, and ocean currents are ignored. The intermediate hidden layer has 15 neurons, and 15 data features are extracted from the four input variables through fully connected operations. Finally, an output result is obtained through a fully connected layer.

[0108] In the training process of the load prediction model, X1, X2, X3, and X4 can be predefined as external inputs, and W... ij To randomly initialize the model's weights, the bias vector is... Input value and weight W ij A fully connected operation yields 15 feature values. These are then nonlinearized using the Sigmoid activation function and weighted by V. j Perform a fully connected operation, V j The weights connecting the j-th hidden layer neuron to the output layer O are also nonlinearized using the Sigmoid function.

[0109] The calculation formula from the input layer to the hidden layer includes:

[0110]

[0111] Among them, X i Given a 1x4 input vector, W ij It is a 4*15 weight matrix. is a 1*15 weighted bias parameter vector. f is the Sigmoid activation function.

[0112] The formula for the calculation process from the hidden layer to the output layer is:

[0113] y = f(H i *V j -β o )

[0114] Among them, V j It is a 15*4 weight matrix, β o Let be the bias vector of 'o'.

[0115] The model loss is calculated using a loss function based on the model's output values ​​and labels. The loss value is then applied to the weight matrix W. ij and V j and bias vector and β o The model is then updated to perform backpropagation training. The loss function formula includes:

[0116]

[0117] Where Loss is the loss function, P m Let be the actual total load power value, and let O be the total load power predicted by the load prediction model at the next time step.

[0118] refer to Figure 8 As shown, based on the comparison between the predicted motor power value and the true motor power value, the load prediction model provided by this invention has a relatively accurate prediction effect.

[0119] For example, action policy network models corresponding to each operating mode can also be pre-built.

[0120] Specifically, for cruise mode, a motion strategy network model of the underwater vehicle under cruise mode conditions can be constructed. First, the loss function for long-duration cruise mode is constructed. Long-duration cruise mode refers to the operating state when performing long-distance, long-cycle cruise missions; it requires the energy system to achieve energy self-sufficiency through a multi-energy hybrid energy storage system, primarily to improve the vehicle's endurance, with less emphasis on maneuverability. Therefore, long-duration cruise mode focuses on the state of charge (SOC) value of the multi-energy hybrid energy storage system. Thus, the expression for the loss function under long-duration cruise mode is:

[0121] LOSS1 = -MAX(L1)

[0122] Where L1 = v*t; LOSS1 represents the loss function value corresponding to the long-term cruise mode. MAX(L1) represents the farthest travel distance in the long-term cruise mode, and v represents the current speed of the aircraft.

[0123] Based on the loss function corresponding to the long-term cruise mode, the initial action policy network model corresponding to the long-term cruise mode can be constructed.

[0124] The initial action policy network model for long-duration cruise mode is trained to obtain the corresponding action policy network model for long-duration cruise mode. For example, the action policy network model can also be a backpropagation (BP) neural network model. Specifically, the state of charge values ​​of the fuel cell, lithium battery, and supercapacitor, the speed, and the actions taken at the current moment in long-duration cruise mode can be used as inputs to the initial action policy network model for long-duration cruise mode. The actions to be taken at the next moment in long-duration cruise mode can be used as outputs to train the initial action policy network model for long-duration cruise mode, resulting in a trained action policy network model for long-duration cruise mode. The aforementioned loss function is then used to perform backpropagation training on the model.

[0125] The actions taken by the underwater vehicle in cruise mode can be expressed by the following formula:

[0126]

[0127] in, S111 represents the action taken during cruise mode operation; S112 represents the action value when the fuel cell generates electricity during long-term self-sustaining mode operation; S112 represents the action value when the fuel cell power generation module is turned off during long-term cruise mode operation; S121 represents the action value when the lithium battery power generation module generates electricity during long-term cruise mode operation; S122 represents the action value when the lithium battery power generation module is turned off during long-term cruise mode operation; S131 represents the action value when the supercapacitor module generates electricity during long-term cruise mode operation.

[0128] For example, an action policy network model corresponding to the acceleration mode can also be constructed. Specifically, a loss function for the underwater vehicle in acceleration mode can be constructed first, considering that acceleration mode requires the underwater vehicle to have a fast dynamic response capability, and the hybrid energy storage system must provide all the power required by the load in a short time (i.e., the predicted power output by the load prediction model); while the requirement for endurance is not high, the expression of the loss function corresponding to acceleration mode can be configured as follows:

[0129] LOSS2 = ∑[(P1 + P2 + P3) - O] 2

[0130] Where LOSS2 represents the loss function value corresponding to the acceleration mode; O represents the total load power of the underwater vehicle predicted by the load prediction model at the next moment; (P1+P2+P3) represents the total output power of the underwater vehicle at the next moment under the acceleration mode; P1 represents the output power of the fuel cell; P2 represents the output power of the lithium battery; and P3 represents the output power of the supercapacitor.

[0131] Based on the loss function corresponding to the acceleration mode, an initial action strategy network model for the underwater vehicle under acceleration mode can be constructed. This initial action strategy network model is then trained to obtain the actual action strategy network model for the underwater vehicle under acceleration mode. Specifically, the state of charge values ​​of the fuel cell, lithium battery, and supercapacitor, the vehicle's speed, and the action taken by the underwater vehicle at the current moment can be used as inputs to the initial action strategy network model under acceleration mode. The action taken by the underwater vehicle at the next moment under acceleration mode is used as the output. This initial action strategy network model under acceleration mode is then trained to obtain the trained action strategy network model for acceleration mode. The expression for the action taken by the underwater vehicle under acceleration mode is:

[0132]

[0133] in, S211 represents the action taken in acceleration mode; S212 represents the action value when the fuel cell generates electricity in acceleration mode; S221 represents the action value when the fuel cell power generation module is turned off in acceleration mode; S222 represents the action value when the lithium battery power generation module generates electricity in acceleration mode; S222 represents the action value when the lithium battery power generation module is turned off in acceleration mode; S231 represents the action value when the supercapacitor module generates electricity in acceleration mode; S232 represents the action value when the supercapacitor module is turned off in acceleration mode.

[0134] For example, for deceleration mode, a corresponding action policy network model can be constructed. First, the loss function for the underwater vehicle in deceleration mode can be constructed. This can be done by considering that the maneuverability requirements are not high in deceleration mode; the expression for the loss function in deceleration mode is as follows:

[0135] LOSS3 = -MAX(L3)

[0136]

[0137] Where LOSS3 represents the loss function value under the deceleration mode, represents the farthest travel distance of the underwater vehicle under the deceleration mode, and 'a' represents the current acceleration of the vehicle.

[0138] Based on the loss function corresponding to the deceleration mode, an initial action strategy network model for the underwater vehicle under deceleration mode can be constructed. This initial action strategy network model is then trained to obtain the actual action strategy network model for the underwater vehicle under deceleration mode. Specifically, the state of charge values ​​of the fuel cell, lithium battery, and supercapacitor, the vehicle's speed, and the action taken by the underwater vehicle at the current moment can be used as inputs to the initial action strategy network model under deceleration mode. The action taken by the underwater vehicle at the next moment under deceleration mode can be used as the output. This initial action strategy network model under deceleration mode is then trained to obtain the trained action strategy network model for deceleration mode.

[0139] The expression for the actions taken by the underwater vehicle in deceleration mode is as follows:

[0140]

[0141] in, S311 represents the action taken by the underwater vehicle in deceleration mode; S312 represents the action value when the fuel cell generates electricity in deceleration mode; S321 represents the action value when the fuel cell shuts down in deceleration mode; S322 represents the action value when the lithium battery generates electricity in deceleration mode; S331 represents the action value when the supercapacitor generates electricity in deceleration mode; S332 represents the action value when the supercapacitor shuts down in deceleration mode.

[0142] For example, after selecting the action to be performed by the underwater vehicle through the action policy network model, it can interact with the environment in which the underwater vehicle is located. For instance, the underwater vehicle model can be used to calculate the state change of the underwater vehicle after taking a certain action; appropriate observation variables can be selected according to the task characteristics of the underwater vehicle under different modal conditions, and the score of the action can be calculated based on the variable and the reward evaluation system; if the state change of the underwater vehicle after taking the action does not meet the ideal situation or exceeds the constraints, the score is deducted.

[0143] Action selection in an action policy network model is an ongoing optimization process. It can continuously search for the best action through a trial-and-error mechanism and improve its behavioral patterns to obtain the maximum cumulative reward. (Reference) Figure 7 Reinforcement learning algorithms can be used to find the optimal action (the policy with the highest score). The initial action can be randomly initialized; since it has no learning value, the evaluation of the initial action is 0. The action policy network model for each modal condition is optimized using reinforcement learning algorithms to obtain the target action policy network model. Specifically, when optimizing the action policy network model for the long-cruise mode, the expression for the reward function of the reinforcement learning algorithm is:

[0144]

[0145] Where G1 represents the reward function of the reinforcement learning algorithm when optimizing the action policy network model under the long cruise mode; η represents the discount factor, η 0 Let η represent the power of the discount factor, where η ∈ [0, 1]. The value represents the reward obtained by the underwater vehicle after taking actions and interacting with the environment at time t in the long-duration cruise mode; k represents the kth time. This represents the reward value obtained by the underwater vehicle after taking actions and interacting with the environment at time t+k in the long cruise mode.

[0146] In the long-cruise mode of the underwater vehicle, the expression for the reward value after the actions taken and the interaction with the environment is:

[0147]

[0148] in, This represents the reward value obtained by the underwater vehicle after taking actions and interacting with the environment at the system state at time 1 in long cruise mode. This indicates the trend of change in the travel distance of an underwater vehicle under long-cruise mode. This refers to the bonus items for underwater vehicles operating under long-term self-sustaining modal conditions. This represents the detailed penalty term for the underwater vehicle in long cruise mode; SOC represents the state of charge value of the battery module; in order to ensure that the underwater vehicle's action strategy network is continuously updated and avoid getting trapped in local optima, the case where the underwater vehicle's travel distance remains unchanged in long cruise mode is set as a deduction term; at the same time, a detailed penalty term for the state of charge value (SOC) is added.

[0149] For example, when optimizing the action policy network model in accelerated mode, the expression for the reward function of the reinforcement learning algorithm is:

[0150]

[0151] Where G2 represents the reward function of the reinforcement learning algorithm when optimizing the action policy network model in accelerated mode; η represents the discount factor, η 0 Let η represent the power of the discount factor, where η ∈ [0, 1]. The value represents the reward obtained by the underwater vehicle after taking actions and interacting with the environment at time t in the accelerated mode; k represents the kth time. This represents the reward value obtained by the underwater vehicle after taking actions and interacting with the environment at time t+k in the accelerated mode.

[0152] For example, in the acceleration mode of an underwater vehicle, the expression for the reward value after the action taken and the interaction with the environment is:

[0153]

[0154] in, This represents the reward value obtained by the underwater vehicle after taking actions and interacting with the environment at the system state at time 1 in the acceleration mode. This indicates the changing trend of the loss function of the action strategy network model corresponding to the underwater vehicle in acceleration mode.

[0155] For example, when optimizing the action policy network model under deceleration mode conditions, the expression for the reward function of the reinforcement learning algorithm is as follows:

[0156]

[0157] Where G3 represents the reward function of the reinforcement learning algorithm when optimizing the action policy network model under deceleration mode conditions; η represents the discount factor, η 0 Let η represent the power of the discount factor, where η ∈ [0, 1]. The value represents the reward obtained by the underwater vehicle after taking actions and interacting with the environment at time t in the deceleration mode; k represents the kth time. This represents the reward value obtained by the underwater vehicle after taking actions and interacting with the environment at time t+k in the deceleration mode.

[0158] The expression for the reward value of an underwater vehicle after its actions and environmental interactions during deceleration mode is as follows:

[0159]

[0160] in, This represents the reward value obtained by the underwater vehicle after taking actions and interacting with the environment at the system state at time 1 in deceleration mode. This indicates the trend of the underwater vehicle's travel distance under deceleration mode. This indicates the bonus items for underwater vehicles operating in deceleration mode. This indicates the detailed penalty items for the underwater vehicle in deceleration mode; SOC indicates the state of charge value of the battery pack module.

[0161] Based on the above, we can obtain the target action policy network model for each modal condition after optimizing the action policy network model for each modal condition using reinforcement learning algorithms.

[0162] In step S11, the current motion parameters and current energy parameters of the underwater vehicle are obtained.

[0163] For example, before operation, an underwater vehicle can first configure parameters such as destination coordinates, average speed, and depth corresponding to the current navigation mission. Furthermore, the operating mode can also be configured. For instance, if a cruise mission is currently being performed, it can be configured to initially be in cruise mode; or, if a pursuit mission is currently underway, it can be configured to initially be in acceleration mode. During operation, the underwater vehicle can collect its current motion parameters and current energy parameters in real time. Additionally, during the operation of the underwater instrument, the operating mode can be adjusted in real time when adjusting the operating speed. For example, if the underwater vehicle needs to decelerate to avoid an obstacle or turn during the current phase of operation, it can be configured to decelerate mode; or, if the underwater vehicle needs to increase its speed during the current phase of operation, it can be configured to accelerate mode accordingly.

[0164] For example, current motion parameters include the underwater vehicle's current speed and current motor speed. For instance, the motor speed could refer to the propeller or the motor speed of a propeller thruster. Current energy parameters include current parameters. This current parameter could be the current output from the DC bus.

[0165] In step S12, the current motion parameters, current energy parameters, and expected speed are input into the load prediction model based on a BP neural network to estimate the expected total power at the next moment.

[0166] For example, the current motion parameters, current energy parameters, and desired speed are input into a load prediction model based on a BP neural network to estimate the desired total power at the next moment, including:

[0167] Step S121: Input each parameter into the input layer of the load prediction model;

[0168] Step S122: Each neuron in the intermediate hidden layer of the load prediction model performs a fully connected operation on each input parameter to obtain feature data.

[0169] Step S123: The output layer of the load prediction model processes the feature data, weight matrix, and bias vector, and outputs the expected total power at the next time step.

[0170] Specifically, the current speed, current motor speed, current current parameters, and desired speed can be used as input parameters for the model. The desired speed can be the speed expected to be achieved during the current navigation task in the current operating mode. These parameters are then input into the trained load prediction model, and the expected total power for the next time step is obtained from the model's output.

[0171] In step S13, a corresponding action strategy is determined based on the current operating mode of the underwater vehicle, and the expected total power is allocated based on the action strategy.

[0172] For example, when collecting the motion parameters of the underwater vehicle at the current moment, the operating mode for the current stage can also be determined. After obtaining the expected total power for the next moment from the model output, the expected total power can be allocated according to the action strategy corresponding to the current operating mode, determining the power supply that needs to be powered at the current moment. Furthermore, the power allocation ratio of the power supply can be configured. Alternatively, the operating mode for the current moment can be configured based on the current speed and the expected speed for the next moment. For example, if the expected speed for the next moment is greater than the current speed, i.e., an acceleration state is required, it can be configured as an acceleration mode; or, if the expected speed for the next moment is equal to the current speed, it can be configured as a cruise mode; or, if the expected speed for the next moment is less than the current speed, it can be configured as a deceleration mode.

[0173] For example, the method further includes: obtaining the system state of the underwater vehicle in the underwater environment after executing the current action strategy, and calculating the reward score of the interaction between the current action strategy and the environment;

[0174] If the current reward score exceeds the preset constraints, an update action strategy will be triggered.

[0175] Specifically, the aforementioned action strategy can be the power allocation strategy executed at the current moment. After executing the power allocation strategy with the desired total power, the system status of the underwater vehicle at the next moment can be collected; the system status can include: speed, motor speed, position, DC bus current, and operating status parameters of each power source, etc.

[0176] Based on the current operating mode, the reward score for the interaction between the current action policy and the environment can be recalculated using methods such as those described in the above embodiments. If the score exceeds the preset constraints, an update to the action policy and the target action policy network model can be triggered.

[0177] For example, during the operation of an underwater vehicle, within a preset statistical period, under a certain operating mode, such as acceleration mode, if the range of speed increase is less than the expected evaluation value for one minute or 15 seconds consecutively, this can trigger the optimization of the target action policy network model, updating the action policy network model again. The updated model can then be used to reallocate power.

[0178] The system and method provided in this invention establish a multi-energy hybrid energy storage system composed of a proton exchange membrane fuel cell, a lithium-ion battery, and a supercapacitor. This system uses renewable energy and is green, environmentally friendly, and pollution-free. The lithium-ion battery and supercapacitor compensate for the insufficient output power of the fuel cell. This improves the efficiency and power quality of the energy storage system and saves costs. By establishing a load prediction model based on a neural network, the required power at the next moment is predicted, and the operating mode and power allocation are switched and adjusted in advance, improving the dynamic response capability of the fuel cell-lithium-ion battery-supercapacitor multi-energy storage system. This method can switch operating modes and allocate power according to different operating conditions encountered during navigation, and optimize the power output range through a differential optimization algorithm. This enables the hybrid energy system to smoothly switch operating modes and achieve optimal power allocation under different operating conditions. The differential optimization algorithm can autonomously determine the optimal power allocation strategy of the hybrid energy system without manual setting, thus reducing design costs and error probability compared to traditional logical strategies.

[0179] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may, for example, be executed synchronously or asynchronously in multiple modules.

[0180] It should be noted that although several modules or units of the device for performing actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0181] Embodiments of the present invention include a computer program product comprising a computer program carried on a storage medium, the computer program containing program code for performing the methods shown in the flowchart.

[0182] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein computer-readable program code is carried. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0183] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0184] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0185] In one embodiment, this application provides a computer program product including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0186] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0187] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.

[0188] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A coordinated energy management strategy, characterized in that, A multi-energy hybrid energy storage system for underwater vehicles is proposed, comprising: a lithium-ion battery, a supercapacitor, and a fuel cell using a proton exchange membrane, all connected to a DC bus; the fuel cell provides average power to the underwater vehicle; the lithium-ion battery supplements power or absorbs additional power to smooth the power of the fuel cell; the supercapacitor smooths power fluctuations on the DC bus to achieve power smoothing for the fuel cell; wherein the DC bus is connected to the power components of the underwater vehicle. Methods for coordinating energy management strategies include: Acquire the current motion parameters and current energy parameters of the underwater vehicle; the current motion parameters include the current speed of the underwater vehicle at the current moment and the current motor speed; the current energy parameters include: the current current. The current motion parameters, current energy parameters, and desired speed are input into a load prediction model based on a BP neural network to estimate the desired total power at the next moment. The corresponding action strategy is determined based on the current operating mode of the underwater vehicle, and the expected total power is allocated based on the action strategy; wherein, the operating modes of the underwater vehicle include: acceleration mode, deceleration mode, and cruise mode.

2. The coordinated energy management strategy according to claim 1, characterized in that, The current motion parameters, current energy parameters, and desired speed are input into a load prediction model based on a backpropagation neural network to estimate the desired total power at the next moment, including: Input the parameters into the input layer of the load prediction model; In the load prediction model, each neuron in the intermediate hidden layer performs fully connected operations on the input parameters to obtain feature data. The output layer of the load prediction model processes the feature data, weight matrix, and bias vector to output the expected total power at the next time step.

3. The coordinated energy management strategy according to claim 2, characterized in that, The method further includes: pre-training a load prediction model based on a BP neural network, including: Randomly initialize the model's weight matrix and bias vector; Input samples, weight matrix, and bias vector are input into the input layer of the model, and the model's intermediate hidden layers are used to perform fully connected operations on the input samples to obtain multiple feature values. The feature values ​​are nonlinearized by an activation function, and then fully connected with the weight matrix. The input values ​​are then output by the model's output layer. The model loss is calculated based on the output value and the labels of the input samples, and the weight matrix and bias vector are updated using the model loss to complete the backpropagation training of the model.

4. The coordinated energy management strategy according to claim 1, characterized in that, The method further includes: A composite topology model corresponding to a multi-energy hybrid energy storage system is constructed; the composite topology model includes: fuel cell, lithium-ion battery, and supercapacitor; Considering the activation loss, ohmic loss, and concentration loss of the fuel cell, a fuel cell model is defined based on the equivalent circuit form; where, according to the operating current... Voltage output characteristics are used to configure the fuel cell into three operating regions: activation loss, ohmic loss, and concentration loss, and the corresponding equivalent resistance capacitor is set. A lithium-ion battery model is defined based on a dynamic first-order RC equivalent circuit. The supercapacitor model is defined based on a fractional-order model, and fractional-order elements are used to describe the dynamic behavior of the supercapacitor.

5. The coordinated energy management strategy according to claim 1, characterized in that, The method further includes: Obtain the system state of the underwater vehicle in the underwater environment after executing the current action strategy, and calculate the reward score for the interaction between the current action strategy and the environment; If the current reward score exceeds the preset constraints, an update action strategy will be triggered.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the coordinated energy management strategy as described in any one of claims 1 to 5.