Hybrid vessel energy management method and system powered by dual three-phase generators

CN122801199APending Publication Date: 2026-09-22HUANGGANG POLYTECHNIC COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610887635.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]本发明实施例提供一种双三相发电机供电的混合动力船舶能量管理方法、系统、电子设备和存储介质,以解决现有技术中面对DTP-PMSG供电的混合动力船舶时,存在无法兼顾复杂工况适应性、燃油经济性、储能设备寿命与系统运行稳定性的多目标协同优化难题,且缺乏客观、自适应的权重决策机制的技术问题

Benefits of technology

(1)实现了燃油经济性与储能设备寿命的协同优化,既能保证短期节能效益,又能有效延缓蓄电池和超级电容的老化,降低全生命周期运维成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801199A_ABST
    Figure CN122801199A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a kind of dual three-phase generator power supply's hybrid power ship energy management method, system, electronic equipment and storage medium, by constructing the power system model including dual three-phase permanent magnet synchronous generator, battery, super capacitor and ship load, and specially establishing the health state attenuation model of battery and super capacitor, provide accurate quantitative basis for subsequent optimization;Multi-objective optimization model is constructed to cover seven dimensions such as fuel consumption, energy storage life, voltage stability, power smoothness and state of charge deviation, the Pareto optimal solution set is solved using multi-objective evolutionary algorithm, and the weight of each target is objectively deduced by multi-objective decision method, which completely eliminates the drawbacks of subjective weighting in traditional method;Based on deep deterministic policy gradient algorithm, an energy management strategy model is constructed, the foregoing objective weight is integrated into the reward function, and the optimal control strategy is obtained by iterative training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine new energy hybrid power system control technology, and in particular to a method, system, electronic equipment and storage medium for energy management of hybrid ships powered by dual three-phase generators. Background Technology

[0002] Currently, with the green transformation of the shipping industry, dual three-phase permanent magnet synchronous generators (DTP-PMSGs) are gradually becoming the core power supply equipment for hybrid-powered ships due to their high power density and high fault tolerance. For energy management of such ships, existing technologies mainly fall into three categories. The first category is rule-based strategies, such as power-following control and fuzzy control, which rely on pre-set engineering experience rules for power allocation. The second category is traditional optimization strategies, such as model predictive control (MPC) and dynamic programming (DP), which solve for the optimal power allocation scheme by establishing a mathematical model of the system. The third category is reinforcement learning-based strategies that have emerged in recent years, such as the Deep Deterministic Policy Gradient (DDPG) algorithm. This algorithm learns the optimal control strategy in a continuous action space through model-free self-learning to cope with the complexity of ship navigation conditions.

[0003] However, the aforementioned existing technologies still have significant limitations when applied to hybrid-powered ships powered by DTP-PMSG. First, rule-based strategies lack adaptability and struggle to cope with frequent strong disturbances such as step loads and pulse loads during ship navigation, easily leading to unreasonable power distribution. Traditional optimization strategies, on the other hand, heavily rely on accurate system mathematical models, but ship propulsion systems exhibit strong nonlinearity and time-varying parameters. Model mismatch will result in decreased control effectiveness, and their high algorithm complexity makes it difficult to meet real-time control requirements. Second, existing strategies often prioritize fuel economy as the sole optimization objective, completely ignoring the State of Health (SOH) degradation problem of hybrid energy storage systems (such as batteries and supercapacitors). This leads to accelerated aging of energy storage devices due to prolonged exposure to high current and deep charge / discharge conditions, significantly increasing the ship's total lifecycle maintenance costs. Secondly, the few strategies that attempt to consider multiple objectives generally employ subjective methods such as the Analytic Hierarchy Process (AHP) or expert scoring to determine the weights of these objectives. This results in highly subjective and arbitrary decision-making, failing to objectively balance multiple interdependent objectives such as fuel consumption, energy storage life, and system stability. Consequently, the final energy management solution is difficult to align with the actual operational needs of ships. Finally, existing reinforcement learning strategies based on DDPG have not been specifically optimized for ship hybrid power systems. Their action space constraints, experience playback mechanisms, and convergence criteria all follow conventional settings, lacking effective modeling and penalty for energy storage SOH decay. This leads to poor algorithm training convergence stability and insufficient control precision.

[0004] In summary, existing technologies face the challenge of achieving multi-objective collaborative optimization when dealing with hybrid ships powered by DTP-PMSG, as they cannot simultaneously consider adaptability to complex operating conditions, fuel economy, energy storage device lifespan, and system operational stability. Furthermore, they lack an objective and adaptive weighted decision-making mechanism. Summary of the Invention

[0005] This invention provides an energy management method, system, electronic device, and storage medium for hybrid ships powered by dual three-phase generators, to solve the technical problem in the prior art that when dealing with hybrid ships powered by DTP-PMSG, it is impossible to simultaneously consider the multi-objective collaborative optimization problem of adaptability to complex operating conditions, fuel economy, energy storage device lifespan, and system operation stability, and that there is a lack of objective and adaptive weight decision-making mechanism.

[0006] In a first aspect, embodiments of the present invention provide an energy management method for a hybrid-powered ship powered by dual three-phase generators, comprising: S1. Construct a hybrid power ship propulsion system model, which includes a dual three-phase permanent magnet synchronous generator model, a battery model, a supercapacitor model, and a ship load model; and construct a battery health state decay model and a supercapacitor health state decay model, which are used to quantify the battery health state decay during operation and the supercapacitor health state decay during operation, respectively. S2. Construct a multi-objective optimization model. The optimization objectives of the multi-objective optimization model include: fuel consumption, cumulative health state decay of the battery, cumulative health state decay of the supercapacitor, DC bus voltage stability index, generator power stability index, battery state of charge deviation index, and supercapacitor state of charge deviation index. Solve the multi-objective optimization model using a multi-objective evolutionary algorithm to obtain a set of Pareto optimal solutions. Use a multi-objective decision-making method to select the solution with the highest proximity from the Pareto optimal solution set as the optimal compromise solution, and deduce the objective weights of each optimization objective based on the optimal compromise solution. S3. Construct an energy management strategy model based on a deep deterministic policy gradient algorithm. The state space of the energy management strategy model includes ship navigation state parameters, energy storage device state parameters, and power system operating parameters. The action space of the energy management strategy model includes generator power adjustment, battery power adjustment, and supercapacitor power adjustment. Construct a multi-objective reward function based on the objective weights of each optimization objective. Iteratively train the energy management strategy model using the multi-objective reward function until the preset convergence condition is met to obtain the trained energy management strategy model. S4. During the actual navigation of the ship, the current state variables are collected in real time, input into the trained energy management strategy model, and output real-time power adjustment commands for the generator, the battery, and the supercapacitor to perform online control of the power distribution of the hybrid power ship propulsion system.

[0007] In a second aspect, embodiments of the present invention provide a hybrid power ship energy management system powered by dual three-phase generators, comprising: The system modeling module is used to construct a hybrid power ship propulsion system model, which includes a dual three-phase permanent magnet synchronous generator model, a battery model, a supercapacitor model, and a ship load model; and to construct a battery health state decay model and a supercapacitor health state decay model, which are used to quantify the battery health state decay during operation and the supercapacitor health state decay during operation, respectively. A multi-objective optimization module is used to construct a multi-objective optimization model. The optimization objectives of the multi-objective optimization model include: fuel consumption, cumulative health state degradation of the battery, cumulative health state degradation of the supercapacitor, DC bus voltage stability index, generator power stability index, battery state of charge deviation index, and supercapacitor state of charge deviation index. A multi-objective evolutionary algorithm is used to solve the multi-objective optimization model to obtain a set of Pareto optimal solutions. A multi-objective decision-making method is used to select the solution with the highest closeness from the Pareto optimal solution set as the optimal compromise solution, and the objective weights of each optimization objective are derived from the optimal compromise solution. The model training module is used to construct an energy management strategy model based on a deep deterministic policy gradient algorithm. The state space of the energy management strategy model includes ship navigation state parameters, energy storage device state parameters, and power system operating parameters. The action space of the energy management strategy model includes generator power adjustment, battery power adjustment, and supercapacitor power adjustment. A multi-objective reward function is constructed based on the objective weights of each optimization objective. The energy management strategy model is iteratively trained using the multi-objective reward function until a preset convergence condition is met, thereby obtaining a trained energy management strategy model. The online control module is used to collect current state variables in real time during the actual navigation of the ship, input them into the trained energy management strategy model, and output real-time power adjustment commands for the generator, the battery, and the supercapacitor, so as to perform online control of the power distribution of the hybrid power ship propulsion system.

[0008] Thirdly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the hybrid power ship energy management method powered by dual three-phase generators as described in the first aspect of the present invention.

[0009] Fourthly, embodiments of the present invention provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the hybrid power ship energy management method powered by dual three-phase generators as described in the first aspect of the present invention.

[0010] This invention provides an energy management method, system, electronic device, and storage medium for hybrid-powered ships powered by dual three-phase generators. First, a power system model is constructed, including dual three-phase permanent magnet synchronous generators, batteries, supercapacitors, and ship loads. A health state decay model for the batteries and supercapacitors is specifically established to provide a precise quantitative basis for subsequent optimization. Second, a multi-objective optimization model is constructed, encompassing seven dimensions: fuel consumption, energy storage life, voltage stability, power stability, and state of charge deviation. A Pareto optimal solution set is obtained using a multi-objective evolutionary algorithm, and the weights of each objective are objectively derived through a multi-objective decision-making method, completely eliminating the drawbacks of subjective weighting in traditional methods. Then, an energy management strategy model is constructed based on a deep deterministic strategy gradient algorithm, incorporating the aforementioned objective weights into the reward function, and obtaining the optimal control strategy through iterative training. Finally, state variables are collected in real-time during actual ship navigation and input into the trained model, outputting power adjustment commands for each power source online to achieve closed-loop energy allocation. Compared with existing technologies, this invention has the following advantages: (1) It achieves synergistic optimization of fuel economy and energy storage equipment lifespan, which can not only ensure short-term energy saving benefits, but also effectively delay the aging of batteries and supercapacitors and reduce the operation and maintenance costs throughout the entire life cycle.

[0011] (2) By using data-driven objective weight allocation, the multi-objective decision-making results are more in line with the actual navigation conditions of the ship, avoiding power allocation mismatch caused by human experience bias.

[0012] (3) The action space constraints and reward function of the deep deterministic strategy gradient algorithm were optimized for the characteristics of the ship power system, which significantly improved the convergence stability and control accuracy of the algorithm training and enabled it to respond quickly to complex disturbances such as step load and pulse load.

[0013] (4) It fully adapts to the high power density and high fault tolerance of the dual three-phase permanent magnet synchronous generator. Through system-level coordinated control, it effectively reduces the DC bus voltage fluctuation rate and generator power fluctuation, and comprehensively improves the operational stability, economy and intelligence level of hybrid power ships. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating the energy management method for a hybrid power ship with dual three-phase generators provided in this embodiment of the invention; Figure 2 A structural topology diagram of a hybrid power ship system provided in an embodiment of the present invention; Figure 3 This is a topology diagram of the DTP-PMSG structure provided in an embodiment of the present invention; Figure 4 A flowchart of the training process for the DDPG energy management strategy based on SOH collaborative optimization provided in this embodiment of the invention; Figure 5 A flowchart of an energy management algorithm for DDPG hybrid-powered ships based on multi-objective optimization and TOPSIS provided in an embodiment of the present invention; Figure 6 A schematic diagram of the structure of a hybrid power ship energy management system powered by dual three-phase generators provided in an embodiment of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Figure 1 This is a flowchart of a hybrid power ship energy management method powered by dual three-phase generators according to an embodiment of the present invention, with reference to... Figure 1 , Figure 2 The method includes: S1. Construct a hybrid power ship propulsion system model, which includes a dual three-phase permanent magnet synchronous generator model, a battery model, a supercapacitor model, and a ship load model; and construct a battery health state decay model and a supercapacitor health state decay model, which are used to quantify the battery health state decay during operation and the supercapacitor health state decay during operation, respectively.

[0018] This invention focuses on a marine hybrid power system powered by a dual three-phase permanent magnet synchronous generator (DTP-PMSG). The hybrid power ship system structure and topology are as follows: Figure 2 As shown. It includes four core units: Main power generation unit: Utilizing a DTP-PMSG as the main power supply unit, it converts the AC output from the generator into DC power via a six-phase rectifier and connects in parallel to the DC bus. Hybrid energy storage unit: Composed of batteries and supercapacitors, both connected in parallel to the DC bus via independent bidirectional DC / DC converters, enabling bidirectional power exchange with the DC bus and balancing long-term power support with pulsed load buffering. Load unit: Includes two typical ship operating conditions: step load and pulsed load, all connected to the DC bus. Control unit: Achieves dynamic power allocation and coordinated control of all units in the system through an energy management strategy module.

[0019] Among them, the dual three-phase permanent magnet synchronous generator is a high power density motor with two sets of three-phase windings, which is suitable for ship propulsion systems; the battery is used for long-term energy buffering, and the supercapacitor is used for short-term pulse power compensation; the health status is a quantitative indicator that characterizes the aging degree of energy storage equipment, with a value range of 0 to 1. The greater the attenuation, the more serious the aging; cycle aging refers to the capacity decay caused by repeated charging and discharging, rate aging refers to the accelerated aging caused by high current charging and discharging; pulse impact aging refers to the increase in the internal resistance of the supercapacitor caused by instantaneous high current pulses.

[0020] In this embodiment, S1 first constructs a complete mathematical model of the ship's propulsion system, including the generator, battery, supercapacitor, and load, providing a simulation basis for subsequent optimization. Based on this, semi-empirical models for the health state degradation of the battery and supercapacitor are creatively established for typical step load and pulse load conditions of ships: the degradation of the battery is composed of the superposition of cyclic aging and rate aging; the degradation of the supercapacitor is composed of the superposition of pulse impact aging and cyclic aging.

[0021] Existing technologies lack quantitative models for energy storage aging that address the complex load characteristics of ships, resulting in energy management strategies being unable to accurately assess equipment wear and tear. In this embodiment, S1 constructs a State of Health (SOH) decay model adapted to ship operating conditions, achieving for the first time real-time quantification of the state of health decay during the operation of batteries and supercapacitors. This provides a scientific basis for the quantitative calculation of the "energy storage life" index in subsequent multi-objective optimization, thus solving the fundamental problem that existing technologies cannot quantify energy storage loss and cannot achieve synergistic optimization of fuel economy and equipment life.

[0022] S2. Construct a multi-objective optimization model. The optimization objectives of the multi-objective optimization model include: fuel consumption, cumulative health state decay of the battery, cumulative health state decay of the supercapacitor, DC bus voltage stability index, generator power stability index, battery state of charge deviation index, and supercapacitor state of charge deviation index. Solve the multi-objective optimization model using a multi-objective evolutionary algorithm to obtain a set of Pareto optimal solutions. Use a multi-objective decision-making method to select the solution with the highest proximity from the Pareto optimal solution set as the optimal compromise solution, and deduce the objective weights of each optimization objective based on the optimal compromise solution.

[0023] Among them, the Pareto optimal solution set refers to the set of all feasible solutions that cannot improve other objectives without compromising any of the objectives under multiple conflicting objectives; proximity is an indicator in multi-objective decision-making that measures how close each solution is to the ideal solution; objective weights are weights calculated based on the data itself, which are different from subjective weighting.

[0024] In this embodiment, S1 first constructs a multi-objective optimization model containing seven optimization objectives: fuel consumption (economy), cumulative health state decay of batteries and supercapacitors (equipment lifespan), DC bus voltage fluctuation rate (stability), generator power fluctuation standard deviation (operational smoothness), and the maximum deviation of the state of charge of batteries and supercapacitors from the target value (energy storage dispatchability). Then, the third-generation non-dominated sorting genetic algorithm (NSGA-III) is used to solve this seven-objective model. Through operations such as non-dominated sorting, reference point association, and environment selection, a set of uniformly distributed Pareto optimal solutions is obtained after 500 iterations. Finally, the Technique for Order Preference by Similarity to IdealSolution (TOPSIS) is used to make decisions on the Pareto solution set: after normalizing the values ​​of each solution for each objective, the Euclidean distance from each solution to the positive and negative ideal solutions is calculated to obtain the proximity score, and the solution with the highest proximity score is selected as the optimal compromise solution; then, based on the normalized values ​​of each objective in the optimal compromise solution, the objective weights of the seven optimization objectives are derived by inverse weighting. Existing technologies generally use subjective assignment methods such as the analytic hierarchy process (AHP) or expert scoring for multi-objective weights, which are subject to human bias and difficult to adapt to different navigation conditions.

[0025] This step utilizes the objective weight back-calculation technology that integrates NSGA-III and TOPSIS to automatically generate weights based entirely on system operation data. This enables the power allocation scheme to objectively reflect the real balance between multiple objectives such as fuel economy, energy storage life, and system stability in actual operation, thus solving the problems of subjective weight allocation and decision results deviating from engineering requirements.

[0026] S3. Construct an energy management strategy model based on a deep deterministic policy gradient algorithm. The state space of the energy management strategy model includes ship navigation state parameters, energy storage device state parameters, and power system operating parameters. The action space of the energy management strategy model includes generator power adjustment, battery power adjustment, and supercapacitor power adjustment. Construct a multi-objective reward function based on the objective weights of each optimization objective. Iteratively train the energy management strategy model using the multi-objective reward function until the preset convergence condition is met to obtain the trained energy management strategy model.

[0027] Among them, the deep deterministic policy gradient algorithm is a reinforcement learning algorithm based on the Actor-Critic architecture, which is suitable for control problems in continuous action spaces; the state space is the set of environmental information perceived by the agent; the action space is the set of decision variables that the agent can execute; and the reward function is the feedback signal that measures the quality of the decision.

[0028] This step first constructs an energy management strategy model based on a deep deterministic policy gradient algorithm. This model includes an Actor network (for outputting power allocation actions) and a Critic network (for evaluating the value of actions). Next, a state space is designed, containing eight state variables: ship speed, battery state of charge (SOC), supercapacitor SOC, battery health status, supercapacitor health status, total ship power demand, generator output power, and DC bus voltage, comprehensively reflecting the real-time operating conditions of the ship's propulsion system. Then, an action space is designed, containing three continuous action variables: generator power adjustment, battery power adjustment, and supercapacitor power adjustment. Finally, based on the seven objective weights derived in step S2, a negative-penalty multi-objective reward function is constructed. This function performs a weighted summation of fuel consumption, battery instantaneous health status decay, supercapacitor instantaneous health status decay, DC bus voltage fluctuation rate, generator power fluctuation standard deviation, maximum battery SOC deviation, and maximum supercapacitor SOC deviation, and then takes a negative value. A larger reward value indicates better overall performance. Finally, the reward function is used to iteratively train the policy model. The network parameters are continuously optimized through experience replay and soft update mechanisms until the preset convergence conditions are met (such as reward fluctuation less than 1% and energy storage decay less than the threshold), and the trained energy management policy model is obtained.

[0029] Existing DDPG (Deep Deterministic Policy Gradient) strategies often adopt settings from other domains, lacking specific optimizations for ship operating conditions and energy storage health, resulting in slow convergence and poor stability. This paper addresses these issues by introducing objective weights to reconstruct the reward function, incorporating the energy storage state of health (SOH) into the state space, and setting dedicated convergence conditions for energy storage. These optimization measures significantly improve the algorithm's adaptability and control accuracy for ship hybrid power systems, resolving the problem that existing reinforcement learning strategies struggle to achieve multi-objective collaborative optimization.

[0030] S4. During the actual navigation of the ship, the current state variables are collected in real time, input into the trained energy management strategy model, and output real-time power adjustment commands for the generator, the battery, and the supercapacitor to perform online control of the power distribution of the hybrid power ship propulsion system.

[0031] Among them, the current state variable refers to the parameters that are collected in real time from the actual operation of the ship and correspond to the state space definition in step S3; the real-time power adjustment command refers to the power value that the generator, battery and supercapacitor need to adjust (positive means increase output or discharge, negative means decrease output or charge); online control refers to the real-time execution of energy distribution decisions during actual navigation.

[0032] This step involves deploying the offline-trained energy management strategy model into the execution phase of the actual ship control system. Specifically, during actual ship navigation, sensors and a data acquisition system acquire real-time state variables, including navigation speed, state of charge and health of energy storage devices, total power demand, current generator output power, and DC bus voltage. These state variables are then fed into the energy management strategy model trained in step S3. The Actor network in the model directly outputs three continuous power adjustment quantities based on the current state. After action constraint verification (e.g., dynamically limiting the power range according to SOH), these adjustment quantities are converted into real-time power adjustment commands for the generator, battery, and supercapacitor, which are then sent to the generator controller and the bidirectional DC / DC converter, respectively, thereby achieving online closed-loop control of the hybrid power system's power distribution.

[0033] Existing rule-based strategies cannot adapt to complex operating conditions, while optimization strategies suffer from high computational costs and poor real-time performance. This step achieves millisecond-level online decision-making response by loading a pre-trained deep reinforcement learning model, efficiently handling frequent step loads and pulse load disturbances during ship navigation. Simultaneously, by utilizing energy storage health status information contained in the state space, the control strategy can automatically adjust power allocation to protect aging equipment. The resulting technical effects include: an overall system operating efficiency improvement of over 8%, DC bus voltage fluctuation rate controlled within 2.6%, and a reduction in the cumulative health status decay of batteries and supercapacitors by 41.8% and 37.9%, respectively, truly achieving multi-objective synergistic optimization of fuel economy, equipment lifespan, and system stability.

[0034] Furthermore, the DTP-PMSG is the core power supply component of the marine hybrid power system. Driven coaxially by the prime mover, it converts mechanical energy into alternating current (AC). The electrical energy undergoes AC-DC conversion via a fully controlled rectifier, and then is filtered and stabilized by the bus capacitor before being fed into the DC bus, continuously supplying power to variable loads and energy storage devices. The topology of this power generation system is as follows: Figure 3 As shown.

[0035] The DTP-PMSG mathematical model is established using the Vector Space Decomposition (VSD) method, which has the following effects: d - q The voltage equation in the coordinate system is:

[0036] In the formula, u d , u q For stator d , q Shaft voltage; i d , i q For stator d , q shaft current; R s Stator resistance; L d , L q for d , q Shaft inductance; ω The electric angular velocity of the motor; ψ f It is a permanent magnet flux linkage.

[0037] The electromagnetic torque equation is:

[0038] In the formula, T e Electromagnetic torque; p n This represents the number of pole pairs of the motor. i d , i q For stator d , q shaft current; L d , L q for d , q Shaft inductance; ψ f It is a permanent magnet flux linkage.

[0039] Based on the above embodiments, as a preferred implementation, in step S1, the battery health state degradation model includes a cycle aging degradation sub-model and a rate aging degradation sub-model.

[0040] The cyclic aging attenuation sub-model is used to output the cyclic aging attenuation amount, which is determined by the charge / discharge depth and the number of cycles; the rate aging attenuation sub-model is used to output the rate aging attenuation amount, which is determined by the charge / discharge current rate; the battery health state attenuation amount is the sum of the cyclic aging attenuation amount and the rate aging attenuation amount.

[0041] The supercapacitor health state decay model includes a pulse impact aging decay sub-model and a cyclic aging decay sub-model. The pulse impact aging decay sub-model is used to output the pulse impact aging decay amount, which is determined by the instantaneous current peak value and the pulse frequency. The supercapacitor cyclic aging decay sub-model is used to output the supercapacitor cyclic aging decay amount, which is determined by the number of cycles. The supercapacitor health state decay amount is the sum of the pulse impact aging decay amount and the supercapacitor cyclic aging decay amount.

[0042] Specifically, hybrid energy storage system models include: (1) The battery model adopts the internal resistance equivalent model, and the terminal voltage and SOC equations are as follows:

[0043]

[0044] In the formula, V bat ( t )for t The terminal voltage of the battery at all times; Voc This is the open-circuit voltage; R in For equivalent internal resistance, and SOC and temperature T Related; I bat This refers to the charging and discharging current. SOC bat This refers to the state of charge of the battery. R in ( SOC bat , T The internal resistance of the battery is... SOC and temperature T The function; T This refers to the battery's operating temperature. C rated Rated capacity; Δ t For time step.

[0045] (2) The supercapacitor model, based on the RC equivalent circuit, has the following equations for terminal voltage and energy storage:

[0046]

[0047] In the formula, V sc ( t )for t The terminal voltage of the supercapacitor at any given time; C sc This refers to the capacitance of the supercapacitor. I sc This refers to the charging and discharging current. V sc ( 0 () represents the initial voltage; E sc The total electrostatic energy stored in the supercapacitor.

[0048] (3) Modeling of bidirectional DC / DC converter A bidirectional DC / DC converter is used to enable bidirectional energy flow between the energy storage unit and the DC bus. Boost Mode (discharge) and Buck The output voltage equation for the charging mode is as follows: ( Boost model) ( Buck model) In the formula, V out For output voltage, V inInput voltage; D For duty cycle, by PI Controller adjustment.

[0049]

[0050] in, V ref Reference voltage; K p , K i for PI Controller parameters.

[0051] Ship loads mainly include step loads and pulse loads. Conventional loads are modeled using a resistive model; step loads are simulated using a step function to represent the equipment start-up and shutdown process; pulse loads are modeled using a periodic pulse model, and their power characteristics are as follows:

[0052] In the formula, P pulse ( t )for t The instantaneous power of the pulse load at any given moment; P k This represents the peak power of the pulsed load. T s The working cycle of the pulse load; D The duty cycle of the pulse load; k This is the pulse cycle number. k =0,1,2,...

[0053] DC bus voltage fluctuation rate is used to evaluate voltage stability under load disturbances:

[0054] In the formula, δ u DC bus voltage fluctuation rate; L Number of sampling periods; U k For the first k Periodic average voltage; U kmax , U kmin For the first k Maximum and minimum values ​​of periodic voltage.

[0055] The state of harmonic decay (SOH) of marine batteries is mainly driven by cycle aging and rate aging, and an engineering semi-empirical model is adopted:

[0056] In the formula, SOHbat ( t )for t Monitor the battery's health status at all times; SOH bat ( t- 1) For t- The battery's health status at moment 1 (previous moment); SOH bat,cycle This refers to the aging caused by battery cycles within a single time step. SOH Attenuation; SOH bat,rate This represents the amount of SOH decay caused by rate aging of the battery within a single time step.

[0057] Among them, the amount of degradation due to cyclic aging SOH bat,cycle Positively correlated with depth of charge / discharge (DOD) and number of cycles:

[0058] In the formula, SOH bat,cycle This is caused by battery cycle aging. SOH Attenuation; k 1 represents the cyclic aging factor. (Marine lithium batteries); DOD bat For the depth of charge and discharge of the battery, ; This refers to the cumulative number of charge-discharge cycles for the battery.

[0059] Rate-dependent aging degradation SOH bat,rate It is positively correlated with the charge / discharge current ratio:

[0060] In the formula, SOH bat,rate This is caused by battery rate aging. SOH Attenuation k 2 represents the rate aging factor, calibrated based on the electrochemical and temperature characteristics of marine lithium batteries. Typical values ​​for marine lithium batteries are... k 2 = 5 × 10⁻⁷ (Marine lithium battery); I bat, rated This refers to the rated charge and discharge current of the battery. I bat This refers to the charge / discharge rate of the battery.

[0061] Supercapacitor SOH Attenuation is mainly due to internal resistance aging and capacity decay, and the model is as follows:

[0062] In the formula, for t The state of health of a supercapacitor at any given time characterizes the current degree of aging of the supercapacitor. for t- The health status of the supercapacitor at time 1 (previous time); This refers to the aging caused by pulse impact on the supercapacitor within a single time step. SOH Attenuation; This is due to the cyclic aging of the supercapacitor within a single time step. SOH Attenuation.

[0063] Among them, pulse current attenuation Positively correlated with instantaneous current peak value and pulse frequency:

[0064] In the formula, Caused by pulsed current surges in the supercapacitor. SOH Attenuation; is the pulse aging coefficient, and the typical calibration value for marine supercapacitors is: ; This represents the peak value of the pulse charge / discharge current. This refers to the rated charge and discharge current of the supercapacitor. This is the pulse load frequency.

[0065] Cyclic aging degradation :

[0066] In the formula, Caused by the charge-discharge cycle of the supercapacitor. SOH Attenuation; The cyclic aging factor is the typical calibration value for marine supercapacitors. ; This represents the cumulative number of charge-discharge cycles for the supercapacitor.

[0067] Based on the above embodiments, as a preferred implementation, in step S2, a multi-objective evolutionary algorithm is used to solve the multi-objective optimization model to obtain a set of Pareto optimal solutions, specifically including: Using the power allocation values ​​of the dual three-phase permanent magnet synchronous generator, the battery, and the supercapacitor at each operating moment of the ship as decision variables, an initial population containing multiple power allocation schemes is randomly generated. Decision variables refer to a set of independent parameters whose values ​​need to be determined in an optimization problem; in this invention, they are the output power of the dual three-phase permanent magnet synchronous generator at each operating moment of the ship. P gen ( t ), Battery output power P bat ( t (Discharging is positive, charging is negative) and supercapacitor output power Psc ( t (Discharge is positive, charging is negative), and these three factors together constitute the power allocation scheme of the system. The initial population is the first set of feasible solutions randomly generated when using a multi-objective evolutionary algorithm. Each individual in the population corresponds to a complete power allocation sequence. In this embodiment, the size of the initial population is set to 200, and the decision variables cover all sampling moments of the entire navigation condition. This step solves the problem in the prior art of lacking a systematic search for the power allocation space of the DTP-PMSG hybrid power system: rule-based strategies rely only on empirically preset thresholds and cannot explore the globally optimal allocation; traditional optimization methods can search but often get stuck in local optima. By randomly generating an initial population with a wide coverage, a diverse solution set basis is provided for subsequent non-dominated sorting, ensuring that the algorithm can approach the Pareto front from multiple directions. The initial population of 200 individuals can fully represent the diversity of the power allocation space, avoid the algorithm from converging to the suboptimal region too early, and lay a data foundation for multi-objective collaborative optimization.

[0068] For each power allocation scheme in the current population, calculate the value of each optimization objective, and based on the calculation results, perform non-dominated sorting of the power allocation schemes in the population and divide them into different non-dominated layers.

[0069] Taking typical ship navigation conditions as the optimization scenario, a 7-objective optimization model is constructed, with the objective function as follows: minF ( x )=[ f 1( x ), f 2( x ), f 3( x ), f 4( x ), f 5( x ), f 6( x ), f 7( x )).

[0070] in, , is the total fuel consumption (g) of the ship under all navigation conditions, and is the target for minimization; This refers to the cumulative battery capacity of the battery under all operating conditions. SOH The attenuation amount is used to minimize the target. For the cumulative total of supercapacitors under all operating conditions SOH The attenuation amount is used to minimize the target. , where is the DC bus voltage fluctuation rate (%), and the target is to minimize it; , which is the standard deviation of generator power fluctuation ( kW Minimize the objective; For storage batteries SOC The maximum deviation from the target value is used to minimize the target. f Supercapacitor SOC The maximum deviation from the target value is minimized.

[0071] Non-dominated ranking refers to dividing the population into hierarchical levels based on Pareto dominance: if solution A is not inferior to solution B on all objectives and is superior to solution B on at least one objective, then A dominates B; solutions not dominated by any other solution constitute the first non-dominated layer (Pareto front), and so on. This step solves the problem that existing technologies cannot simultaneously evaluate multi-dimensional performance such as fuel economy, energy storage life, and voltage stability: single-objective optimization completely ignores equipment aging, while subjective weighting methods lack scientific comparative basis. Through non-dominated ranking, the algorithm can objectively identify which power allocation schemes have a relative advantage on all objectives, thereby guiding the search towards the true Pareto boundary. Clearly separating power allocation schemes according to their superiority or inferiority provides a clear evolutionary direction for subsequent reference point association and environment selection, avoiding decision-making biases caused by different dimensions or subjective weights in traditional weighting methods.

[0072] Furthermore, based on the above embodiments, the constraints include: (1) Power balance constraint:

[0073] In the formula, for t Generator output power at any given time; for t The battery output power at all times (negative for charging, positive for discharging). for t The output power of the supercapacitor at all times (negative when charging, positive when discharging); for t Power required by ship load at any given time.

[0074] (2) Equipment power constraints:

[0075] In the formula, This is the minimum allowable output power (idle power) of the generator. This is the maximum allowable output power (rated power) of the generator. for t The battery's output power at all times (negative for charging, positive for discharging). This is the minimum allowable output power (maximum charging power) of the battery. This refers to the maximum allowable output power (maximum discharge power) of the battery.

[0076] for t The output power of the supercapacitor at all times (negative when charging, positive when discharging). This is the minimum allowable output power (maximum charging power) for a supercapacitor. This represents the maximum allowable output power (maximum discharge power) of the supercapacitor.

[0077] (3) Equipment power constraints: ,

[0078] In the formula, for t The state of charge of the battery at all times; For the minimum allowable battery SOC Lower limit; For the maximum allowable capacity of the battery SOC Upper limit; for t The state of charge of the supercapacitor at all times; Minimum allowable for supercapacitors SOC Lower limit; For the maximum allowable value of supercapacitors SOC Upper limit.

[0079] (4) SOH constraint: ≥0.5, ≥0.8 In the formula, for t The state of charge of the battery at all times; for t The state of charge of the supercapacitor at any given time.

[0080] (5) Bus voltage constraint:

[0081] In the formula, for t Real-time voltage of the ship's DC bus; This is the minimum allowable voltage lower limit for the DC bus. This is the maximum allowable voltage limit for the DC bus.

[0082] Reference points are set for each of the optimization objectives, and each power allocation scheme is associated with the reference point closest to the corresponding power allocation scheme.

[0083] Based on the distance between each power allocation scheme and its associated reference point, a power allocation scheme is selected from the current population to form the next generation population.

[0084] Repeat the above steps until the number of iterations reaches a preset threshold, then stop iterating and output the non-dominated power allocation scheme in the current population as the Pareto optimal solution set. The preset threshold is 500 generations; it can also be terminated early if there is no significant change in the Pareto front. Repeat the evolutionary operations such as non-dominated sorting, reference point association, and environment selection. The overall performance of the population in each generation (such as the convergence of the non-dominated layer and the distribution of the solution set) will gradually improve. When the iteration reaches 500 generations, the population can be considered to have stabilized near the true Pareto front, and at this point, all individuals in the first non-dominated layer of the current population constitute the Pareto optimal solution set.

[0085] Based on the above embodiments, as a preferred implementation, in step S2, a multi-objective decision-making method is used to select the solution with the highest proximity from the Pareto optimal solution set as the optimal compromise solution, specifically including: Each non-dominated power allocation scheme in the Pareto optimal solution set is taken as a decision-making scheme, and a decision matrix is ​​constructed using the values ​​of the optimization objectives as decision indicators. The Pareto optimal solution set refers to a set of N non-dominated solutions obtained by solving the seven-objective optimization model using the NSGA-III algorithm. Each solution corresponds to the power allocation sequence of the dual three-phase permanent magnet synchronous generator, battery, and supercapacitor at each time point under all navigation conditions of the ship. A non-dominated power allocation scheme is one in the solution set where no other solution is superior to the solution in all seven optimization objectives. The decision-making schemes are these non-dominated solutions, and the decision indicators are the seven optimization objectives used to evaluate the merits of the schemes: fuel consumption. M fuel Δ SOH bat,total Δ, cumulative health status decay of supercapacitor SOH sc,total DC bus voltage fluctuation rate δ u Standard deviation of generator power fluctuation Δ P gen,std Maximum deviation of battery state of charge | SOC bat-0.6 | max Maximum deviation from the state of charge of supercapacitors | SOC sc- 0.5 | max The original values ​​of each decision-making option on each decision indicator are then assigned. x ij ( i =1,2,…, N , j Arrange the numbers = 1, 2, ..., 7 in rows to form a... N Decision matrix with 7 rows and 7 columns X =( x ij ) N×7 This step addresses the problem in existing technologies that cannot systematically compare the comprehensive performance of various schemes from multiple non-dominated solutions: traditional methods often rely directly on expert experience or simple ranking, lacking a structured data organization. By constructing a decision matrix, the seven-dimensional objective values ​​of all candidate schemes are summarized into a unified data table, providing a standardized data foundation for subsequent normalization processing and distance calculation.

[0086] The decision matrix is ​​then subjected to range normalization to obtain a normalized decision matrix. Since the dimensions and magnitudes of the objectives differ, range normalization is used.

[0087] Obtain the normalized matrix .

[0088] In the formula, After normalization, the first i The solution is at the th solution. j Standardized values ​​for each objective; The j-th optimization objective is the maximum value among all Pareto solutions; For the first j The optimization objective is to find the minimum value among all Pareto solutions. For the first i The Pareto solution, in the... j The original values ​​for each optimization objective; N is the total number of non-dominated solutions in the Pareto front obtained by the NSGA-III algorithm; The normalized decision matrix is ​​used to map all objectives to the interval [0,1].

[0089] Based on the normalized decision matrix, positive ideal solutions and negative ideal solutions are determined, wherein the positive ideal solutions are formed by combining the optimal values ​​of each optimization objective in the Pareto optimal solution set, and the negative ideal solutions are formed by combining the worst values ​​of each optimization objective in the Pareto optimal solution set.

[0090] Specifically, the ideal solution The maximum value of each objective after normalization, i.e.:

[0091] Negative ideal solution The minimum value after normalization of each objective is:

[0092] In the formula, The ideal solution is the theoretically "perfect optimal solution," which is composed of the optimal values ​​of each objective. Negative ideal solution: The theoretical "worst solution", which is composed of the worst values ​​of each objective; For the first i The Pareto solution, in the... j Normalized values ​​on each target; For the first j The normalized maximum value of an objective in all Pareto solutions (the optimal value of the objective). For the first j The objective is the normalized minimum (worst value of the objective) among all Pareto solutions.

[0093] Calculate the first Euclidean distance from each of the decision-making schemes to the positive ideal solution, and the second Euclidean distance from each of the decision-making schemes to the negative ideal solution.

[0094] Specifically, calculate the Euclidean distance from each Pareto solution to the positive and negative ideal solutions:

[0095]

[0096] Based on the first Euclidean distance and the second Euclidean distance of each of the proposed solutions, a proximity degree is calculated for each proposed solution. The proximity degree is equal to the second Euclidean distance divided by the sum of the first and second Euclidean distances. for:

[0097] In the formula, For the first i The Euclidean distance from each Pareto solution to the ideal solution; For the first i The Euclidean distance from a Pareto solution to a negative ideal solution. For the positive ideal solution in the th case j The value of each target (normalized to 1); For the negative ideal solution in the th... j The value of each target (normalized to 0); For the first i The proximity of a Pareto solution represents the overall performance of the solution.

[0098] Select the decision-making scheme with the highest proximity value. C max As the optimal compromise solution, let it be denoted as .

[0099] Based on the above embodiments, as a preferred implementation method, and by back-calculating the objective weights of each optimization objective according to the optimal compromise solution, the specific methods include: Obtain the normalized values ​​of each optimization objective in the normalized decision matrix of the optimal compromise solution.

[0100] Calculate the reciprocal of the normalized value of each optimization objective to obtain the reciprocal weight value of each optimization objective.

[0101] Calculate the sum of the reciprocal weight values ​​of each optimization objective to obtain the total reciprocal weight value.

[0102] For each optimization objective, the reciprocal weight value of the optimization objective is divided by the reciprocal total weight value, and the resulting quotient is used as the objective weight of the optimization objective.

[0103] Specifically, based on the objective contribution of the optimal compromise solution, the weights of each objective are derived in reverse. :

[0104] The objective weights of the seven optimization objectives are obtained and used to reconstruct the DDPG reward function.

[0105] In the formula, For the first j The objective weights of each optimization objective; The optimal compromise solution is at the th j Normalized values ​​on each target; The optimal compromise solution is at the th k Normalized values ​​on an optimization objective.

[0106] Based on the above embodiments, as a preferred implementation, in step S3, the energy management strategy model is iteratively trained using the multi-objective reward function, specifically including: Obtain the real-time battery health status degradation amount, and determine the battery health status value based on the battery health status degradation amount.

[0107] When the battery health status value is lower than a first preset threshold, the range of the battery power adjustment amount is narrowed from the first range to the second range; wherein the second range is a proper subset of the first range, and the upper limit of the second range is less than the upper limit of the first range, and the lower limit of the second range is greater than the lower limit of the first range.

[0108] Obtain the real-time health status decay of the supercapacitor, and determine the supercapacitor health status value based on the supercapacitor health status decay.

[0109] When the health status value of the supercapacitor is lower than the second preset threshold, the range of the supercapacitor power adjustment is narrowed from the third range to the fourth range; wherein the fourth range is a proper subset of the third range, and the upper limit of the fourth range is less than the upper limit of the third range, and the lower limit of the fourth range is greater than the lower limit of the third range.

[0110] The energy management strategy model is iteratively trained using the multi-objective reward function within the second and fourth value ranges.

[0111] Specifically, such as Figure 4 As shown, the DDPG algorithm employs an Actor-Critic dual-network architecture, combining experience replay and soft update mechanisms to achieve optimal control of the continuous action space. The Actor network maps states to actions through a deterministic policy and outputs the optimal power allocation scheme; the Critic network evaluates the value of actions and guides the Actor network optimization. The priority experience replay mechanism adjusts the sample sampling probability based on the temporal difference error (TD-error) to improve training efficiency; the soft update mechanism ensures training stability by slowly updating the target network parameters.

[0112] When designing state variables, the state vector is selected by combining the characteristics of the ship system with the requirements of SOH (State of Health) co-optimization:

[0113] In the formula, S is the state vector; v is the ship's speed (km / h). , These refer to the state of charge of the battery and the supercapacitor, respectively. , These are the health statuses of the battery and the supercapacitor, respectively. This represents the total power demand of the ship (kW). Output power (kW) for DTP-PMSG; This is the DC bus voltage (V).

[0114] The action variables are the power adjustment amounts of each power source, forming a continuous action space: A=

[0115] in, The power change (kW) of DTP-PMSG is taken in the range of [-10, 10]. This represents the change in battery charging and discharging power (kW), with charging being positive and discharging being negative, and the value range is [-20, 20]. The change in supercapacitor charging and discharging power (kW) ranges from -30 to 30.

[0116] Integration SOH The multi-objective reward function is reconstructed based on the weights derived from multi-objective optimization and TOPSIS, resulting in a negative penalty reward function:

[0117] In the formula, R represents the immediate reward; ~ ω 7 represents the multi-objective weights obtained by back-engineering TOPSIS; This represents the normalized fuel consumption. , These are the normalized battery and supercapacitor instantaneous values, respectively. SOH Attenuation; This represents the normalized DC bus voltage fluctuation rate. This represents the normalized standard deviation of generator power fluctuation. , These are the normalized versions of the battery and supercapacitor. SOC maximum deviation.

[0118] based on SOH Dynamic hard constraints limit the charging and discharging behavior of energy storage units: Battery constraints: when hour, The value range is narrowed to [-10, 10] kW; when At that time, it shrinks to [-5,5]kW.

[0119] Supercapacitor confinement: when hour, The value range is narrowed to [-20, 20]kW.

[0120] Based on the above embodiments, as a preferred implementation, in step S3, iterative training of the energy management strategy model includes: Set the health status degradation thresholds for batteries and supercapacitors.

[0121] When collecting experience samples, the instantaneous health state degradation of the battery and the instantaneous health state degradation of the supercapacitor are obtained in each experience sample.

[0122] When the instantaneous health state degradation of the battery exceeds the battery health state degradation threshold, or the instantaneous health state degradation of the supercapacitor exceeds the supercapacitor health state degradation threshold, the empirical sample is removed.

[0123] For each empirical sample that was not removed, calculate the temporal difference error of that empirical sample.

[0124] Based on the instantaneous health state decay of the battery, the instantaneous health state decay of the supercapacitor, and the objective weights of each optimization objective, the weighted total health state decay of the empirical sample is calculated, wherein the weighted total health state decay is equal to the product of the instantaneous health state decay of the battery and the objective weights corresponding to the cumulative health state decay of the battery, plus the product of the instantaneous health state decay of the supercapacitor and the objective weights corresponding to the cumulative health state decay of the supercapacitor.

[0125] Based on the time-series differential error and the weighted total health status decay, the comprehensive sampling priority of the empirical sample is calculated.

[0126] Based on the comprehensive sampling priority of each experience sample, experience samples are sampled from the experience replay pool to update the network parameters of the energy management strategy model.

[0127] Specifically, such as Figure 5 As shown, the sample selection rules are set. SOH Attenuation threshold Remove instantaneous SOH Damaged samples with attenuation exceeding the threshold.

[0128] Priority calculation optimization: Integrating TD-error and SOH Overall priority for attenuation calculation:

[0129] in, .

[0130] In the formula, P The overall priority of the empirical samples; These are the two-dimensional balanced weighting coefficients; This refers to timing difference error; Weighted total of batteries and supercapacitors SOH Attenuation; Weighted total for all samples SOH The maximum value of the attenuation; The target weights for batteries derived from TOPSIS; For supercapacitors SOH Target weight; For the instantaneous battery SOH Attenuation; For the instantaneous supercapacitor SOH Attenuation.

[0131] The dual convergence criterion states that an algorithm is considered to have converged when it simultaneously satisfies both of the following conditions: (1) The average reward value fluctuation of 50 consecutive training rounds is ≤1%.

[0132] (2) 50 consecutive rounds of training .

[0133] The training process includes: (1) Initialize the Actor-Critic network and target network parameters, set the experience replay pool capacity to 10000, and the training rounds to 1300.

[0134] (2) Collect actual ship navigation data and generate training datasets, including different load disturbance scenarios.

[0135] (3) In each round of training, the agent outputs an action through the Actor network according to the current state, executes it after the action constraint verification, and obtains the reward and the next state from the environmental feedback.

[0136] (4) Store the experience samples (S,A,R,S′) into the experience replay pool, and update the network parameters by sampling samples according to the optimized priority.

[0137] (5) Update the target network parameters using a soft update mechanism:

[0138] (6) Repeat the training until the double convergence condition is met, and output the optimal energy management strategy.

[0139] In a second aspect, embodiments of the present invention provide a hybrid power ship energy management system powered by dual three-phase generators. Based on the methods in the above embodiments, the system 600 includes: The system modeling module 610 is used to construct a hybrid power ship propulsion system model, which includes a dual three-phase permanent magnet synchronous generator model, a battery model, a supercapacitor model, and a ship load model; and to construct a battery health state decay model and a supercapacitor health state decay model, which are used to quantify the battery health state decay during operation and the supercapacitor health state decay during operation, respectively.

[0140] The multi-objective optimization module 620 is used to construct a multi-objective optimization model. The optimization objectives of the multi-objective optimization model include: fuel consumption, cumulative health state decay of the battery, cumulative health state decay of the supercapacitor, DC bus voltage stability index, generator power stability index, battery state of charge deviation index, and supercapacitor state of charge deviation index. The multi-objective evolutionary algorithm is used to solve the multi-objective optimization model to obtain a set of Pareto optimal solutions. A multi-objective decision-making method is used to select the solution with the highest closeness from the Pareto optimal solution set as the optimal compromise solution, and the objective weights of each optimization objective are derived from the optimal compromise solution.

[0141] The model training module 630 is used to construct an energy management strategy model based on a deep deterministic policy gradient algorithm. The state space of the energy management strategy model includes ship navigation state parameters, energy storage device state parameters, and power system operating parameters. The action space of the energy management strategy model includes generator power adjustment, battery power adjustment, and supercapacitor power adjustment. A multi-objective reward function is constructed based on the objective weights of each optimization objective. The energy management strategy model is iteratively trained using the multi-objective reward function until a preset convergence condition is met, thereby obtaining a trained energy management strategy model.

[0142] The online control module 640 is used to collect current state variables in real time during the actual navigation of the ship, input them into the trained energy management strategy model, and output real-time power adjustment commands for the generator, the battery, and the supercapacitor to perform online control of the power distribution of the hybrid power ship propulsion system.

[0143] This invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the hybrid power ship energy management method powered by dual three-phase generators as described in the above embodiments of this invention.

[0144] This invention provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the steps of the hybrid power ship energy management method powered by dual three-phase generators as described in the above embodiments of this invention.

[0145] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A method for energy management of hybrid-powered ships powered by dual three-phase generators, characterized in that, include: S1. Construct a hybrid power ship propulsion system model, which includes a dual three-phase permanent magnet synchronous generator model, a battery model, a supercapacitor model, and a ship load model; and construct a battery health state decay model and a supercapacitor health state decay model, which are used to quantify the battery health state decay during operation and the supercapacitor health state decay during operation, respectively. S2. Construct a multi-objective optimization model. The optimization objectives of the multi-objective optimization model include: fuel consumption, cumulative health state decay of the battery, cumulative health state decay of the supercapacitor, DC bus voltage stability index, generator power stability index, battery state of charge deviation index, and supercapacitor state of charge deviation index. Solve the multi-objective optimization model using a multi-objective evolutionary algorithm to obtain a set of Pareto optimal solutions. Use a multi-objective decision-making method to select the solution with the highest proximity from the Pareto optimal solution set as the optimal compromise solution, and deduce the objective weights of each optimization objective based on the optimal compromise solution. S3. Construct an energy management strategy model based on a deep deterministic strategy gradient algorithm. The state space of the energy management strategy model includes ship navigation state parameters, energy storage device state parameters, and power system operating parameters. The action space of the energy management strategy model includes generator power adjustment, battery power adjustment, and supercapacitor power adjustment. A multi-objective reward function is constructed based on the objective weights of each optimization objective; the energy management strategy model is iteratively trained using the multi-objective reward function until the preset convergence condition is met, thereby obtaining the trained energy management strategy model. S4. During the actual navigation of the ship, the current state variables are collected in real time, input into the trained energy management strategy model, and output real-time power adjustment commands for the generator, the battery, and the supercapacitor to perform online control of the power distribution of the hybrid power ship propulsion system.

2. The method according to claim 1, characterized in that, In S1, the battery health status degradation model includes a cycle aging degradation sub-model and a rate aging degradation sub-model. The cyclic aging degradation sub-model is used to output the cyclic aging degradation amount, which is determined by the charge-discharge depth and the number of cycles; the rate aging degradation sub-model is used to output the rate aging degradation amount, which is determined by the charge-discharge current rate; the battery health state degradation amount is the sum of the cyclic aging degradation amount and the rate aging degradation amount. The supercapacitor health state decay model includes a pulse impact aging decay sub-model and a cyclic aging decay sub-model. The pulse impact aging decay sub-model is used to output the pulse impact aging decay amount, which is determined by the instantaneous current peak value and the pulse frequency. The supercapacitor cyclic aging decay sub-model is used to output the supercapacitor cyclic aging decay amount, which is determined by the number of cycles. The decay of the supercapacitor's health status is the sum of the decay caused by pulse shock aging and the decay caused by cyclic aging of the supercapacitor.

3. The method according to claim 1, characterized in that, In step S2, a multi-objective evolutionary algorithm is used to solve the multi-objective optimization model to obtain a set of Pareto optimal solutions, specifically including: Using the power allocation values ​​of the dual three-phase permanent magnet synchronous generator, the battery, and the supercapacitor at each operating moment of the ship as decision variables, an initial population containing multiple power allocation schemes is randomly generated. For each power allocation scheme in the current population, calculate the value of each optimization objective, and based on the calculation results, perform non-dominated sorting of the power allocation schemes in the population and divide them into different non-dominated layers; Reference points are set for each of the optimization objectives, and each power allocation scheme is associated with the reference point closest to the corresponding power allocation scheme. Based on the distance between each power allocation scheme and its associated reference point, a power allocation scheme is selected from the current population to form the next generation population; Repeat the above steps until the number of iterations reaches a preset threshold, then stop the iteration and output the non-dominated power allocation scheme in the current population as the Pareto optimal solution set.

4. The method according to claim 1, characterized in that, In step S2, a multi-objective decision-making method is used to select the solution with the highest proximity from the Pareto optimal solution set as the optimal compromise solution, specifically including: Each non-dominated power allocation scheme in the Pareto optimal solution set is taken as a decision-making scheme, and a decision matrix is ​​constructed using the values ​​of each optimization objective as decision indicators. The decision matrix is ​​normalized by range normalization to obtain a normalized decision matrix; Based on the normalized decision matrix, positive ideal solutions and negative ideal solutions are determined, wherein the positive ideal solution is formed by combining the optimal values ​​of each optimization objective in the Pareto optimal solution set, and the negative ideal solution is formed by combining the worst values ​​of each optimization objective in the Pareto optimal solution set. Calculate the first Euclidean distance from each of the decision-making schemes to the positive ideal solution, and the second Euclidean distance from each of the decision-making schemes to the negative ideal solution; Based on the first Euclidean distance and the second Euclidean distance of each of the proposed solutions, the proximity of each proposed solution is calculated, wherein the proximity is equal to the second Euclidean distance divided by the sum of the first Euclidean distance and the second Euclidean distance; The solution with the highest proximity value is selected as the optimal compromise solution.

5. The method according to claim 4, characterized in that, And based on the optimal compromise solution, the objective weights of each optimization objective are derived, specifically including: Obtain the normalized values ​​of each optimization objective in the normalized decision matrix of the optimal compromise solution; Calculate the reciprocal of the normalized value of each optimization objective to obtain the reciprocal weight value of each optimization objective; Calculate the sum of the reciprocal weight values ​​of each optimization objective to obtain the total reciprocal weight value; For each optimization objective, the reciprocal weight value of the optimization objective is divided by the reciprocal total weight value, and the resulting quotient is used as the objective weight of the optimization objective.

6. The method according to claim 1, characterized in that, In step S3, the energy management strategy model is iteratively trained using the multi-objective reward function, specifically including: Obtain the real-time battery health status degradation amount, and determine the battery health status value based on the battery health status degradation amount; When the battery health status value is lower than the first preset threshold, the range of the battery power adjustment amount is narrowed from the first range to the second range; wherein the second range is a proper subset of the first range, and the upper limit of the second range is less than the upper limit of the first range, and the lower limit of the second range is greater than the lower limit of the first range. Obtain the real-time health status decay of the supercapacitor, and determine the health status value of the supercapacitor based on the supercapacitor health status decay. When the health status value of the supercapacitor is lower than the second preset threshold, the range of the supercapacitor power adjustment is narrowed from the third range to the fourth range; wherein the fourth range is a proper subset of the third range, and the upper limit of the fourth range is less than the upper limit of the third range, and the lower limit of the fourth range is greater than the lower limit of the third range. The energy management strategy model is iteratively trained using the multi-objective reward function within the second and fourth value ranges.

7. The method according to claim 6, characterized in that, In step S3, iterative training of the energy management strategy model includes: Set the battery health status degradation threshold and the supercapacitor health status degradation threshold; When collecting experience samples, the instantaneous health state degradation of the battery and the instantaneous health state degradation of the supercapacitor are obtained in each experience sample. When the instantaneous health state degradation of the battery exceeds the battery health state degradation threshold, or the instantaneous health state degradation of the supercapacitor exceeds the supercapacitor health state degradation threshold, the empirical sample is removed. For each empirical sample that was not removed, calculate the temporal difference error of that empirical sample; Based on the instantaneous health state decay of the battery, the instantaneous health state decay of the supercapacitor, and the objective weights of each optimization objective, the weighted total health state decay of the empirical sample is calculated, wherein the weighted total health state decay is equal to the product of the instantaneous health state decay of the battery and the objective weights corresponding to the cumulative health state decay of the battery, plus the product of the instantaneous health state decay of the supercapacitor and the objective weights corresponding to the cumulative health state decay of the supercapacitor. Based on the time-series differential error and the weighted total health status decay, the comprehensive sampling priority of the empirical sample is calculated. Based on the comprehensive sampling priority of each experience sample, experience samples are sampled from the experience replay pool to update the network parameters of the energy management strategy model.

8. A hybrid power ship energy management system powered by dual three-phase generators, characterized in that, include: The system modeling module is used to construct a hybrid power ship propulsion system model, which includes a dual three-phase permanent magnet synchronous generator model, a battery model, a supercapacitor model, and a ship load model; and to construct a battery health state decay model and a supercapacitor health state decay model, which are used to quantify the battery health state decay during operation and the supercapacitor health state decay during operation, respectively. A multi-objective optimization module is used to construct a multi-objective optimization model. The optimization objectives of the multi-objective optimization model include: fuel consumption, cumulative health state degradation of the battery, cumulative health state degradation of the supercapacitor, DC bus voltage stability index, generator power stability index, battery state of charge deviation index, and supercapacitor state of charge deviation index. A multi-objective evolutionary algorithm is used to solve the multi-objective optimization model to obtain a set of Pareto optimal solutions. A multi-objective decision-making method is used to select the solution with the highest closeness from the Pareto optimal solution set as the optimal compromise solution, and the objective weights of each optimization objective are derived from the optimal compromise solution. The model training module is used to construct an energy management strategy model based on a deep deterministic policy gradient algorithm. The state space of the energy management strategy model includes ship navigation state parameters, energy storage device state parameters, and power system operating parameters. The action space of the energy management strategy model includes generator power adjustment, battery power adjustment, and supercapacitor power adjustment. A multi-objective reward function is constructed based on the objective weights of each optimization objective. The energy management strategy model is iteratively trained using the multi-objective reward function until a preset convergence condition is met, thereby obtaining a trained energy management strategy model. The online control module is used to collect current state variables in real time during the actual navigation of the ship, input them into the trained energy management strategy model, and output real-time power adjustment commands for the generator, the battery, and the supercapacitor, so as to perform online control of the power distribution of the hybrid power ship propulsion system.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the hybrid power ship energy management method powered by dual three-phase generators as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the hybrid power ship energy management method powered by dual three-phase generators as described in any one of claims 1 to 7.