Adaptive dynamic programming hydrogen energy supply chain optimization decision-making method and system

By using an adaptive dynamic programming method, the hydrogen energy supply chain is transformed into a two-player zero-sum game problem. Neural networks are used to optimize decision-making, which solves the problems of dynamic balance between supply and demand and asymmetric constraints in the hydrogen energy supply chain, thus achieving a stable and economical hydrogen energy supply.

CN121809878APending Publication Date: 2026-04-07SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The hydrogen energy supply chain faces high uncertainty and disturbances, complex dynamic characteristics and asymmetric constraints, making it difficult for traditional optimization methods to achieve dynamic supply and demand balance and effective decision-making.

Method used

An adaptive dynamic programming method is adopted to abstract the hydrogen energy supply chain into a discrete-time affine nonlinear system. A utility function is introduced to transform it into a two-player zero-sum game problem, which is solved by approximation using a neural network. The decision is optimized through the H∞ control framework and online learning mechanism.

Benefits of technology

It achieves stable hydrogen storage in the hydrogen energy supply chain under worst-case demand disturbances, satisfies asymmetric constraints, reduces system modeling complexity, and enables global dynamic optimization and economic decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809878A_ABST
    Figure CN121809878A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of energy management and intelligent decision, and provides a self-adaptive dynamic programming hydrogen energy supply chain optimization decision method and system, which abstracts a hydrogen energy supply chain into a discrete-time affine nonlinear system, introduces a utility function, constructs a performance index function considering asymmetric constraint, and improves the performance of the hydrogen energy supply chain. The hydrogen storage state value is maintained at an expected level, and a control problem is converted into a double-person zero-sum game problem; a self-adaptive dynamic programming method based on a neural network is adopted to approximately solve the double-person zero-sum game problem, according to data collected along an actual running track of a system, the weight of each network is optimized and updated by taking error function minimization of each network in the neural network as a target, and a decision result after solving is obtained. The method does not depend on an accurate model, can adapt to environment changes online, effectively processes asymmetric constraints and can suppress demand disturbance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy management and intelligent decision-making technology, specifically relating to an adaptive dynamic programming method and system for optimizing hydrogen energy supply chain decisions. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Hydrogen energy, as a clean and efficient secondary energy source, is an important component of the future energy system. The hydrogen energy industry chain includes upstream hydrogen production (such as industrial by-product hydrogen and hydrogen production from renewable energy water electrolysis), midstream storage and transportation (such as high-pressure gaseous hydrogen tubular vehicles, liquid hydrogen tank trucks, and hydrogen pipelines), and downstream applications (such as hydrogen refueling stations, fuel cell power generation, and chemical raw materials). The key to achieving the safe, economical, and efficient operation of the hydrogen energy industry chain lies in achieving a dynamic balance between supply and demand.

[0004] However, optimizing the hydrogen energy supply chain faces many challenges: High degree of uncertainty and disturbance: Hydrogen production, especially water electrolysis relying on renewable energy sources such as wind and solar power, exhibits significant fluctuations in output. Demand from hydrogen users, such as for hydrogen refueling stations in the transportation sector and industrial users, also varies in real-time and randomly. These factors constitute strong disturbances to the system.

[0005] Complex dynamic characteristics: The storage and transportation of hydrogen energy involves physical and chemical changes, making the entire system a complex, nonlinear dynamic system. Traditional static optimization methods (such as linear programming and mixed integer programming) struggle to capture the real-time dynamics of the system and are unable to make effective forward-looking decisions.

[0006] Real-world physical and operational constraints: Each piece of equipment in the supply chain has its own physical and operational limitations. For example, the hydrogen production rate of an electrolyzer has upper and lower limits, and its ramp-up and deceleration rates may differ; the transport capacity and cycle time of the tube bundle vehicle are fixed; and the capacity and safe pressure of the hydrogen storage tank are finite. In particular, many constraints are asymmetric; for example, the minimum technical output of a hydrogen production plant cannot be zero, while the maximum output is limited by equipment capacity. Such contingent conditions are difficult to handle with traditional control methods.

[0007] Model dependency issue: It is very difficult to accurately establish a mathematical model for the entire hydrogen energy industry chain, as system parameters may change at any time (such as transportation efficiency, equipment aging, etc.). Traditional control methods that rely on accurate models have poor adaptability. Summary of the Invention

[0008] To address the aforementioned problems, this invention proposes an adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method and system. This invention does not rely on precise models, can adapt to environmental changes online, effectively handles asymmetric constraints, and can suppress demand disturbances.

[0009] According to some embodiments, the present invention adopts the following technical solution: An adaptive dynamic programming method for optimizing the hydrogen energy supply chain includes the following steps: The hydrogen energy supply chain is abstracted as a discrete-time affine nonlinear system, where the system's state vector represents the hydrogen storage state value of each key hydrogen storage node in the supply chain, and the control input vector represents the decision variables subject to asymmetric constraints. By introducing a utility function and constructing a performance index function that considers asymmetric constraints, the hydrogen storage state value is maintained at the desired level, thus transforming the control problem into a two-person zero-sum game problem. An adaptive dynamic programming method based on neural networks is used to approximate the solution of the two-player zero-sum game problem. The neural network includes an evaluation network, an execution network, and a perturbation network. Based on data collected along the actual running trajectory of the system, the weights of each network are optimized and updated with the goal of minimizing the error function of each network in the neural network, and the solution result is obtained.

[0010] As an alternative implementation, the state vector is used to represent the amount or pressure of hydrogen storage at each key hydrogen storage node in the supply chain at each discrete moment, including hydrogen production plant storage tanks, regional hydrogen storage centers and hydrogen refueling station storage tanks.

[0011] As an alternative implementation, the control input vector It includes m controllable operations, encompassing the hydrogen production rate of each hydrogen production plant and the hydrogen transport rate of each transportation route. The control input vector is subject to asymmetric constraints, i.e. ,in and These are the minimum and maximum allowed values ​​for the i-th control input, and .

[0012] As an alternative implementation method, the system is represented as follows:

[0013] in, , is the state vector. To control the input vector, Let q represent q uncontrollable external disturbances. Represents the internal dynamics of the system. Let be the control input function, and be the function with respect to the state vector. Nonlinear functions.

[0014] As an alternative implementation, the process of constructing a performance index function considering asymmetric constraints includes: the infinite time-domain performance index function is:

[0015] in It is a penalty term for deviations in hydrogen storage state, with the goal of maximizing hydrogen storage capacity. Maintaining at the expected level, weight matrix It is positive definite. It is a penalty term for the control input, used to handle asymmetric constraints;

[0016] in,

[0017] in It controls the upper limit of input. It controls the lower limit of the input. It is an integral variable; This functional form ensures that the cost increases sharply as the control quantity approaches its boundary, thus naturally limiting the solution to the feasible region. It is a measure of the disturbance energy, and a preset disturbance suppression level. It is a positive definite weighted matrix.

[0018] As an alternative implementation method, the process of transforming the control problem into a two-player zero-sum game problem includes: control strategy As the minimizing party, the goal is to minimize the performance metric function; while the perturbation As the maximizing side, representing the worst-case hydrogen demand, the goal is to maximize performance index J, with the objective of finding a control pair. That is, a saddle point, satisfying:

[0019] in and They are respectively and The optimal value.

[0020] As an alternative implementation, the process of approximating the solution using a neural network-based adaptive dynamic programming method includes: constructing an evaluation network for approximating the function. That is, from the state The initial future accumulated cost has the following structure:

[0021] in, It is the weight vector used to evaluate the network. It is its activation function; Construct an execution network for approximate optimal control strategy Its output is the optimal hydrogen production and transportation decision, with the following structure:

[0022] It is the weight vector of the execution network. It is its activation function; Due to the existence of asymmetric constraints, the actual output control strategy is obtained through a hyperbolic tangent function transformation to satisfy the constraints:

[0023] in It is a constant used to handle asymmetric input constraints. It is the offset of an asymmetric constraint; The goal of the execution network is to learn to generate this policy; A perturbation network is constructed to approximate the worst-case perturbation, i.e., the most severe hydrogen demand pattern. Its structure is as follows: .

[0024] It is the weight vector of the perturbation network. It is its activation function; As an alternative implementation, the process of optimizing and updating the weights of each network with the goal of minimizing the error function of each network in the neural network includes using an online adaptive policy learning algorithm, based on data collected along the actual operating trajectory of the system. Simultaneously update the weights of the evaluation, execution, and perturbation networks. The goal of the update is to minimize their respective error functions.

[0025] As an alternative implementation, the process of minimizing the respective error functions is as follows: the update of the evaluation network aims to minimize the Bellman error, and the updates of the execution network and the perturbation network aim to make their policies approximate the optimal policy under the current value function.

[0026] An adaptive dynamic programming-based hydrogen energy supply chain optimization decision system includes: The affine nonlinear system representation module is configured to abstract the hydrogen energy supply chain as a discrete-time affine nonlinear system, where the system's state vector represents the hydrogen storage state values ​​of each key hydrogen storage node in the supply chain, and the control input vector represents the decision variables subject to asymmetric constraints. The control conversion module is configured to introduce a utility function and construct a performance index function that considers asymmetric constraints in order to maintain the hydrogen storage state value at the desired level, thus transforming the control problem into a two-person zero-sum game problem. The adaptive dynamic programming solution module is configured to approximate the two-player zero-sum game problem using an adaptive dynamic programming method based on a neural network. The neural network includes an evaluation network, an execution network, and a perturbation network. Based on data collected along the actual running trajectory of the system, the module optimizes and updates the weights of each network with the goal of minimizing the error function of each network in the neural network, thereby obtaining the solution decision result.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing the H∞ control framework and zero-sum game, the decision-making strategy generated by this invention can ensure the stability of hydrogen storage in the hydrogen energy supply chain under the worst demand disturbance, and avoid system collapse caused by sudden increase in demand or sudden decrease in supply, thus achieving a robust balance between supply and demand.

[0028] This invention introduces a non-quadratic performance index to handle asymmetric constraints, enabling the model to accurately reflect the asymmetric operational limitations of hydrogen production and transportation equipment (such as electrolyzers and compressors), providing precise modeling of real-world constraints, and resulting in more realistic and economical decisions.

[0029] This invention employs adaptive dynamic programming and online learning mechanisms, enabling the decision-making system to "learn as it runs," without relying on a precise system model. It can automatically adapt to changes in supply chain structure or operating parameters, greatly reducing the complexity of system modeling and maintenance.

[0030] This invention provides a complete closed-loop decision-making framework from system modeling and performance index definition to online solution, realizing end-to-end optimization decision-making. It can automatically generate optimal hydrogen production and transportation instructions based on real-time hydrogen storage status, achieving global, dynamic and forward-looking optimization of the entire hydrogen energy industry chain.

[0031] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0032] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0033] Figure 1 This is a flowchart of one embodiment. Detailed Implementation

[0034] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0035] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0036] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0037] Where there is no conflict, the embodiments and features described in this application may be combined with each other.

[0038] Example 1 As described in the background section, existing hydrogen energy supply chain optimization decision-making schemes face the following challenges: how to achieve dynamic supply and demand balance in the presence of unknown system dynamics and external demand disturbances; how to effectively integrate the asymmetric physical and operational constraints of each link in the industry chain (such as the asymmetric upper and lower limits of hydrogen production and transportation rates) into the optimal decision-making model; and how to design a robust decision-making algorithm capable of online learning, real-time optimization, and minimizing operating costs, ensuring stable hydrogen storage levels, and withstanding worst-case demand fluctuations.

[0039] This invention provides an adaptive dynamic programming decision-making method based on asymmetric constraints. This method transforms the supply and demand balance problem of the hydrogen energy supply chain into a robust control problem of a nonlinear system, and solves it using zero-sum game theory, such as... Figure 1 As shown, the core steps are as follows: Step 1: Dynamic System Modeling of the Hydrogen Energy Supply Chain The hydrogen energy supply chain is abstracted as a discrete-time affine nonlinear system. The system's state, control input, and disturbance are defined as follows: State vector : Represents the amount or pressure of hydrogen stored at n key hydrogen storage nodes in the supply chain (such as hydrogen production plant storage tanks, regional hydrogen storage centers, hydrogen refueling station storage tanks, etc.) at discrete time k.

[0040] Control input vector : Represents the decision variable, including m controllable operations, such as the hydrogen production rate of each hydrogen production plant and the hydrogen transportation rate of each transportation route. This input is subject to asymmetric constraints, i.e. ,in and These are the minimum and maximum allowed values ​​for the i-th control input, and may be... .

[0041] Perturbation input vector: This represents q uncontrollable external disturbances, primarily the hydrogen demand rate of end users.

[0042] The state equation of the system can be expressed as: ; in This represents the internal dynamics of the system, such as the natural loss and leakage of hydrogen. To control the input functions, they can be related to the state. The nonlinear function, where , .

[0043] Step 2: Construct a performance index function considering asymmetric constraints and a zero-sum game. To achieve supply-demand balance and economic optimization, we design a performance index function. This function not only penalizes deviations in hydrogen storage state and energy consumption of control inputs but also suppresses the worst-case effects of external disturbances. Specifically, to handle the asymmetric constraints of the control inputs, we introduce a non-quadratic utility function. Define the following infinite-time domain performance index function:

[0044] in It is a penalty term for deviations in hydrogen storage state, with the goal of maximizing hydrogen storage capacity. Maintaining at the expected level, weight matrix It is positive definite. It is a penalty term for the control input, used to handle asymmetric constraints;

[0045] in,

[0046] in It controls the upper limit of input. It controls the lower limit of the input. It is an integral variable; This functional form ensures that the cost increases sharply as the control quantity approaches its boundary, thus naturally limiting the solution to the feasible region. It is a measure of the disturbance energy, and a preset disturbance suppression level. It is a positive definite weighted matrix.

[0047] The control problem is transformed into a two-player zero-sum game problem: control strategy As the minimizing party, the goal is to minimize the performance metric function; while the perturbation As the maximizing side (representing the worst-case hydrogen demand), the goal is to maximize performance index J. Our objective is to find a control pair... That is, a saddle point, satisfying:

[0048] in and They are respectively and optimal value Step 3: Solving the Hamilton-Jacobi-Isaacs equations based on adaptive dynamic programming The optimal solution to this zero-sum game problem requires solving the discrete-time Hamilton-Jacobi-Isaacs (HJI) equations. However, the HJI equations are highly nonlinear and difficult to solve directly.

[0049] This invention employs an adaptive dynamic programming method based on neural networks to approximate the solution. To this end, three neural networks are constructed: Critic Network: Used for approximating value functions That is, from the state The initial future accumulated cost, its structure is as follows

[0050] in It is the weight vector used to evaluate the network. It is its activation function.

[0051] Actor Network: Used for approximate optimal control strategies Its output is the optimal hydrogen production and transportation decision. Its structure is as follows:

[0052] Due to the existence of asymmetric constraints, the actual output control strategy is obtained through a hyperbolic tangent function transformation to satisfy the constraints:

[0053] in It is a constant used to handle asymmetric input constraints. It is the offset of an asymmetric constraint; The goal of the execution network is to learn and generate this policy.

[0054] Disturbance Network: Used to approximate worst-case perturbations, i.e., the most severe hydrogen demand patterns. Its structure is as follows:

[0055] Step 4: Online Synchronous Update Algorithm for Neural Network Weights This invention proposes an online adaptive policy learning algorithm, which uses data collected along the actual operating trajectory of the system. Simultaneously update the weights of the evaluation, execution, and perturbation networks. The goal of the update is to minimize the respective error function. For example, the update of the evaluation network aims to minimize the Bellman error, while the update of the execution and perturbation network aims to make its policy approach the optimal policy under the current value function.

[0056] The algorithm provided in this embodiment updates without relying on precise knowledge in the system's state equations; it can adapt in real time to changes in hydrogen energy supply chain parameters and environmental dynamics. Compared to traditional alternating update strategies (strategy iteration or value iteration), synchronous updates improve learning efficiency and convergence speed.

[0057] Example 2 An adaptive dynamic programming-based hydrogen energy supply chain optimization decision system includes: The affine nonlinear system representation module is configured to abstract the hydrogen energy supply chain as a discrete-time affine nonlinear system, where the system's state vector represents the hydrogen storage state values ​​of each key hydrogen storage node in the supply chain, and the control input vector represents the decision variables subject to asymmetric constraints. The control conversion module is configured to introduce a utility function and construct a performance index function that considers asymmetric constraints in order to maintain the hydrogen storage state value at the desired level, thus transforming the control problem into a two-person zero-sum game problem. The adaptive dynamic programming solution module is configured to approximate the two-player zero-sum game problem using an adaptive dynamic programming method based on a neural network. The neural network includes an evaluation network, an execution network, and a perturbation network. Based on data collected along the actual running trajectory of the system, the module optimizes and updates the weights of each network with the goal of minimizing the error function of each network in the neural network, thereby obtaining the solution decision result.

[0058] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of one or more computer-usable storage media (including, but not limited to, disk storage, etc.) containing computer-usable program code. CD - ROMIt takes the form of a computer program product implemented on (such as optical memory, etc.).

[0059] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0060] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive dynamic programming method for optimizing the hydrogen energy supply chain, characterized in that, Includes the following steps: The hydrogen energy supply chain is abstracted as a discrete-time affine nonlinear system, where the system's state vector represents the hydrogen storage state value of each key hydrogen storage node in the supply chain, and the control input vector represents the decision variables subject to asymmetric constraints. By introducing a utility function and constructing a performance index function that considers asymmetric constraints, the hydrogen storage state value is maintained at the desired level, thus transforming the control problem into a two-person zero-sum game problem. An adaptive dynamic programming method based on neural networks is used to approximate the solution of the two-player zero-sum game problem. The neural network includes an evaluation network, an execution network, and a perturbation network. Based on data collected along the actual running trajectory of the system, the weights of each network are optimized and updated with the goal of minimizing the error function of each network in the neural network, and the solution result is obtained.

2. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The state vector is used to represent the amount or pressure of hydrogen storage at each key hydrogen storage node in the supply chain at each discrete moment, including hydrogen production plant storage tanks, regional hydrogen storage centers and hydrogen refueling station storage tanks.

3. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The control input vector It includes m controllable operations, encompassing the hydrogen production rate of each hydrogen production plant and the hydrogen transport rate of each transportation route. The control input vector is subject to asymmetric constraints, i.e. ,in and These are the minimum and maximum allowed values ​​for the i-th control input, and .

4. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The system is represented as: in, , is the state vector. To control the input vector, Let q represent q uncontrollable external disturbances. Represents the internal dynamics of the system. Let be the control input function, and be the function with respect to the state vector. Nonlinear functions.

5. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The process of constructing a performance index function considering asymmetric constraints includes: the infinite time-domain performance index function is: in It is a penalty term for deviations in hydrogen storage state, with the goal of maximizing hydrogen storage capacity. Maintaining at the expected level, weight matrix It is positive definite. It is a penalty term for the control input, used to handle asymmetric constraints; in, in It controls the upper limit of input. It controls the lower limit of the input. It is an integral variable; This functional form ensures that the cost increases sharply as the control quantity approaches its boundary, thus naturally limiting the solution to the feasible region. It is a measure of the disturbance energy, and a preset disturbance suppression level. It is a positive definite weighted matrix.

6. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The process of transforming a control problem into a two-player zero-sum game includes: control strategies. As the minimizing party, the goal is to minimize the performance metric function; while the perturbation As the maximizing side, representing the worst-case hydrogen demand, the goal is to maximize performance index J, with the objective of finding a control pair. That is, a saddle point, satisfying: 。 in and They are respectively and The optimal value.

7. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The process of approximating the solution using a neural network-based adaptive dynamic programming method includes: constructing an evaluation network for approximating the function. That is, from the state The initial future accumulated cost has the following structure: in, It is the weight vector used to evaluate the network. It is its activation function; Construct an execution network for approximate optimal control strategy Its output is the optimal hydrogen production and transportation decision, with the following structure: It is the weight vector of the execution network. It is its activation function; Due to the existence of asymmetric constraints, the actual output control strategy is obtained through a hyperbolic tangent function transformation to satisfy the constraints: in It is a constant used to handle asymmetric input constraints. It is the offset of an asymmetric constraint; The goal of the execution network is to learn to generate this policy; A perturbation network is constructed to approximate the worst-case perturbation, i.e., the most severe hydrogen demand pattern. Its structure is as follows: It is the weight vector of the perturbation network. It is its activation function.

8. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The process of optimizing and updating the weights of each network in a neural network, with the goal of minimizing the error function of each network, includes using an online adaptive policy learning algorithm and data collected along the actual operating trajectory of the system. Simultaneously update the weights of the evaluation, execution, and perturbation networks. The goal of the update is to minimize their respective error functions.

9. The adaptive dynamic programming-based hydrogen energy supply chain optimization decision-making method as described in claim 1, characterized in that, The process of minimizing their respective error functions is as follows: the update of the evaluation network aims to minimize the Bellman error, and the updates of the execution network and the perturbation network aim to make their policies approximate the optimal policy under the current value function.

10. An adaptive dynamic programming-based hydrogen energy supply chain optimization decision system, characterized in that, include: The affine nonlinear system representation module is configured to abstract the hydrogen energy supply chain as a discrete-time affine nonlinear system, where the system's state vector represents the hydrogen storage state values ​​of each key hydrogen storage node in the supply chain, and the control input vector represents the decision variables subject to asymmetric constraints. The control conversion module is configured to introduce a utility function and construct a performance index function that considers asymmetric constraints in order to maintain the hydrogen storage state value at the desired level, thus transforming the control problem into a two-person zero-sum game problem. The adaptive dynamic programming solution module is configured to approximate the two-player zero-sum game problem using an adaptive dynamic programming method based on a neural network. The neural network includes an evaluation network, an execution network, and a perturbation network. Based on data collected along the actual running trajectory of the system, the module optimizes and updates the weights of each network with the goal of minimizing the error function of each network in the neural network, thereby obtaining the solution decision result.