Hybrid regional motor train unit configuration and control deep collaborative optimization method
By optimizing the parameter matching and energy management control of multi-stack fuel cell systems, the stability and life issues of the fuel cell system in urban EMUs were solved, and efficient energy conversion and durability of the system were improved.
Patent Information
- Application Number
- CN202411231188.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The existing multi-stack fuel cell hybrid system in urban EMUs has problems such as the parameter matching process affecting the energy management results, low energy conversion efficiency, poor system durability, and poor internal balance of the multi-stack fuel cell system, resulting in insufficient system stability and life.
A deep collaborative optimization method for hybrid urban EMU configuration and control is adopted. By constructing a system constraint penalty model, a state-action pair value function and a flexible action-evaluation algorithm based on direction guidance, combined with a three-objective non-dominated sorting genetic algorithm with a scaling factor, the parameter matching and energy management control of the multi-stack fuel cell system are optimized to achieve overall optimization of the system.
It improves the performance coordination of multiple fuel cell stacks, extends the system life, reduces hydrogen consumption, reduces the configuration cost of the fuel cell hybrid system, and improves the stability and durability of the system.
Smart Images

Figure CN119099374B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of hybrid city regional multiple units, in particular to a configuration and control deep synergy optimization method for a hybrid city regional multiple unit. BACKGROUND
[0002] Proton exchange membrane fuel cells have been applied in the fields of automobiles, trucks, shunting locomotives, etc. due to their zero pollution, low noise, high efficiency and other advantages. In order to solve the problems that fuel cells cannot recover braking energy and have low energy conversion efficiency, a storage battery is usually provided for the fuel cells to form a fuel cell hybrid power system to achieve flexible conversion of energy. Nowadays, a complete fuel cell hybrid power system usually includes a fuel cell system, a storage battery system, a one-way direct-direct converter, a bidirectional direct-direct converter, a signal collection system, an energy management and control system and other parts. With the application of fuel cell hybrid power systems in high-power scenarios (such as city regional multiple units), a hybrid power system with only a single fuel cell cannot meet the high-power demand of system operation. Therefore, multiple fuel cell monomers are connected in parallel to form a multi-stack fuel cell hybrid power system to solve the above problems. At present, the multi-stack fuel cell hybrid power system has the following defects which need to be further optimized:
[0003] The parameter matching process affects the energy management result: the parameter design process of the traditional city regional multiple unit is a sequential process based on vehicle design parameters and line conditions to determine the hybrid power system parameter configuration according to power and energy constraints, which considers less constraints and ignores some actual factors, and does not consider the influence of parameter matching result on the energy management process. The power of the equipped fuel cell or storage battery is too large, which may cause the output power of the fuel cell or storage battery to change frequently during the operation of the hybrid power system, affecting its service life and reducing the stability of the system.
[0004] Low energy conversion efficiency: the traditional multi-stack fuel cell system distributes the demand power of the city regional multiple unit to each stack through an average distribution or chain distribution strategy. These methods do not consider the differences between the stacks and cannot finely adjust the output power of the single stack. The performance decay of the stack seriously affects its energy conversion efficiency, and the energy conversion efficiency of the stack with serious performance decay is low, which further leads to low average energy conversion efficiency of the multi-stack fuel cell system.
[0005] Poor system durability: the short service life and poor durability of fuel cells are one of the important reasons hindering their large-scale application. The performance decay of fuel cells is largely affected by their operating conditions. Excessive decay of fuel cell performance will affect the safe and stable operation of the system. Therefore, it is necessary to model and analyze the performance decay of fuel cells. Unreasonable operating conditions will cause the service life of fuel cells to decrease significantly, thereby reducing the durability of the system.
[0006] Poor internal balance of multi-stack fuel cell system: current research on multi-stack fuel cell system focuses on the performance change of each fuel cell in the system, and is usually optimized to reduce the performance degradation of individual fuel cells. It does not consider that the overuse or premature degradation of any fuel cell in the multi-stack fuel cell system as a whole will disrupt the balance of the entire system, leading to system instability and affecting the safe operation of the system. SUMMARY
[0007] To solve the above problems, the present application proposes a hybrid power city regional motor train unit configuration and control deep collaborative optimization method, which considers the influence of hybrid power system power source configuration on energy management control results, optimizes the parameter matching process and multi-stack fuel cell hybrid power system energy management control process as a whole, reduces fuel consumption and fuel cell performance degradation, ensures the performance coordination of multi-stack fuel cells, prolongs the service life of multi-stack fuel cells, and improves the durability of the system. In addition, this method also helps to reduce the cost of power source configuration.
[0008] To achieve the above purpose, the technical scheme adopted by the present application is: a hybrid power city regional motor train unit configuration and control deep collaborative optimization method, which is used for the fuel cell hybrid power system of a city regional motor train unit, including a multi-stack fuel cell system, a battery system, a direct current bus, a signal collection system and an energy management and control system; wherein the multi-stack fuel cell system is composed of a plurality of fuel cell monomers, each fuel cell monomer is connected to a unidirectional direct-direct converter and then connected in parallel with the remaining monomers to access the direct current bus; the battery is connected to a bidirectional direct-direct converter and then connected to the direct current bus; the signal collection system is used to collect the output voltage, output current and other information of the multi-stack fuel cell and the battery, and to provide a basis for subsequent energy management and control system calculation of hydrogen consumption, performance degradation and the like of the fuel cell and the battery; the energy management and control system is designed based on a DSP controller, the corresponding program of the energy management is loaded to distribute the demand power of the city regional motor train unit to the multi-stack fuel cell and the battery, and the power output of each power source is indirectly controlled through the control of the corresponding direct-direct converter.
[0009] Based on the above fuel cell hybrid power system for city regional motor train unit, the hybrid power city regional motor train unit configuration and control deep collaborative optimization method comprises the following steps:
[0010] S100, according to the design parameters of the city regional motor train unit and the basic conditions of the operation line, analyzing the constraint conditions that the motor train unit power system needs to meet, and constructing a constraint penalty model of the system;
[0011] S200, analyzing and establishing a hydrogen consumption model of the system, a performance attenuation model of the fuel cell and the battery, a performance coordination evaluation model of the multi-stack fuel cell system, a fuel cell output power fluctuation limitation model, and a battery state of charge constraint model, and constructing a state-action pair value function of the system based on the above models;
[0012] S300, under the physical constraints of the system operation, estimating the value of the state-action pair of the system and reducing the strategy entropy by using a flexible action-evaluation algorithm based on direction guidance, constructing a combined value to evaluate the pros and cons of the action by the value of the state-action pair and the strategy entropy, and finally selecting the optimal output power distribution scheme of the system at this moment until the end of the system operation process, thereby completing the whole process energy management;
[0013] S400, analyzing the constraint conditions met by the system, the parameter configuration cost of the system, and the energy management control result by using a three-objective non-dominated sorting genetic algorithm with a scaling factor, obtaining the parameter configuration non-dominated condition of the system according to the three-objective non-dominated sorting rule, and finally determining the final parameter matching scheme according to different expected targets.
[0014] Further, in the step S100, the constraint conditions of the city regional motor train unit power system are analyzed, and a constraint penalty model of the system is constructed, including the steps of:
[0015] S101, extracting the parameters for working condition construction in the design parameters of the city regional motor train unit; extracting the route parameters in the operation route of the city regional motor train unit; performing stress analysis on the city regional motor train unit in the operation process, performing traction calculation on multiple special working conditions of the city regional motor train unit according to the stress analysis result, obtaining the maximum power and maximum energy consumption under each working condition, and further comparing to obtain the maximum power and maximum energy consumption required in the operation process of the city regional motor train unit;
[0016] S102, power and energy constraint: the output power and energy consumption of the hybrid power system should meet the maximum power and maximum energy consumption required in the operation process of the city regional motor train unit;
[0017] S103, mass and volume constraint: the hybrid power system should meet the requirements of the axle load and space volume of the motor train unit;
[0018] S104, bus voltage constraint: ensuring that the maximum working voltage of the fuel cell is lower than the rated voltage of the DC bus; at the same time, the maximum output voltage of the battery is lower than the rated voltage of the DC bus;
[0019] S105, economic condition constraint: the purchased hybrid power system should not exceed the maximum budget;
[0020] S106, based on the analysis of the above constraint conditions, the power system configuration each does not satisfy a constraint condition, then the penalty value adds 1, when the power system configuration satisfies all the constraint conditions, the penalty value is 0, the smaller the penalty value is the better.
[0021] Further, in the step S200:
[0022] The hydrogen consumption model of the system during operation is represented as:
[0023]
[0024] Wherein,
[0025] In the formula, C fc is the hydrogen consumption of the multi-stack fuel cell system, which is obtained by adding the hydrogen consumptions of the fuel cell monomers; k is the state of charge correction coefficient of the battery; C bat is the equivalent hydrogen consumption of the battery; n is the number of fuel cell monomers in the multi-stack fuel cell system; m fc is the hydrogen consumption of the multi-stack fuel cell system; m bat is the equivalent hydrogen consumption of the battery system; C fci is the hydrogen consumption of the fuel cell monomer;
[0026] The fuel cell performance attenuation model is:
[0027] D fc =D on / off +D low +D high +D chg ;
[0028] Wherein, D on / off , D low , D high and D chg respectively represent the performance loss caused by the start-stop working condition, high-power operation, low-power operation and variable load working condition;
[0029] The multi-stack fuel cell performance coordination evaluation model is:
[0030] m vs =k vs V s ;
[0031] Wherein,
[0032]
[0033] In the formula, k VS is the multi-stack fuel cell performance coordination coefficient; D fci is the real-time attenuation degree of the i th fuel cell; R fci_initrepresents the initial remaining performance of the i th fuel cell before being put into use; L fc represents the coordinated degradation average of fuel cell performance; V s represents the performance coordination coefficient of variation of multiple stacks of fuel cells; σ represents the standard deviation of performance coordination of multiple stacks of fuel cells; μ represents the expectation of fuel cell performance coordination degradation period;
[0034] The fuel cell output power fluctuation limit model is:
[0035]
[0036] wherein k Dfc represents the fuel cell performance degradation coefficient; P fci,t represents the fuel cell output power at time t; P fci,t-1 represents the fuel cell output power at time t-1;
[0037] The battery state of charge constraint model is:
[0038] m soc =k soc (SOC-SOC init );
[0039] wherein k soc represents the battery state of charge constraint coefficient; SOC represents the state of charge of the battery; SOC init represents the set target state of charge of the battery.
[0040] Further, in order to quantitatively evaluate the performance degradation condition of the fuel cell, the remaining performance of the fuel cell needs to be evaluated, and the fuel cell remaining performance evaluation formula is:
[0041]
[0042] wherein ΔU rated_max is the maximum voltage drop allowed when the fuel cell is operated at rated current, defining that when the rated output voltage of the fuel cell drops by 10%, it is considered as the end of its life; ΔU rated_degraded is the voltage drop of the fuel cell due to performance degradation; U rated_init is the rated voltage of the fuel cell.
[0043] Further, the state-action pair value function of the system includes fuel cell and battery hydrogen consumption, fuel cell output power fluctuation constraint, battery state of charge constraint, and multiple stacks of fuel cell performance coordination constraint;
[0044] The system comprehensive evaluation function is:
[0045] J=m fc +m bat +mvs +m Dfc +m soc .
[0046] Furthermore, in step S300, a flexible action-evaluation algorithm based on direction guidance is used to estimate the value of the system state-action pair and reduce the policy entropy. The flexible action-evaluation algorithm based on direction guidance includes an actor network and four critic networks, namely, state value estimation v and Target v networks; action-state value estimation Q0 and Q1 networks; the input of the actor network is state s t , the output is the probability of action π(a t |s t ), the input of the critic network is the state, the output of the vcritic network is the estimate of the state value v(s), and the output of the Qcritic network is the estimate of the action-state value q(s,a).
[0047] Furthermore, before the direction-guided flexible action-evaluation algorithm is applied to energy management, the network needs to be trained, including the following steps:
[0048] (1) Generate experience pool: a state s is known t , we get the probability of all actions π(a|s t ), and then obtain action a by probability sampling t , then a t Input into the environment and get s t+1 and r t+1 , so you get an experience: (s t ,a t ,s t+1 ,r t+1 ), and then put the experience into the experience pool. When training the network, a batch of experiences is selected from the experience pool to eliminate the strong correlation between each experience;
[0049] (2) Qcritic network update: extract data from the experience pool (s t ,a t ,s t+1 ,r t+1 ) The Qcritic network is updated. Based on the optimal Bellman equation, the output of the Target vcritic network is used for true value estimation. Together with the output of the Qcritic network before the update as the predicted value estimation, the mean squared error loss is constructed as the loss function to train the Qcritic network;
[0050] (3) vcritic network update: extract data from the experience pool (s t ,at t+1 t+1 ) update the vcritic network, use the state value estimation with entropy as the true value of the output of the vcritic network, use the output of the vcritic network as the predicted value to construct the mean square error loss as the loss function, and train the vcritic network;
[0051] (4) actor network update: combine the probability of the policy taking action a t in state s t π(a t |s t ), the action-state pair value q(s t ,a t ), the temperature adjustment coefficient a, the state value v(s t ), etc. to construct a combined value as a loss function to perform gradient descent training on the actor;
[0052] (5) repeat the above steps until the network converges.
[0053] Further, the flexible action-evaluation algorithm based on direction guidance guides the network training process through three behaviors of real-time data cleaning, evolutionary direction sorting, and adaptive temperature coefficient adjustment based on the experience pool, including steps:
[0054] (1) Real-time data cleaning of experience pool: instead of processing the entire experience pool, the real-time extracted experience is processed. In the real-time training process of the network, the selected experience for training needs to be identified and processed according to the constraint conditions in the urban motor train operation process, and the data that does not meet the constraint conditions is removed from the experience pool, and experience is supplemented and extracted from the experience pool;
[0055] (2) Evolutionary direction sorting: based on the error between the predicted value and the true value during sample training, the important samples are sorted after normalization;
[0056] (3) Adaptive temperature adjustment coefficient: the average error of the training experience is output after fuzzy control to output the temperature adjustment coefficient, so that the temperature adjustment coefficient is adaptively modified as the training process proceeds.
[0057] Further, the physical constraints of the system operation include the maximum and minimum limits of the output power of the fuel cell and the battery, the maximum and minimum limits of the state of charge of the battery, the fuel cell output power change rate not exceeding 12.5% of its rated power, and the power source output power and load power matching.
[0058] Further, in the step S400, the role of the scaling factor is to gradually tighten the limit of the parameter matching constraint condition along with the iteration process, which helps to expand the search range in the initial stage and find the solution satisfying the constraint condition in the later iteration stage;
[0059] A three-objective fast non-dominated sorting rule needs to be constructed to meet the constraint condition, the system parameter configuration cost and the energy management control result, a scaling factor ε is defined, the maximum value of the Γ function is 7, therefore the initial value of ε is set to 6, and the rule is constructed as follows:
[0060] (1) If the Γ function value of x1 in the two parameter configurations x1 and x2 is less than or equal to the scaling factor ε, and the Γ function value of x2 is greater than the scaling factor ε, then x1 dominates x2;
[0061] (2) When the Γ function values of x1 and x2 are both greater than the scaling factor ε, if the Γ function value of x1 is less than the Γ function value of x2, then x1 dominates x2;
[0062] (3) When the Γ function values of x1 and x2 are both less than or equal to the scaling factor ε, then the dominance of x1 and x2 is determined according to the system parameter configuration cost and the energy management control result: when there is a configuration between the two that makes the system parameter configuration cost and the energy management control result the same as the other and one of them is better, then the configuration dominates the other; If x1 dominates x2, then the relationship is expressed as:
[0063]
[0064] n represents the nth objective function; f1(x) and f2(x) represent the system parameter configuration cost and the energy management result corresponding to the x parameter configuration, respectively.
[0065] The beneficial effects of the technical scheme are as follows:
[0066] The application discloses a kind of hybrid power city area motor train unit configuration and intelligent control depth collaborative optimization method.The method fully considers the influence of power source configuration of hybrid power system on energy management control result, the parameter matching process and the energy management control process of multi-stack fuel cell hybrid power system are regarded as a whole to optimize, the value of system state-action pair is estimated and the strategy entropy is reduced by using flexible action-evaluation algorithm based on direction guide, the merits of action are evaluated by the size of combination value constructed from the value of state-action pair and strategy entropy, the optimal output power distribution scheme of system is selected, and the effective distribution of city area motor train unit demand power is completed;The constraint condition satisfied by system, system parameter configuration cost and energy management control result are analyzed by using three-objective non-dominated sorting genetic algorithm with scaling factor, three-objective non-dominated sorting rules are set, the scaling factor is continuously reduced in solving process to make system satisfy constraint condition, the non-dominated condition of parameter configuration solution of system is obtained, and finally the final parameter matching scheme is determined according to different expected target, and the depth collaborative optimization of parameter matching and energy management control of fuel cell hybrid power city area motor train unit is completed.The method can improve the performance coordination of multi-stack fuel cell, effectively prolong the life of multi-stack fuel cell system while reducing hydrogen consumption, improve system durability, and also help to reduce the configuration cost of fuel cell hybrid power system.
[0067] The hybrid power city area motor train unit configuration and intelligent control depth collaborative optimization method provided by the application builds a system state-action pair value function minimum optimization target, which comprehensively considers multiple system indicators, including system hydrogen consumption, fuel cell performance attenuation, multi-stack fuel cell performance coordination, fuel cell output power fluctuation and battery state of charge and other factors, to realize the overall optimization of system operation.
[0068] The application considers the integrity of multi-stack fuel cell system, and includes the multi-stack fuel cell performance coordination into one of the optimization targets, to effectively reduce the situation that any fuel cell in the multi-stack fuel cell system is excessively used or prematurely decays to the end of life.The life of the multi-stack fuel cell reaches the end at the same time, which makes the life of the multi-stack fuel cell system no longer limited by any single stack, effectively prolongs the durability of the system, and improves the stability of system operation. BRIEF DESCRIPTION OF DRAWINGS
[0069] Figure 1 It is a flowchart of the hybrid power city area motor train unit configuration and control depth collaborative optimization method of the application;
[0070] Figure 2 It is a schematic diagram of the fuel cell hybrid power system for city area motor train unit in the embodiment of the application. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.
[0072] In this embodiment, see Figure 2 As shown, the present invention proposes a method for deep collaborative optimization of the configuration and control of a hybrid urban EMU. The fuel cell hybrid system for the urban EMU includes a multi-stack fuel cell system, a battery system, a DC bus, a signal collection system, and an energy management and control system; wherein the multi-stack fuel cell system is composed of a plurality of fuel cell monomers, each fuel cell monomer is connected to a unidirectional DC-DC converter and then connected to the DC bus in parallel with the remaining monomers; the battery is connected to a bidirectional DC-DC converter and then connected to the DC bus; the signal collection system is used to collect information such as the output voltage and output current of the multi-stack fuel cells and batteries, providing a basis for the subsequent energy management and control system to calculate the hydrogen consumption, performance attenuation, etc. of the fuel cells and batteries; the energy management and control system is designed based on a DSP controller, and the corresponding energy management program is loaded to distribute the required power of the urban EMU to the multi-stack fuel cells and batteries, and indirectly controls the power output of each power source by controlling the corresponding DC-DC converter, taking into account the fuel economy and durability of the system.
[0073] This method fully considers the impact of the hybrid power system power source configuration on the energy management control results, and optimizes the parameter matching process and the multi-stack fuel cell hybrid power system energy management control process as a whole. First, based on the design parameters of the urban EMU and the basic conditions of the operating line, the constraints are analyzed to meet the system's dynamic performance, safe operation, and economic budget. A constraint penalty model for the system is constructed using symbolic functions. A state-action pair value function is established by fully considering the hydrogen consumption of fuel cells and batteries, fuel cell performance degradation, battery state-of-charge constraints, fuel cell output power fluctuation limits, and the performance coordination of multiple fuel cell stacks. A flexible action-evaluation algorithm based on direction guidance is used to estimate the value of the system's state-action pairs and reduce the policy entropy. The combination of the state-action pair value and the policy entropy is used to evaluate the performance of the actions. Finally, the optimal output power allocation scheme is selected to achieve energy management control. A three-objective non-dominated sorting genetic algorithm with a scaling factor is used to analyze the system's constraints, system parameter configuration costs, and energy management control results. A three-objective non-dominated sorting rule is set. The scaling factor is continuously reduced during the solution process to ensure that the system meets the constraints. The non-dominated solution of the system's parameter configuration solution is obtained. Finally, the final parameter matching scheme is determined based on different expected goals, completing the deep collaborative optimization of the urban EMU parameter matching and the energy management control of the multi-stack fuel cell hybrid system.
[0074] Furthermore, during the implementation of the proposed method, the operation of the fuel cell hybrid system is subject to constraints imposed by the power source's output characteristics and physical conditions. These include limitations on the output power of the fuel cell and battery, limitations on the battery's state of charge range, limitations on the rate of change of fuel cell power, and limitations on matching the power source's output power with the load. The proposed method must be applied while satisfying these constraints.
[0075] Based on the above fuel cell hybrid system for urban EMU, the hybrid urban EMU configuration and control deep collaborative optimization method is proposed. Figure 1 As shown, the steps include:
[0076] Based on the design parameters of the urban EMU and the basic conditions of the operating line, S100 analyzes the constraints that the EMU power system needs to meet and builds a constraint penalty model for the system;
[0077] S200 analyzes and establishes a system hydrogen consumption model, a fuel cell and battery performance degradation model, a multi-stack fuel cell system performance coordination evaluation model, a fuel cell output power fluctuation limitation model, and a battery state of charge constraint model. Based on the above models, the system's state-action value function is constructed;
[0078] Under the physical constraints of system operation, S300 uses a flexible action-evaluation algorithm based on direction guidance to estimate the value of system state-action pairs and reduce policy entropy. The value of the state-action pair and the policy entropy are used to construct a combined value to evaluate the quality of the action. Finally, the optimal output power allocation scheme for the system at that moment is selected until the system operation ends, completing the whole process of energy management.
[0079] S400 uses a three-objective non-dominated sorting genetic algorithm with a scaling factor to analyze the constraints satisfied by the system, the system parameter configuration cost, and the energy management control results. According to the three-objective non-dominated sorting rules, the system parameter configuration non-domination situation is obtained, and finally the final parameter matching solution is determined according to different expected goals.
[0080] The specific operation of the deep collaborative optimization method of hybrid urban EMU configuration and intelligent control includes three steps.
[0081] The first step is to determine the fuel cell hybrid system parameter matching constraints, as follows:
[0082] S101, power and energy consumption calculation, extract the parameters in the design parameters of the city regional motor train unit for the construction of working conditions, such as load, maximum continuous traction, maximum speed, etc.; extract the route parameters in the running line of the city regional motor train unit, such as slope, line length, maximum climbing distance, average station spacing parameters; force analysis is carried out on the city regional motor train unit in the running process, according to the force analysis result, the maximum power and maximum energy consumption of the city regional motor train unit in the running process are obtained through traction calculation of multiple special working conditions of the city regional motor train unit, and the maximum power and maximum energy consumption required in the running process of the city regional motor train unit are further compared;
[0083] S102, power and energy constraint: the output power and energy consumption of the hybrid power system should meet the maximum power and maximum energy consumption required in the running process of the city regional motor train unit;
[0084] The power and energy constraints that the hybrid power system needs to meet are as follows:
[0085]
[0086] n fc and n bat are the number of fuel cells and storage batteries, respectively, which are independent variables in the optimization process; P fc and P bat are the rated output power of fuel cells and storage batteries, respectively; η fcDC , η batDC and η DC / AC are the efficiencies of one-way DC / DC converter, bidirectional DC / DC converter and inverter, respectively; E bat is the energy storage of a single storage battery; P loadmax and E loadmax represent the maximum power and maximum energy consumption required in the running process of the city regional motor train unit, respectively.
[0087] S103, construct quality and volume constraints: when the city regional motor train unit is equipped with a hybrid power system, the actual conditions need to be considered; the axle load and internal space of the motor train unit are limited, so the hybrid power system should meet the requirements of the axle load and space volume of the motor train unit;
[0088] The constraints of the mass and volume of the power source of the city regional motor train unit are as follows:
[0089]
[0090] The additional mass K mfc and volume coefficient K Vfc of the fuel cell; the additional mass K mbat and volume coefficient K Vbat of the storage battery; the mass M fc and volume V fc; the mass M of the battery monomer bat and the volume V bat ; the maximum mass M of the power system allowed by the EMU max , the maximum volume V max .
[0091] S104, bus voltage constraint: ensure that the maximum operating voltage of the fuel cell is lower than the rated voltage of the DC bus; at the same time, the maximum output voltage of the battery is lower than the rated voltage of the DC bus;
[0092] The bus voltage constraint of the city EMU is as follows:
[0093]
[0094] U batmax and U fcmax are the maximum output voltages of the battery and the fuel cell respectively.
[0095] S105, economic condition constraint: the purchased hybrid power system should not exceed the maximum budget;
[0096] The economic condition constraint of the city EMU is as follows:
[0097] C tra < C sys ;
[0098] C tra and C sys are the cost of the purchased hybrid power system and the maximum purchase budget of the system respectively.
[0099] S106, determine the parameter matching feasible region: based on the analysis of the above constraint conditions, the power system configuration is punished by 1 for each constraint condition not met, and the punishment value is 0 when the power system configuration meets all the constraint conditions, and the smaller the punishment value is, the better.
[0100] The constraint penalty model of the system is constructed as follows:
[0101] Γ=sign(min(P sys -P tra ,0))+sign(min(E sys -E tra ,0))+sign(min(M tra -M sys ,0))+sign(min(V tra -V sys ,0))+sign(min(U batmax -U bus ,0))+sign(min(U fcmax -U bus,0))+sign(min(C sys -C tra ,0));
[0102] P sys 、E sys 、M sys 、V sys and C sys They represent the maximum required power, maximum energy consumption, maximum allowable mass, maximum allowable volume and maximum purchase budget of the system respectively; P tra 、E tra 、M tra 、V tra and C tra They represent the maximum required power, maximum energy consumption, maximum mass, maximum volume and purchase budget under the current power system configuration; U fcmax 、U batmax and U bus They represent the maximum voltage of the fuel cell, the maximum voltage of the battery, and the bus voltage respectively.
[0103] The second step is to construct a state-action value function for the urban EMU and use a flexible action-evaluation algorithm based on direction guidance to obtain the optimal output power allocation plan for the system. The implementation process of this step is as follows:
[0104] (1) Constructing a hydrogen consumption model for the system: This includes the hydrogen consumption of the fuel cell and the equivalent hydrogen consumption of the battery. Since the hybrid system cannot be supplied with energy externally during operation, all energy must be obtained by consuming hydrogen through the fuel cell. Therefore, the battery's electrical energy consumption must be equivalent to the fuel cell's hydrogen consumption. The hydrogen consumption model during the system's operation is expressed as:
[0105]
[0106] in,
[0107] Where C fc is the hydrogen consumption of the multi-stack fuel cell system, which is obtained by adding the hydrogen consumption of the fuel cell monomers; k is the charge state correction coefficient of the battery; C bat is the equivalent hydrogen consumption of the battery; n is the number of fuel cell monomers in the multi-stack fuel cell system; m fc is the hydrogen consumption of the multi-stack fuel cell system; m bat is the equivalent hydrogen consumption of the battery system; C fci is the hydrogen consumption of a fuel cell.
[0108] (2) Construct a performance degradation model of fuel cell: To control the durability of fuel cell, it is necessary to evaluate the performance degradation of fuel cell in real time, which is caused by start-stop operation, high-power operation, low-power operation and variable load operation. The performance degradation model of fuel cell is as follows:
[0109] D fc on / off +D low +D high +D chg ;
[0110] Wherein, D on / off , D low , D high and D chg represent the performance loss caused by start-stop operation, high-power operation, low-power operation and variable load operation, respectively.
[0111] (3) Construct a formula for evaluating the remaining performance of fuel cell: In order to quantitatively evaluate the performance degradation of fuel cell, it is necessary to evaluate the remaining performance of fuel cell. The formula for evaluating the remaining performance of fuel cell is as follows:
[0112]
[0113] Wherein, ΔU rated_max is the maximum pressure drop allowed when the fuel cell operates at rated current, which defines the end of life when the rated output voltage of the fuel cell drops by 10%; ΔU rated_degraded is the pressure drop of the fuel cell caused by performance degradation; U rated_init is the rated voltage of the fuel cell.
[0114] (4) Construct a performance coordination evaluation model of multi-stack fuel cell: In order to control any one of the multi-stack fuel cell not to be used excessively or to decay to the end of life prematurely to prolong the service life of the multi-stack fuel cell system, the performance coordination evaluation model of multi-stack fuel cell is as follows:
[0115] m vs = k vs V s ;
[0116] Wherein,
[0117]
[0118] In the formula, k VS is the performance coordination coefficient of multi-stack fuel cell; D fci is the real-time degradation degree of the i-th fuel cell; R fci_init represents the initial remaining performance of the i-th fuel cell before use; Lfc is the fuel cell performance coordination attenuation average degree; V s is the fuel cell performance coordination variability coefficient of multiple stacks; σ is the fuel cell performance coordination standard deviation; μ is the fuel cell performance coordination expectation.
[0119] (5) A fuel cell output power fluctuation limiting model is constructed: frequent fluctuation of fuel cell output power is an important reason for performance attenuation of the fuel cell, therefore, limiting the output power fluctuation of the fuel cell can effectively reduce the performance attenuation of the fuel cell, and the fuel cell output power fluctuation limiting model is:
[0120]
[0121] wherein, k Dfc is the fuel cell performance attenuation coefficient; P fci,t is the fuel cell output power at t moment; P fci,t-1 is the fuel cell output power at t-1 moment.
[0122] (6) A battery state of charge constraint model is constructed: the state of charge of the battery will affect the dynamic performance of the system when the system runs at high power, and keeping the state of charge of the battery within a small range is beneficial to the system to have enough power as support when facing possible high power operation, and the battery state of charge constraint model is:
[0123] m soc = k soc (SOC-SOC init );
[0124] wherein, k soc is the battery state of charge constraint coefficient; SOC represents the state of charge of the battery; SOC init represents the set target state of charge of the battery.
[0125] (7) A state-action pair value function of the system is constructed: considering the overall optimization of the system, the state-action pair value function of the system includes fuel cell and battery hydrogen consumption, fuel cell output power fluctuation constraint, battery state of charge constraint and multiple stack fuel cell performance coordination constraint, and the state-action pair value function of the system is:
[0126] J = m fc +m bat +m vs +m Dfc +m soc .
[0127] (8) Analysis solution constraints: the fuel cell hybrid system needs to be subjected to physical constraints during actual operation, including the maximum and minimum limits on the output power of the fuel cell and the battery, the maximum and minimum limits on the state of charge of the battery, the rate of change of the output power of the fuel cell not exceeding 12.5% of its rated power, and the matching of the output power of the power source and the load power, which can be expressed as:
[0128]
[0129] (9) Training the network:
[0130] The direction-guided flexible action-evaluation algorithm is used to estimate the value of the system state-action pair and reduce the policy entropy, and the direction-guided flexible action-evaluation algorithm includes an actor network and four critic networks, which are state value estimation v and Target v networks, action-state value estimation Q0 and Q1 networks, the input of the actor network is the state s t , and the output is the probability of action t |s t , the input of the critic network is the state s t , and the output of the v critic network is the estimation of the state value v(s), and the output of the Q critic network is the estimation of the action-state pair value q(s, a); in order to make the policy more balanced and diversified, the concept of entropy is added in the direction-guided flexible action-evaluation algorithm to measure the index of uncertainty of the policy.
[0131] Before the direction-guided flexible action-evaluation algorithm is applied to energy management, the network needs to be trained, including the following steps: (1) generating an experience pool: given a state s t , the probability of all actions t is obtained through the actor network, then the action a t is obtained by sampling according to the probability, then a t is input into the environment to obtain s t+1 and r t+1 , so as to obtain an experience: (s t , a t , s t+1 , r t+1 ), then the experience is put into the experience pool, and when training the network, a batch of experiences is selected from the experience pool to eliminate the strong correlation between each experience; (2) Q critic network update: data (s t , a t , s t+1 , r t+1) Update the Qcritic network. Based on the optimal Bellman equation, the output of the Target vcritic network is used to estimate the true value. Together with the output of the Qcritic network before the update as the predicted value estimate, the mean square error loss is constructed as the loss function to train the Qcritic network; (3) vcritic network update: take data (s) from the experience pool t ,a t ,s t+1 ,r t+1 ) Update the vcritic network, use the entropy-containing formula to estimate the state value as the true value of the vcritic network output, use the output of the vcritic network as the predicted value to construct the mean square error loss as the loss function, and train the vcritic network; (4) actor network update: Combine the strategy in state s t Next take action a t The probability π(a t |s t ), action-state pair value q(s t ,a t ), temperature adjustment coefficient α, state value v(s t ) etc. to construct the combined value as the loss function to perform gradient descent training on the actor; (5) Repeat the above steps until the network converges.
[0132] The flexible action-evaluation algorithm based on direction guidance guides the network training process through three behaviors: real-time data cleaning in the experience pool, evolutionary direction sorting, and adaptive temperature coefficient adjustment. This ensures that the algorithm has good search capabilities while improving the convergence of the algorithm. The steps include:
[0133] ① Real-time data cleaning of the experience pool: There is a lot of experience about a particular system state, and the data quality needs to be improved to ensure the accuracy and consistency of the data. In order to improve the speed and efficiency of data cleaning, instead of processing the entire experience pool, the real-time extracted experience is processed. During the real-time training of the network, the selected experience for training needs to be identified and processed according to the constraints of the urban EMU operation process. Data that does not meet the constraints is removed from the experience pool, and additional experience is extracted from the experience pool to ensure sufficient training data.
[0134] ② Sorting evolutionary directions: After determining the evolutionary direction during algorithm training, it is necessary to converge it as quickly as possible to improve the convergence speed of the algorithm. To this end, it is necessary to adjust the algorithm's search capabilities and focus the algorithm training process more on the convergence process. Therefore, based on the error between the predicted value and the true value during sample training, the important samples are normalized and sorted, so that the experience with small error has a greater probability of being adopted to accelerate network convergence;
[0135] ③Adaptive temperature adjustment coefficient: The temperature adjustment coefficient is used to balance the search ability and convergence ability of the algorithm. The training of the algorithm needs to ensure sufficient search ability to find a more suitable convergence direction in the early stage, and needs to accelerate the convergence of the algorithm as much as possible in the later stage. Therefore, the fixed temperature adjustment coefficient is not conducive to balancing the search ability and convergence ability. The average error of training experience is output after fuzzy control to output the temperature adjustment coefficient, so that the temperature adjustment coefficient is adaptively modified with the training process.
[0136] (10) Flexible action-evaluation algorithm based on direction guidance is used to solve:
[0137] The third link is the deep collaborative optimization of parameter matching and energy management control of multi-stack fuel cell hybrid power system. A three-objective non-dominated sorting genetic algorithm with a scaling factor is used to analyze the constraints satisfied by the system, the parameter configuration cost of the system, and the energy management control result, to obtain the non-dominated situation of the parameter configuration of the system. Finally, according to different expected targets, such as focusing on energy management results or focusing on power source configuration cost or both weighted as a judgment standard, the final parameter matching scheme is determined; the role of the scaling factor is to gradually tighten the constraints of parameter matching in the iteration process, which helps to expand the search range in the early stage and find solutions that meet the constraint conditions in the later stage.
[0138] When using a three-objective non-dominated sorting genetic algorithm with a scaling factor to solve, a three-objective fast non-dominated sorting rule for the constraints satisfied by the system, the parameter configuration cost of the system, and the energy management control result needs to be constructed. A scaling factor ε is defined. Since the maximum value of the Γ function is 7, the initial value of ε is set to 6. The rule is constructed as follows:
[0139] (1) If the Γ function value of x1 is less than or equal to the scaling factor ε in two parameter configurations x1 and x2, and the Γ function value of x2 is greater than the scaling factor ε, then x1 dominates x2;
[0140] (2) When the Γ function values of x1 and x2 are both greater than the scaling factor ε, if the Γ function value of x1 is less than the Γ function value of x2, then x1 dominates x2;
[0141] (3) When the Γ function values of x1 and x2 are both less than or equal to the scaling factor ε, then the dominance of x1 and x2 is determined according to the system parameter configuration cost and the energy management control result: when there is a configuration between them that makes the system parameter configuration cost and the energy management control result the same as the other one and one of them is better, then the configuration dominates the other one. If x1 dominates x2, then the relationship is expressed as:
[0142]
[0143] n represents the nth objective function; f1(x) and f2(x) represent the parameter configuration cost and energy management result of the corresponding system under the x parameter configuration, respectively.
[0144] The hybrid regional rail motor train unit configuration and intelligent control deep collaborative optimization method can fully consider the influence of parameter matching result on energy management control result, comprehensively optimize the operation of the system, balance the difference between the multiple fuel cell systems to prolong the service life of the multiple fuel cell systems, and reduce the operation cost of the system. In addition, the method also maintains the state of charge of the battery in a small range of fluctuation, which is beneficial to the system to have sufficient power to support the system operation when facing possible high power demand.
[0145] The above shows and describes the basic principles and main features of the present application and the advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above examples, and the above examples and descriptions in the specification are only to illustrate the principles of the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for deep collaborative optimization of hybrid urban EMU configuration and control, characterized in that: The fuel cell hybrid system for urban EMUs includes a multi-stack fuel cell system, a battery system, a DC bus, a signal collection system, and an energy management and control system. The multi-stack fuel cell system consists of multiple fuel cell units, each of which is connected to a unidirectional DC-DC converter and then connected to the DC bus in parallel with the remaining units. The battery is connected to a bidirectional DC-DC converter and then connected to the DC bus. The signal collection system is used to collect output voltage and output current information from multiple fuel cell and battery stacks, providing a basis for subsequent energy management and control system calculations of hydrogen consumption and performance degradation of fuel cells and batteries. The energy management and control system is designed based on a DSP controller, which loads the corresponding energy management program to distribute the required power of the urban EMU to multiple fuel cell and battery stacks, and indirectly controls the power output of each power source by controlling the corresponding DC-DC converter. Based on the above-mentioned fuel cell hybrid power system for urban EMUs, a hybrid urban EMU configuration and control deep collaborative optimization method includes the following steps: Based on the design parameters of the urban EMU and the basic conditions of the operating line, S100 analyzes the constraints that the EMU power system needs to meet and builds a constraint penalty model for the system; S200 analyzes and establishes a system hydrogen consumption model, a fuel cell and battery performance degradation model, a multi-stack fuel cell system performance coordination evaluation model, a fuel cell output power fluctuation limitation model, and a battery state of charge constraint model. Based on the above models, the system's state-action value function is constructed; Under the physical constraints of system operation, the S300 uses a flexible action-evaluation algorithm based on direction guidance to estimate the value of system state-action pairs and reduce policy entropy. The value of the state-action pair and the policy entropy are used to construct a combined value to evaluate the quality of the action. Finally, the optimal output power allocation scheme for the system at the current moment is selected until the system operation ends, completing the whole process of energy management. The flexible action-evaluation algorithm based on direction guidance guides the network training process through three behaviors: real-time data cleaning in the experience pool, evolutionary direction sorting, and adaptive temperature coefficient adjustment. The steps include: (1) Real-time data cleaning of the experience pool: Instead of processing the entire experience pool, the real-time extracted experience is processed. During the real-time training of the network, the selected experience for training needs to be identified and processed according to the constraints of the urban train operation process, and the data that does not meet the constraints is removed from the experience pool, and additional experience is extracted from the experience pool; (2) Evolutionary direction sorting: Based on the error between the predicted value and the true value during sample training, the important samples are sorted after normalization; (3) Adaptive temperature adjustment coefficient: The temperature adjustment coefficient is output after fuzzy control using the average error of training experience, so that the temperature adjustment coefficient can be adaptively modified as the training process progresses; S400 uses a three-objective non-dominated sorting genetic algorithm with a scaling factor to analyze the constraints satisfied by the system, the system parameter configuration cost, and the energy management control results. According to the three-objective non-dominated sorting rules, the system parameter configuration non-domination situation is obtained, and finally the final parameter matching solution is determined according to different expected goals.
2. A hybrid urban EMU configuration and control depth collaborative optimization method according to claim 1, characterized in that: In step S100, the constraints of the urban EMU power system are analyzed and a constraint penalty model of the system is constructed, including the following steps: S101, extracting parameters for operating condition construction from the design parameters of the urban EMU; extracting route parameters from the operating route of the urban EMU; Conduct stress analysis on the urban EMUs in operation. Based on the stress analysis results, perform traction calculations for multiple special operating conditions of the urban EMUs to obtain the maximum power and maximum energy consumption under each operating condition. Further comparison is performed to obtain the maximum power and maximum energy consumption required during the operation of the urban EMUs. S102, Power and Energy Constraints: The output power and energy consumption of the hybrid system should meet the maximum power and energy consumption required during the operation of the urban EMU; S103, Mass and Volume Constraints The hybrid system shall comply with the axle weight and space volume requirements of the EMU; S104, bus voltage constraint: ensure that the maximum operating voltage of the fuel cell is lower than the rated voltage of the DC bus; at the same time, the maximum output voltage of the battery is lower than the rated voltage of the DC bus; S105, Economic Constraint: The purchased hybrid system should not exceed the maximum budget; S106 , based on the analysis of the above constraints, a penalty value is increased by 1 for each constraint that the power system configuration fails to satisfy. When the power system configuration satisfies all constraints, the penalty value is 0. The smaller the penalty value, the better.
3. A hybrid urban EMU configuration and control depth collaborative optimization method according to claim 1, characterized in that: In step S200: The hydrogen consumption model during the operation of the system is expressed as: in, Where C fc is the hydrogen consumption of the multi-stack fuel cell system, which is obtained by adding the hydrogen consumption of the fuel cell monomers; k is the charge state correction coefficient of the battery; C bat is the equivalent hydrogen consumption of the battery; n is the number of fuel cell monomers in the multi-stack fuel cell system; m fc is the hydrogen consumption of the multi-stack fuel cell system; m bat is the equivalent hydrogen consumption of the battery system; C fci is the fuel cell single hydrogen consumption; The fuel cell performance attenuation model is: D fc =D on / off +D low +D high +D chg ; Among them, D on / off 、D low 、D high and D chg Respectively represent the performance loss caused by start-stop conditions, high power operation, low power operation and variable load conditions; The multi-stack fuel cell performance coordination evaluation model is: m vs =k vs In s ; in, Where k VS is the performance coordination coefficient of multiple fuel cell stacks; D fci is the real-time attenuation degree of the i-th fuel cell; R fci_init represents the initial residual performance of the i-th fuel cell before it is put into use; L fc Coordinate the average degree of attenuation for fuel cell performance; V s is the coefficient of variation of the performance coordination of multiple fuel cells; σ is the standard deviation of the performance coordination of multiple fuel cells; μ is the expected attenuation of the fuel cell performance coordination; The fuel cell output power fluctuation limitation model is: Among them, k Dfc is the fuel cell performance attenuation coefficient; P fci,t is the fuel cell output power at time t; P fci,t-1 is the fuel cell output power at time t-1; The battery state of charge constraint model is: m soc =k soc (SOC-SOC init ); Among them, k soc SOC is the battery state of charge constraint coefficient; SOC represents the battery state of charge; SOC init Indicates the set target state of charge of the battery.
4. A hybrid urban EMU configuration and control depth collaborative optimization method according to claim 3, characterized in that: In order to quantitatively evaluate the performance degradation of the fuel cell, it is necessary to evaluate the residual performance of the fuel cell. The residual performance evaluation formula of the fuel cell is: Among them, ΔU rated_max It is the maximum voltage drop allowed when the fuel cell is running at rated current. It is defined that when the rated output voltage of the fuel cell drops by 10%, it is considered as the end of its life. rated_degraded is the pressure drop caused by fuel cell performance degradation; U rated_init is the rated voltage of the fuel cell.
5. A hybrid electric vehicle configuration and control depth collaborative optimization method according to claim 4, characterized in that: The system's state-action value function includes fuel cell and battery hydrogen consumption, fuel cell output power fluctuation constraints, battery state of charge constraints, and multi-stack fuel cell performance coordination constraints; The comprehensive evaluation function of the system is: J=m fc +m bat +m vs +m Dfc +m soc 。 6. A hybrid electric vehicle configuration and control depth collaborative optimization method according to claim 1, characterized in that: In step S300, a flexible action-evaluation algorithm based on direction guidance is used to estimate the value of the system state-action pair and reduce the policy entropy. The flexible action-evaluation algorithm based on direction guidance includes an actor network and four critic networks, namely, the state value estimation v and Target v networks; the action-state value estimation Q0 and Q1 networks; the input of the actor network is the state s t , the output is the probability of action π(a t |s t ), the input of the critic network is the state, the output of the vcritic network is the estimate of the state value v(s), and the output of the Qcritic network is the estimate of the action-state value q(s,a).
7. A hybrid electric vehicle configuration and control depth collaborative optimization method according to claim 6, characterized in that: Before applying the direction-guided flexible action-evaluation algorithm to energy management, the network needs to be trained, including the following steps: (1) Generate experience pool: a state s is known t , we get the probability of all actions π(a|s t ), and then obtain action a by probability sampling t , then a t Input into the environment and get s t+1 and r t+1 , so you get an experience: (s t ,a t ,s t+1 ,r t+1 ), and then put the experience into the experience pool. When training the network, a batch of experiences is selected from the experience pool to eliminate the strong correlation between each experience; (2) Qcritic network update: extract data from the experience pool (s t ,a t ,s t+1 ,r t+1 ) The Qcritic network is updated. Based on the optimal Bellman equation, the output of the Target vcritic network is used for true value estimation. Together with the output of the Qcritic network before the update as the predicted value estimation, the mean squared error loss is constructed as the loss function to train the Qcritic network; (3) vcritic network update: extract data from the experience pool (s t ,a t ,s t+1 ,r t+1 ) Update the vcritic network, use the entropy-containing formula to estimate the state value as the true value of the vcritic network output, use the output of the vcritic network as the predicted value to construct the mean square error loss as the loss function, and train the vcritic network; (4) Actor network update: Combined strategy in state s t Next take action a t The probability π(a t |s t ), action-state pair value q(s t ,a t ), temperature adjustment coefficient α, state value v(s t ), construct the combined value as the loss function to perform gradient descent training on the actor; (5) Repeat the above steps until the network converges.
8. A hybrid urban EMU configuration and control depth collaborative optimization method according to claim 1, characterized in that: The physical constraints of system operation include maximum and minimum limits on the output power of fuel cells and batteries, maximum and minimum limits on the battery state of charge, the fuel cell output power change rate not exceeding 12.5% of its rated power, and matching of power source output power and load power.
9. A hybrid electric vehicle configuration and control depth collaborative optimization method according to claim 1, characterized in that: In step S400, the role of the scaling factor is to gradually tighten the restrictions on the parameter matching constraints as the iteration process progresses, which helps to expand the search range in the initial stage and find a solution that meets the constraints in the later stages of the iteration; It is necessary to construct a three-objective fast non-dominated sorting rule for the constraints satisfied by the system, the system parameter configuration cost, and the energy management control results. A scaling factor ε is defined. The maximum value of the Γ function is 7, so the initial value of ε is set to 6. The construction rule is: (1) If the value of the Γ function of x1 in two parameter configurations x1 and x2 is less than or equal to the scaling factor ε, and the value of the Γ function of x2 is greater than the scaling factor ε, then x1 dominates x2; (2) When the Γ function values of x1 and x2 are both greater than the scaling factor ε, if the Γ function value of x1 is less than the Γ function value of x2, then x1 dominates x2; (3) When the Γ function values of x1 and x2 are both less than or equal to the scaling factor ε, the dominance of x1 and x2 is determined based on the system parameter configuration cost and the energy management control result: when there is a configuration between the two that makes the system parameter configuration cost and the energy management control result the same as the other and one of them is better, then the current configuration dominates the other; if x1 dominates x2, then the relationship is expressed as: n represents the nth objective function; f1(x) and f2(x) represent the parameter configuration cost and energy management result of the system under the parameter configuration x, respectively.
Citation Information
Patent Citations
Heterogeneous and homogeneous group coevolution method for improving evolution ability of swarm robot
CN113485119A
Hybrid power multi-target layered energy management method for urban motor train unit
CN118003986A