Fuel cell automobile energy management method based on working condition identification and noise network exploration reinforcement learning

Through the method of building identification and noise network reinforcement learning based on working conditions, combining K-mean clustering and particle swarm algorithm to optimize the support vector machine, the multi-objective optimization problem of fuel cell vehicle energy management in complex working conditions is solved, and efficient energy utilization and fuel cell life are achieved.

CN120270120APending Publication Date: 2025-07-08ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510351685.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing fuel cell vehicle energy management methods are difficult to achieve multi-objective optimization under complex operating conditions, have poor adaptability, high energy consumption, and lack effective utilization of actual operating conditions, which affects the life of fuel cell and the maintenance of power battery charge state.

Method used

The method of building recognition based on working conditions and strengthened learning of noise networks is adopted, and typical working conditions are constructed by combining the K-mean clustering algorithm and short-stroke method. The reinforcement learning strategy for noise network exploration is designed, and the online working conditions are recognized through the particle swarm algorithm optimization support vector machine to realize the power distribution optimization of fuel cells and power batteries.

Benefits of technology

It significantly improves the energy utilization efficiency of fuel cell vehicles, reduces energy consumption, extends fuel cell life, and improves the overall performance of the vehicle under different driving conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120270120A_ABST
    Figure CN120270120A_ABST
Patent Text Reader

Abstract

The invention provides a fuel cell automobile energy management method based on working condition construction recognition and noise network reinforcement learning, which comprises the following steps: firstly, constructing typical working conditions by using a K-means clustering algorithm and a short-stroke method; then, taking the energy consumption, the service life and the state of charge maintenance of the fuel cell as instant return, developing a reinforcement learning energy management strategy based on a noise network to perform control optimization, and performing offline training to obtain optimal control parameters of various typical working conditions; and finally, designing a support vector machine working condition recognizer optimized based on a particle swarm optimization algorithm, carrying out online recognition on driving working conditions, and mapping optimal control parameters under typical working conditions to realize power distribution. According to the method, working condition construction, working condition recognition and reinforcement learning are fused into fuel cell automobile energy management control, the method has the advantages of high optimization precision and high adaptability, energy consumption can be effectively reduced, and the service life of the fuel cell can be effectively prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fuel cell vehicle energy management, and specifically is a fuel cell vehicle energy management method based on operating condition recognition and noise network exploration reinforcement learning. Background Art

[0002] As a zero-emission, high-efficiency system, the development of fuel cell vehicles is a key way to deal with environmental pollution and energy shortages. How to develop a reasonable energy management method to allocate energy between fuel cells, power batteries and motors under different driving conditions to give full play to their energy-saving potential is an important topic in the field of fuel cell vehicles, which is of great significance for improving the overall energy efficiency of vehicles and promoting environmentally friendly travel.

[0003] Some existing patents, such as the invention patent with patent number CN119348505A, have designed a whole vehicle energy management method for hydrogen fuel cell vehicles, which can accurately analyze the power demand of the whole vehicle and the power request of the fuel cell system, ensure that the power demand of the whole vehicle is met when the load change rate is not high, and avoid charging the power battery with excess power, resulting in increased hydrogen consumption and loss during charging and discharging, but the patent lacks exploration of the algorithm optimization potential. The invention patent with patent number CN118386949B builds a system model for the whole vehicle energy source management of fuel cell vehicles based on the driving power of fuel cell vehicles. By designing the Double Deep Q-network algorithm, under the premise of considering the power following effect, it solves multi-objective control problems including delaying the attenuation of the power source, reduces the hydrogen consumption while ensuring power following, and delays the aging of the fuel cell, but lacks effective use of the input information of the actual working conditions.

[0004] Some existing fixed rules or simple optimization algorithms are difficult to dynamically adapt to changing driving conditions, and the energy management method of fuel cell vehicles needs to consider multiple goals at the same time, such as energy consumption, fuel cell life, and battery state of charge (SOC) maintenance. Under complex working conditions, these goals may have a mutually restrictive relationship, which increases the complexity of energy management control. In addition, the driving conditions of fuel cell vehicles are highly uncertain and diverse. How to accurately identify the current working conditions and dynamically adjust the control strategy according to the working conditions is also a technical problem that needs to be solved urgently.

[0005] Therefore, how to develop an energy control algorithm for a fuel cell vehicle hybrid power system with multi-objective optimization, self-learning, and operating condition adaptability to increase energy utilization efficiency and extend fuel cell life is a difficulty that needs to be solved. Summary of the invention

[0006] The present invention proposes a fuel cell vehicle energy management method based on driving cycle construction, recognition, and noise network reinforcement learning. This algorithm integrates driving cycle construction, recognition, and reinforcement learning into the energy management control of fuel cell vehicles. By using the K-means clustering algorithm and the short-trip method, typical driving cycles are constructed. A reinforcement learning energy management strategy based on the noise network (NoisyNet-DQN) is developed for control optimization, and the optimal control parameters for various typical driving cycles are obtained through offline training. A driving cycle recognizer based on a support vector machine optimized by the particle swarm algorithm is designed for online recognition of driving cycles, and the optimal control parameters under typical driving cycles are mapped to achieve power distribution, thereby reducing energy consumption and extending the life of fuel cells.

[0007] The object of the present invention can be achieved through the following technical solutions:

[0008] A fuel cell vehicle energy management method based on driving cycle construction, recognition, and noise network reinforcement learning. The powertrain of the fuel cell vehicle consists of a fuel cell, a power battery, a unidirectional DC / DC converter, a bidirectional DC / DC converter, an inverter, and a drive motor. The fuel cell is connected to the DC bus through a unidirectional DC-DC converter, and the power battery is connected to the DC bus through a bidirectional DC-DC converter.

[0009] This energy management method specifically includes:

[0010] S1: Construct typical driving cycles based on the K-means algorithm and the short-trip method;

[0011] S2: According to the actual power distribution control process of the fuel cell vehicle, taking the instantaneous equivalent hydrogen consumption, the fuel cell degradation amount, and the maintenance of the state of charge (SOC) of the power battery as the immediate rewards, and the fuel cell power change as the action variable, a reinforcement learning energy management strategy based on noise network exploration (NoisyNet-DQN) is proposed for control optimization. Using four typical driving cycles as training scenarios, the optimal parameters corresponding to the NoisyNet-DQN strategy for each scenario are obtained respectively;

[0012] S3: Design a driving cycle recognizer based on a support vector machine (SVM) optimized by the particle swarm algorithm (Particle Swarm Optimization, PSO) for online recognition of driving cycles;

[0013] S4: Under actual driving conditions, according to the recognition results of the PSO-SVM, map the optimal parameters of the NoisyNet-DQN strategy corresponding to the current driving cycle. The energy management system distributes power between the fuel cell and the power battery, and then transmits it to the drive motor through the inverter to achieve online adaptive adjustment of the control algorithm parameters and improve the comprehensive performance of the vehicle.

[0014] Further, step S1 is specifically as follows:

[0015] S11: Divide the kinematic segments in combination with the standard working conditions, and select characteristic parameters to describe the kinematic segments;

[0016] S12: According to the characteristic parameters of each kinematic segment, randomly select a sample point as the initial clustering center, calculate the distance D(i) between the current sample i and the nearest clustering center, and then calculate the probability P(i) of the current sample as the next clustering center. When the random number rand is greater than P(i - 1) and less than P(i), select the current sample as the clustering center, and repeat the current step until K types of clustering centers are determined. Based on the Euclidean distance formula, the probability that the current sample is selected as the clustering center is:

[0017]

[0018] In the formula, N is the total number of samples.

[0019] S13: Calculate the distance from each sample to the K types of clustering centers, assign the sample to the cluster where the nearest clustering center is located, and calculate the mean value of each cluster, and set it as the new clustering center until the clustering center does not change;

[0020] S14: Based on the clustering analysis result, arrange each kinematic segment in sequence by the short stroke method to complete the construction of the typical working condition.

[0021] Further, step S2 is specifically as follows:

[0022] S21: Take the SOC of the power battery 2, the output power P fuel of the fuel cell 1, the demand power P dem and the vehicle speed V of the vehicle as the state variables S of the NoisyNet - DON strategy:

[0023] S = {SOC, P fuel , P dem , V}(2)

[0024] S22: Set the action variable as the change in the fuel cell power:

[0025] A = {-4, -3, -2, -1, 0, 1, 2, 3, 4}(3)

[0026] Set the immediate reward as the opposite of the instantaneous equivalent fuel consumption, the degradation of the fuel cell 1, and the maintenance of the SOC of the power battery 2:

[0027]

[0028] Wherein, α′, β′, and γ′ are the weight factors of the economic item, the degradation item of the fuel cell 1, and the SOC maintenance item of the power battery 2, respectively, and SOC ref is the reference value of the SOC of the power battery 2, is the maximum value of the instantaneous fuel cell degradation amount, is the instantaneous equivalent hydrogen consumption, and the subscripts min and max represent the minimum value and the maximum value; in addition, the state constraints during vehicle driving are set as:

[0029]

[0030] Wherein, P bat is the power of the fuel cell 1, and η DC / DC is the efficiency of the unidirectional DC / DC converter 3.

[0031] S23: The NoisyNetDQN method saves the data such as the state S, action a, immediate reward, and the next moment action generated during the interaction process into the experience pool, and conducts intelligent agent Q-network learning through mini-batch random sampling. Based on the traditional DON reinforcement learning, the greedy exploration method is directly adopted to select actions, that is, only select the action with the largest action value function:

[0032] a←argmax a Q(S,a,ε,θ)(6)

[0033] Wherein, ε is the probability, and θ is the network weight.

[0034] The result of adding the maximum action value function of the next moment and the immediate reward r is used together with the action value generated by the evaluation network for the calculation of the loss function, and the evaluation Q network is updated by the backpropagation of the loss function. The loss function update formula of the NoisyNet-DON method is as follows:

[0035]

[0036] Furthermore, step S3 is specifically as follows:

[0037] S31: If a given sample set {(x i ,y i ), i = 1, 2,..., M} is provided, where the input value x i ∈R m , the output value y i ∈R, and M is the total number of samples, a regression function can be constructed as:

[0038]

[0039] Wherein, ω is the weight matrix, α and b are the regression model coefficients, and K(x i ,x j) is the kernel function of the LSSVM model. The radial basis kernel function type is selected, and its expression is:

[0040]

[0041] In the formula, x i is the center of the jth radial basis function, and σ is the kernel function width. According to the structural risk minimization criterion, it is transformed into solving the minimum value J of the objective function:

[0042]

[0043] In the formula, C is the penalty factor, e i is the acceptable error, represents mapping x i to a high-dimensional feature space.

[0044] S32: Select the penalty factor C and the kernel function width σ as the optimization variables of the PSO algorithm, and select the working condition recognition error of the SVM as the optimization objective:

[0045]

[0046] In the formula, n is the total number of samples in the test set, R(i) is the true category of the ith test sample, P(i) is the predicted category of the ith sample, is the fitness function of the algorithm.

[0047] S33: Obtain the optimal penalty factor C and the kernel function width σ, and update the PSO-SVM working condition identifier.

[0048] Compared with the prior art, the advantages of the present invention are as follows:

[0049] 1. For the fuel cell vehicle energy management method based on working condition construction recognition and noise network reinforcement learning described in the present invention, by integrating typical working condition construction, working condition recognition, and reinforcement learning algorithms, it clearly divides the work and solves the problems of insufficient optimization accuracy, poor adaptability, and high energy consumption of the existing energy management system under complex working conditions, significantly improving the energy utilization efficiency of fuel cell vehicles.

[0050] 2. For the fuel cell vehicle energy management method based on working condition construction recognition and noise network reinforcement learning described in the present invention, considering the working condition changes and power demand dynamic characteristics during vehicle driving, a noise network exploration method is newly added on the basis of reinforcement learning to increase randomness and learning rate, and then the optimal control parameters are automatically adjusted according to the immediate reward and action variables, which can greatly reduce energy consumption, maintain the battery SOC, and extend the service life of fuel cells, achieving efficient energy management.

[0051] 3. A fuel cell vehicle energy management method based on operating condition construction recognition and noise network reinforcement learning according to the present invention designs a support vector machine classifier based on the particle swarm optimization algorithm for online operating condition recognition, and then maps the optimized control parameters of the noise network reinforcement learning algorithm in real time to perform online power distribution, which can improve the comprehensive performance of the vehicle under different driving conditions. Description of the Drawings

[0052] Figure 1 Schematic diagram of the topology structure of the fuel cell vehicle used in the embodiment;

[0053] Figure 2 Algorithm flow chart of the fuel cell vehicle based on operating condition construction recognition and reinforcement learning used in the embodiment;

[0054] Figure 3 Flow chart of combining clustering analysis and short - trip method to construct typical operating conditions used in the embodiment;

[0055] Figure 4 Schematic diagram of the noise network exploration reinforcement learning control algorithm used in the embodiment;

[0056] Figure 5 Flow chart of the support vector machine operating condition recognizer optimized based on the particle swarm algorithm used in the embodiment.

[0057] Markings in the figure: 1. Fuel cell, 2. Power battery, 3. Unidirectional DC / DC converter, 4. Bidirectional DC / DC converter, 5. Inverter, 6. Drive motor. Detailed Embodiment

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described examples are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0059] This application proposes a fuel cell vehicle energy management method based on operating condition construction recognition and noise network reinforcement learning. The powertrain of the fuel cell vehicle is as Figure 1 shown, and it consists of a fuel cell 1, a power battery 2, a unidirectional DC / DC converter 3, a bidirectional DC / DC converter 4, an inverter 5, and a drive motor 6. The fuel cell 1 is connected to the DC bus through the unidirectional DC - DC converter 3, and the power battery 2 is connected to the DC bus through the bidirectional DC / DC converter 4. The energy management system distributes power between the fuel cell 1 and the power battery 2, and the drive motor 6 is used to provide the driving power for the vehicle.

[0060] The schematic diagram of the energy management method is as follows Figure 2 shown. First, typical driving conditions are constructed using the K-means clustering algorithm and the short-trip method. Then, taking the fuel cell energy consumption, lifespan, and the maintenance of the state of charge of the battery as immediate rewards, a reinforcement learning energy management strategy based on a noisy network is developed for control optimization, and the optimal control parameters for various typical driving conditions are obtained through offline training. Finally, a support vector machine driving condition identifier optimized by the particle swarm algorithm is designed to perform online identification of the driving conditions, and the optimal control parameters under the typical driving conditions are mapped to achieve power distribution. The specific steps of this energy management method are as follows:

[0061] S1: Construct typical driving conditions based on the K-means algorithm and the short-trip method, as Figure 3 shown, and its basic steps are as follows:

[0062] S11: Divide the kinematic segments in combination with the standard driving conditions, and select characteristic parameters to describe the kinematic segments;

[0063] S12: According to the characteristic parameters of each kinematic segment, randomly select a sample point as the initial clustering center, calculate the distance D(i) between the current sample i and the nearest clustering center, and then calculate the probability P(i) of the current sample becoming the next clustering center. When the random number rand is greater than P(i - 1) and less than P(i), select the current sample as the clustering center, and repeat this step until K types of clustering centers are determined. Based on the Euclidean distance formula, the probability that the current sample is selected as the clustering center is:

[0064]

[0065] where N is the total number of samples.

[0066] S13: Calculate the distance from each sample to the K types of clustering centers, assign the sample to the cluster where the nearest clustering center is located, and calculate the mean of each cluster, and set it as the new clustering center until the clustering center does not change;

[0067] S14: Based on the clustering analysis results, use the short-trip method to arrange the kinematic segments in sequence to complete the construction of typical driving conditions.

[0068] S2: According to the actual power distribution control process of the fuel cell vehicle, taking the instantaneous equivalent hydrogen consumption, the fuel cell degradation amount, and the maintenance of the SOC of the second power battery as immediate rewards, and the fuel cell power change as the action variable, a reinforcement learning energy management strategy based on noisy network exploration (NoisyNet-DQN) is proposed for control optimization, and four typical driving conditions are used as training conditions to obtain the optimal parameters corresponding to the NoisyNet-DQN strategy under each condition, as Figure 4 shown. The specific steps are as follows:

[0069] S21: Take the SOC of the power battery 2, the output power P of the fuel cell 1 fuel , the required power P dem and the speed V of the vehicle as the state variables S of the NoisyNet-DON strategy:

[0070] S = {SOC, P fuel , P dem , V} (2)

[0071] S22: Set the action variable as the change in fuel cell power:

[0072] A = {-4, -3, -2, -1, 0, 1, 2, 3, 4} (3)

[0073] Set the immediate reward as the negative of the instantaneous equivalent fuel consumption, the degradation of the fuel cell 1, and the maintenance of the SOC of the power battery 2:

[0074]

[0075] In the formula, α′, β′, γ′ are the weight factors of the economic term, the fuel cell 1 degradation term, and the power battery 2 SOC maintenance term respectively, SOC ref is the reference value of the SOC of the power battery 2, is the maximum value of the instantaneous fuel cell degradation, is the instantaneous equivalent hydrogen consumption, and the subscripts min and max represent the minimum and maximum values; in addition, the state constraints during vehicle driving are set as:

[0076]

[0077] In the formula, P bat is the power of the fuel cell 1, and η DC / DC is the efficiency of the unidirectional DC / DC converter 3.

[0078] S23: The NoisyNet DQN method saves the data such as the state S, action a, immediate reward, and the action at the next moment generated during the interaction process into the experience pool, and conducts intelligent agent Q-network learning through mini-batch random sampling. Based on the traditional DON reinforcement learning, directly adopt the greedy exploration method to select actions, that is, only select the action with the largest action value function:

[0079] a ← argmax a Q(S, a, ε, θ) (6)

[0080] In the formula, ε is the probability and θ is the network weight.

[0081] The sum of the maximum action value function at the next moment and the immediate reward r is used together with the action value generated by the evaluation network for the calculation of the loss function, and the evaluation Q-network is updated by backpropagation of the loss function. The loss function update formula of the NoisyNet-DON method is as follows:

[0082]

[0083] S3: Design a working condition identifier based on a support vector machine (SVM) optimized by the particle swarm optimization (PSO) algorithm to perform online identification of driving working conditions, as Figure 5 shown. The specific steps are as follows:

[0084] S31: If a given sample set {(x i , y i ), i = 1, 2,..., M} is given, where the input value x i ∈ R m , and the output value yi ∈ R, and M is the total number of samples, a regression function can be constructed as:

[0085]

[0086] In the formula, ω is the weight matrix, α and b are the regression model coefficients, and K(x i , x j ) is the kernel function of the LSSVM model. Select the radial basis kernel function type, and its expression is:

[0087]

[0088] In the formula, x i is the center of the jth radial basis function, and σ is the kernel function width. According to the structural risk minimization criterion, it is transformed into solving the minimum value J of the objective function:

[0089]

[0090] In the formula, C is the penalty factor, e i is the acceptable error, represents mapping x i to a high-dimensional feature space.

[0091] S32: Select the penalty factor C and the kernel function width σ as the optimization variables of the PSO algorithm, and select the working condition identification error of the SVM as the optimization objective:

[0092]

[0093] Where n is the total number of samples in the test set, R(i) is the true category of the i-th test sample, and P(i) is the predicted category of the i-th sample. is the fitness function of the algorithm.

[0094] S33: Obtain the optimal penalty factor C and kernel function width σ, and update the PSO-SVM working condition recognizer.

[0095] S4: Under the actual working conditions, according to the PSO-SVM working condition recognition result, map the optimal parameters of the NoisyNet-DQN strategy corresponding to the current working condition. The energy management system performs power distribution between the fuel cell 1 and the power battery 2, and then transmits it to the drive motor 6 through the inverter 5 to realize the online adaptive adjustment of the control algorithm parameters and improve the comprehensive performance of the vehicle.

[0096] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made by using the content of the specification and drawings of the present invention under the inventive concept of the present invention, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.

Claims

1. A fuel cell vehicle energy management method based on working condition construction, recognition and noise network reinforcement learning, characterized in that The energy management method specifically includes: S1: Construct typical working conditions based on the K-means algorithm and the short-travel method; S2: According to the actual power distribution control process of the fuel cell vehicle, taking the instantaneous equivalent hydrogen consumption, the fuel cell degradation amount, and the SOC maintenance of the power battery (2) as immediate rewards, and the fuel cell power change as the action variable, propose a reinforcement learning energy management strategy based on noisy network exploration for control optimization, and use four typical working conditions as training conditions to obtain the optimal parameters corresponding to the NoisyNet-DQN strategy under each condition; S3: Design a support vector machine working condition identifier optimized by the particle swarm algorithm for online identification of driving working conditions; S4: Under actual working conditions, according to the PSO-SVM working condition identification results, map the optimal parameters of the NoisyNet-DQN strategy corresponding to the current working condition. The energy management system distributes power between the fuel cell (1) and the power battery (2), and then transmits it to the drive motor (6) through the inverter (5) to achieve online adaptive adjustment of the control algorithm parameters and improve the comprehensive performance of the vehicle.

2. The energy management method according to claim 1, wherein The transmission system of the fuel cell vehicle consists of a fuel cell (1), a power battery (2), a unidirectional DC / DC converter (3), a bidirectional DC / DC converter (4), an inverter (5), and a drive motor (6). The fuel cell (1) is connected to the DC bus through a unidirectional DC-DC converter (3), and the power battery (2) is connected to the DC bus through a bidirectional DC-DC converter (4).

Citation Information

Patent Citations

  • A fuel cell vehicle energy management method based on deep reinforcement learning

    CN118386949B

  • Vehicle energy management method for hydrogen fuel cell vehicle

    CN119348505A