Online updating method for intelligent energy management strategy of networked fuel cell passenger car
By obtaining diversified traffic information in real time in the Internet of Vehicles environment and combining dynamic planning and deep reinforcement learning methods, the online update of fuel cell passenger bus energy management strategies is solved, and the problem that strategies in the existing technology cannot adapt to changes in driving conditions is improved, and the accuracy and efficiency of energy management are improved.
Patent Information
- Application Number
- CN202510455160.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
AI Technical Summary
The existing fuel cell bus energy management strategy based on deep reinforcement learning failed to effectively adapt to the dynamic changes in the actual driving conditions of the vehicle during the offline training stage, making it difficult to ensure long-term optimization performance after the deployment of the actual vehicle.
In the environment of Internet of Vehicles, obtain diversified traffic information in real time, and combine dynamic planning algorithms and deep reinforcement learning to update energy management strategies online, and use knowledge transfer and strategy reuse to optimize power distribution.
It significantly improves the accuracy of energy management strategies and long-term optimization performance, and improves the energy management efficiency of fuel cell buses under diversified traffic information.
Smart Images

Figure CN120297431A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of energy management of fuel cell buses, and particularly relates to an online update method for an intelligent energy management strategy of a connected fuel cell bus that integrates speed planning and transfer learning. Background Art
[0002] At the present stage, with the development and popularization of artificial intelligence technology, energy management strategies based on deep reinforcement learning have become the mainstream methods in the energy management field due to their powerful self-learning and self-optimization capabilities. However, such energy management strategies generally only focus on improving the strategy optimization effect in the offline training stage by reasonably setting the reward function and algorithm hyperparameters, but ignore the characteristics of the continuous dynamic change of the actual driving conditions of the vehicle, resulting in untimely update of the trained strategy, so that it cannot well adapt to the change of driving conditions after actual vehicle deployment, and thus it is difficult to ensure long-term optimization performance. Summary of the Invention
[0003] In view of this, aiming at the technical problems existing in the field, the present invention provides an online update method for an intelligent energy management strategy of a connected fuel cell bus, which specifically includes the following steps:
[0004] Step 1: Real-time obtain the multi-source traffic information of the connected fuel cell bus in the vehicle networking environment and upload it to the cloud computing platform;
[0005] For the global optimal speed planning problem, set the spatio-temporal coupling constraint conditions in the traffic information; respectively define the state variables including vehicle speed and driving time, and the vehicle acceleration as the action variable; by discretizing the driving distance of each station interval according to a specific step size, transform the global planning into a finite-step optimization, and design the state transition equation and the speed planning target; on this basis, design a dynamic programming algorithm for solving the global optimal speed of the future bus station interval in the cloud platform;
[0006] Step 2: Fine-tune the target energy management strategy neural network based on deep reinforcement learning for any station interval, use the future optimal vehicle speed solved by the above dynamic programming as the input, and use the fine-tuned target energy management strategy as the output; after fine-tuning this station interval, use the network parameters for initializing the neural network parameters when fine-tuning the next station interval to achieve knowledge transfer and strategy reuse;
[0007] Step 3: After using Step 2 to fine-tune the target energy management strategy of the current station interval, establish a Markov state transition probability matrix of the drive demand power corresponding to the historical vehicle speed and the future vehicle speed, and quantitatively update the target energy management strategy based on the quantization index of the difference rate of the transition probability matrix between adjacent stations;
[0008] Step 4: Download and deploy the updated target energy management strategy from the cloud platform to the vehicle controller, so that the vehicle plans the optimal vehicle speed based on real-time multi-source traffic information, and finally realizes the optimization of power distribution during driving.
[0009] Furthermore, the multi-source traffic information obtained in Step 1 at least includes: signal light position, signal light timing and phase, bus stop position, number of passengers, road surface gradient, road speed limit, driving distance, and driving time.
[0010] Since the spatial constraint is only related to the vehicle position, the time constraint is transformed into a spatial constraint, so that the global optimal planning problem is defined as an optimization problem in the spatial domain, and the state variables are defined as the vehicle speed v(s) and driving time t(s), and the action variable is the vehicle acceleration a(s), that is, the state space x(s) = [v(s), t(s)] T , the action space u(s) = a(s), and s represents the system state;
[0011] Discretize the driving distance of each station interval according to the step size Δs, so that the global planning is transformed into a finite-step optimization. When the discrete step size Δs is short enough, the vehicle is regarded as doing uniform acceleration and deceleration within each discrete step. Therefore, the state transition equation of the optimization problem is defined as:
[0012]
[0013] In the formula, v(k) and v(k + 1) respectively represent the vehicle speeds at the kth (k = 0, 1,..., N - 1) and k + 1th spatial steps, and t(k) and t(k + 1) respectively represent the passing times at the kth and k + 1th spatial steps;
[0014] Define the vehicle speed planning goal as: solve an optimal speed trajectory to minimize the total energy consumption required by the vehicle wheel end drive, and the corresponding cost function is:
[0015]
[0016] In the formula, P m represents the electric power of the drive motor, m represents the vehicle mass, g represents the acceleration due to gravity, f represents the rolling resistance coefficient, represents the road surface gradient, C d represents the air resistance coefficient, A represents the frontal area, δ represents the conversion coefficient of the rotating mass, η m represents the working efficiency of the motor, κ characterizes the working state of the motor, κ = -1 in the driving state, and κ = 1 in the braking state;
[0017] Set various constraint conditions, including:
[0018] Speed constraint:
[0019] v(s0) = 0, v(s f ) = 0
[0020] 0 ≤ v(k) ≤ v limit
[0021] 0 ≤ v(k) ≤ v max
[0022] where s0 represents the starting point of each station section, s f represents the ending point of each station section, v limit represents the maximum vehicle speed of the speed limit section, v max represents the maximum designed vehicle speed of the fuel cell bus;
[0023] Travel time constraint:
[0024] t(s0) = 0, t(s f ) = t f
[0025] where t f represents the travel duration of each station section;
[0026] Signal light safety constraint:
[0027]
[0028] where i represents the signal light number, i = 1, 2,..., n; t represents the absolute running time of the vehicle; T0 represents the offset time of the signal light, that is, the time that the signal light has run within its cycle when the vehicle departs; T, T r and T g respectively represent the cycle duration, red light phase duration and green light phase duration of the signal light;
[0029] Acceleration constraint:
[0030] -a max ≤ a(k) ≤ a max
[0031] where a max represents the maximum acceleration of the vehicle;
[0032] The design of the dynamic programming algorithm includes:
[0033] Using π = {a0, a1,..., a N-1} to represent any solution to the global optimal vehicle speed planning problem, representing the feasible sets of state variables and action variables. In the initial state, x(0) = x0, then redefine the cost function with the control strategy π applied as:
[0034]
[0035] where, g N (x N ) represents the terminal cost, and g k (x k , u k ) represents the cost after applying the action u k at the state x k ;
[0036] Based on the Bellman optimality principle, solve the minimum cost function for the terminal k = N:
[0037]
[0038] Solve the minimum cost functions for the intermediate steps k = N - 1, N - 2, …, 0:
[0039]
[0040] After performing the above reverse calculations, given the initial state value, obtain the complete optimal control sequence through forward optimization to obtain the global optimal vehicle speed trajectory.
[0041] Furthermore, the process of fine-tuning the energy management strategy for adjacent station intervals in step two includes:
[0042] (1) Policy pre-training: Train the initial energy management strategy for each station interval M represents the total number of station intervals on the bus line;
[0043] (2) Policy parameter transfer: Initialize the shallow neural network parameters of the target energy management strategy for the next station interval using the shallow neural network parameters of the initial energy management strategy, and randomly initialize the deep neural network parameters of the target energy management strategy;
[0044] (3) Policy fine-tuning: Use the planned future vehicle speed as the operating condition input, and fine-tune the initialized target energy management strategy for the next station in the cloud platform until stable convergence; Represent the target energy management strategy for each station interval as Represent the fine-tuning process as
[0045] Furthermore, use the same deep reinforcement learning algorithm for the pre-training and fine-tuning processes of the target energy management strategy, and the loss function of its value network is:
[0046]
[0047] where, r t 、s tand a t represent rewards, states, and actions respectively, γ ∈ (0, 1) represents the discount factor, and α ∈ (0, 1) represents the temperature factor. denotes the experience pool. denotes the value network with parameters θ i ; denotes the target value network with parameters ; π represents the policy network, and π φ denotes the policy network with parameters φ;
[0048] The loss function of the policy network is:
[0049]
[0050] Specifically, the state space is defined as:
[0051] s t = [v, a, P fc , SOC]
[0052] where v, a, P fc , and SOC represent speed, acceleration, fuel cell power, and state of charge of the battery respectively;
[0053] The action space is defined as:
[0054] a t = ΔP fc , ΔP fc ∈ [-5kW, 5kW]
[0055] where ΔP fc represents the power change rate of the fuel cell;
[0056] And the reward function is defined as:
[0057]
[0058] where represents the hydrogen consumption rate, SOC tar represents the target SOC value, and ψ is the weight factor.
[0059] Furthermore, the update process of the target energy management strategy corresponding to the entire bus line includes:
[0060] For the initial energy management strategy of the first station section Pre-trained using any operating condition data applicable to fuel cell buses before the vehicle departs; from the first stop interval to the penultimate stop interval, the fuel cell bus obtains the traffic information of the next stop interval at departure and plans the optimal speed trajectory, and then fine-tunes the target energy management strategy for the next stop interval; and so on, making the target energy management strategy for each stop interval serve as the initial energy management strategy for the next stop interval By continuously fine-tuning the target energy management strategy for the next stop interval until reaching the terminal of the bus line; and in the last stop interval, only use the target energy management strategy of the penultimate stop interval to perform real-time energy management, and no longer perform optimal speed trajectory solving and network parameter updating; thus, the complete update process of the energy management strategy on the entire bus line is expressed as:
[0061]
[0062] Furthermore, step three specifically includes the following steps:
[0063] First, establish a Markov state transition probability matrix of the driving demand power using the maximum likelihood estimation method:
[0064]
[0065] In the formula, p xy (λ) represents the one-step transition probability that the driving demand power transfers from state x at time λ to state y at time λ + 1, represents the number of times of transferring from state x to state y, and z represents the total number of times of transferring from state x;
[0066] After that, use the JS divergence value to quantify the matrix difference rate:
[0067]
[0068] Finally, use the moving average method to quantitatively update the policy network parameters:
[0069] φ←δ·φ0+(1-δ)φ
[0070] In the formula, φ0 represents the original parameters of the policy network.
[0071] Furthermore, in step four, the optimal vehicle speed planning result of each stop interval is used as the recommended vehicle speed of the fuel cell bus; each time the fuel cell bus departs, the fine-tuned policy network parameters are downloaded from the cloud platform and deployed to the vehicle-mounted controller, and the power distribution is optimized in real time during driving.
[0072] The online update method for the intelligent energy management strategy of the connected fuel cell bus provided by the present invention combines the ideas of knowledge transfer and policy reuse on the basis of deep reinforcement learning. It can make full use of the results of consecutive station intervals in model training to achieve efficient online update of the intelligent energy management strategy, improve the relevance between the energy management executed by the strategy and multi-source traffic information, and thus can significantly enhance the accuracy and long-term optimization performance of the strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 is the overall flowchart of the method provided by the present invention;
[0074] Figure 2 is the detailed architecture diagram of the method provided by the present invention;
[0075] Figure 3 (a) is the speed-distance curve of the vehicle speed planning result for the entire bus line;
[0076] Figure 3 (b) is the speed-time curve of the vehicle speed planning result for the entire bus line;
[0077] Figure 4 (a) is the distance-time curve of the speed planning result for the second station interval;
[0078] Figure 4 (b) is the speed-distance curve of the speed planning result for the second station interval;
[0079] Figure 4 (c) is the speed-time curve of the speed planning result for the second station interval;
[0080] Figure 5 is the convergence curve of the energy management strategy for the second station interval;
[0081] Figure 6 (a) is the Markov state transition probability matrix of the driving demand power for the first station interval;
[0082] Figure 6 (b) is the Markov state transition probability matrix of the driving demand power for the second station interval;
[0083] Figure 7 is the comparison of hydrogen consumption of different energy management strategies. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0085] The present invention provides an online update method for the intelligent energy management strategy of a connected fuel cell bus, as Figure 1 、 2 shown, specifically including the following steps:
[0086] Step 1: Real-time obtain the multi-source traffic information of the connected fuel cell bus in the vehicle networking environment and upload it to the cloud computing platform;
[0087] For the global optimal vehicle speed planning problem, set the spatio-temporal coupling constraint conditions in the traffic information; respectively define the state variables including vehicle speed and travel time, and the vehicle acceleration as the action variable; by discretizing the driving distance of each station interval according to a specific step size, transform the global planning into a finite-step optimization, and design the state transition equation and the vehicle speed planning objective; on this basis, design a dynamic programming algorithm to solve the global optimal vehicle speed of the future bus station interval in the cloud platform;
[0088] Step 2: Fine-tune the target energy management strategy neural network based on deep reinforcement learning for any station interval, using the future optimal vehicle speed solved by the above dynamic programming as the input and the fine-tuned target energy management strategy as the output; after fine-tuning this station interval, use the network parameters for the initialization of the neural network parameters when fine-tuning the next station interval to achieve knowledge transfer and strategy reuse;
[0089] Step 3: After using Step 2 to fine-tune the target energy management strategy of the current station interval, establish a Markov state transition probability matrix of the drive demand power corresponding to the historical vehicle speed and the future vehicle speed, and quantitatively update the target energy management strategy based on the difference rate quantization index of the transition probability matrix between adjacent stations;
[0090] Step 4: Download and deploy the updated target energy management strategy from the cloud platform to the vehicle-mounted controller, so that the vehicle plans the optimal vehicle speed based on the real-time multi-source traffic information, and finally realizes the optimization of power distribution during driving.
[0091] In a preferred embodiment of the present invention, the multi-source traffic information obtained in Step 1 at least includes: signal light position, signal light timing and phase, bus stop position, number of passengers, road surface slope, road speed limit, driving distance and travel time;
[0092] Since the space constraint is only related to the vehicle position, the time constraint is transformed into a space constraint, so that the global optimal planning problem is defined as an optimization problem in the spatial domain. The state variables are defined as the vehicle speed v(s) and the driving time t(s), and the action variable is the vehicle acceleration a(s), that is, the state space x(s) = [v(s), t(s)] T , and the action space u(s) = a(s), where s represents the system state;
[0093] The driving distance of each station interval is discretized according to the step size Δs, so that the global planning is transformed into a finite-step optimization. When the discrete step size Δs is short enough, the vehicle is regarded as moving with uniform acceleration and deceleration within each discrete step. Therefore, the state transition equation of the optimization problem is defined as:
[0094]
[0095] In the formula, v(k) and v(k + 1) represent the vehicle speeds at the kth (k = 0, 1, …, N - 1) and (k + 1)th spatial steps respectively, and t(k) and t(k + 1) represent the passing times at the kth and (k + 1)th spatial steps respectively;
[0096] The vehicle speed planning objective is defined as: to solve an optimal speed trajectory to minimize the total energy consumption required by the vehicle wheel-end drive. The corresponding cost function is:
[0097]
[0098] In the formula, P m represents the electric power of the drive motor, m represents the vehicle mass, g represents the acceleration due to gravity, f represents the rolling resistance coefficient, represents the road surface gradient, C d represents the air resistance coefficient, A represents the frontal area, δ represents the rotary mass conversion coefficient, η m represents the working efficiency of the motor, and κ characterizes the working state of the motor. When in the driving state, κ = -1, and when in the braking state, κ = 1;
[0099] Set various constraint conditions, including:
[0100] Speed constraint:
[0101] v(s0) = 0, v(s f ) = 0
[0102] 0 ≤ v(k) ≤ v limit
[0103] 0 ≤ v(k) ≤ v max
[0104] In the formula, s0 represents the starting point of each station interval, s f represents the end point of each station interval, vlimit Represents the maximum vehicle speed in the speed - limited section, v max Represents the maximum designed vehicle speed of the fuel - cell bus;
[0105] Driving - time constraint:
[0106] t(s0) = 0, t(s f ) = t f
[0107] In the formula, t f Represents the driving duration of each station interval;
[0108] Signal - light safety constraint:
[0109]
[0110] In the formula, i represents the signal - light number, i = 1, 2, …, n; t represents the absolute running time of the vehicle; T0 represents the offset time of the signal - light, that is, the time that the signal - light has run within its cycle when the vehicle departs; T, T r and T g Represent the cycle duration, red - light phase duration and green - light phase duration of the signal - light respectively;
[0111] Acceleration constraint:
[0112] -a max ≤a(k)≤a max
[0113] In the formula, a max Represents the maximum acceleration of the vehicle;
[0114] Designing the dynamic - programming algorithm includes:
[0115] Using π = {a0, a1, …, a N-1} to represent any solution of the global - optimal vehicle - speed planning problem, Represents the feasible sets of state variables and action variables. In the initial state, x(0) = x0, then re - define the cost function with the control strategy π imposed as:
[0116]
[0117]
[0118] In the formula, g N (x N ) represents the terminal cost, g k (x k , u k ) represents the cost after applying the action u k at the state x k ;
[0119] Based on the Bellman optimality principle, solve the minimum cost function for the terminal \(k = N\):
[0120]
[0121] Solve the minimum cost function for the intermediate steps \(k = N - 1, N - 2, \ldots, 0\):
[0122]
[0123] After performing the above reverse calculation, given the initial state value, obtain the complete optimal control sequence through forward optimization The global optimal vehicle speed trajectory can be obtained.
[0124] In a preferred embodiment of the present invention, the process of fine-tuning the energy management strategy for adjacent station intervals in step two includes:
[0125] (1) Policy pre-training: Train the initial energy management strategy for each station interval M represents the total number of station intervals on the bus line;
[0126] (2) Policy parameter transfer: Initialize the shallow neural network parameters of the target energy management strategy for the next station interval using the shallow neural network parameters of the initial energy management strategy, and randomly initialize the deep neural network parameters of the target energy management strategy;
[0127] (3) Policy fine-tuning: Take the planned future vehicle speed as the operating condition input, and fine-tune the initialized target energy management strategy for the next station in the cloud platform until stable convergence; Represent the target energy management strategy for each station interval as Represent the fine-tuning process as
[0128] Furthermore, use the same deep reinforcement learning algorithm for the target energy management strategy pre-training and fine-tuning processes, and the loss function of its value network is:
[0129]
[0130] In the formula, \(r\) t , \(s\) t and \(a\) t represent the reward, state, and action respectively, \(\gamma\in(0,1)\) represents the discount factor, \(\alpha\in(0,1)\) represents the temperature factor, represents the experience pool, represents the value network with parameters \(\theta\) i , represents the parameters as The target value network, π represents the policy network, π φ represents the policy network with parameters φ;
[0131] The loss function of the policy network is:
[0132]
[0133] Specifically, the state space is defined as:
[0134] s t = [v, a, P fc , SOC]
[0135] In the formula, v, a, P fc , SOC respectively represent speed, acceleration, fuel cell power, and state of charge of the battery;
[0136] The action space is defined as:
[0137] a t = ΔP fc , ΔP fc ∈[-5kW, 5kW]
[0138] In the formula, ΔP fc represents the power change rate of the fuel cell;
[0139] And the reward function is defined as:
[0140]
[0141] In the formula, represents the hydrogen consumption rate, SOC tar represents the target SOC value, and ψ is a weight factor.
[0142] In a preferred embodiment of the present invention, the update process of the target energy management strategy corresponding to the entire bus line includes:
[0143] For the initial energy management strategy of the first station section It is pre-trained using any operating condition data applicable to fuel cell buses before the vehicle departs; from the first station section to the penultimate station section, the fuel cell bus obtains the traffic information of the next station section at departure and plans the optimal speed trajectory, and then fine-tunes the target energy management strategy of the next station section; and so on, making the target energy management strategy of each station section serve as the initial energy management strategy of the next station section By continuously fine-tuning the target energy management strategy of the next station section until reaching the terminal of the bus line; and in the last stop interval, only use the target energy management strategy of the penultimate stop interval Execute real-time energy management, and no longer perform optimal speed trajectory solution and network parameter update; thus, the complete update process of the energy management strategy on the entire bus line is expressed as:
[0144]
[0145] In a preferred embodiment of the present invention, step three specifically includes the following steps:
[0146] First, use the maximum likelihood estimation method to establish the Markov state transition probability matrix of the driving demand power:
[0147]
[0148] In the formula, p xy (λ) represents the one-step transition probability that the driving demand power transfers from state x at time λ to state y at time λ + 1, represents the number of times of transferring from state x to state y, and z represents the total number of transfers from state x;
[0149] After that, use the JS divergence value to quantify the matrix difference rate:
[0150]
[0151] Finally, use the moving average method to quantitatively update the policy network parameters:
[0152] φ←δ·φ0+(1-δ)φ
[0153] In the formula, φ0 represents the original parameters of the policy network.
[0154] In step four, the optimal vehicle speed planning result of each stop interval is used as the recommended vehicle speed of the fuel cell bus; each time the fuel cell bus departs, the fine-tuned policy network parameters are downloaded from the cloud platform and deployed to the vehicle-mounted controller, and the power distribution is optimized in real time during driving.
[0155] In a specific example of the present invention, there are 24 stop intervals on the bus line, the spatial domain discrete step size Δs = 1m, the total line length is 14.70km, and the total travel time is 2610s; Figure 3 (a) and Figure 3 (b) are respectively the speed-distance curve and speed-time curve of the vehicle speed planning result of the entire bus line; it can be seen that the vehicle speed planning result meets the setting of the total line length and the total travel time.
[0156] Taking the second station section as an example, the section length is 1655 m and the passing time is 155 s; there are a total of 3 signalized intersections, and the positions of the signal lights are 210 m, 460 m, and 1350 m respectively. The cycle T i of each signal light, the red light phase duration and the green light phase duration are 60 s, 30 s, and 30 s respectively. The offset time of each signal light is 5 s, 50 s, and 40 s respectively. The speed limit section is located at 700 m to 1100 m, and the speed limit value v limit is 40 km / h; the vehicle speed planning results of the second station section are as Figure 4 shown Figure 4 (a) is the distance-time curve of the speed planning results of this station section. It can be seen that the fuel cell bus can safely pass through each signalized intersection continuously within the green light phase; Figure 4 (b) is the speed-distance curve of the speed planning results of this station section. It can be seen that the fuel cell bus can strictly comply with the road speed limit constraints and drive at the maximum allowable speed in the speed limit section; Figure 4 (c) is the speed-time curve of the speed planning results of this station section. It can be seen that the fuel cell bus can pass through the second station section at a relatively stable vehicle speed within the established passing time. The convergence curve of the energy management strategy for the second station section is as Figure 5 shown. It can be seen that the energy management strategy without transfer learning needs 70 training rounds to converge because it needs to train the agent from scratch. In contrast, the energy management strategy incorporating transfer learning only needs 30 rounds of fine-tuning to converge smoothly, improving the convergence speed by 57.14%; thus, generalizing to other station sections, through calculation, the method of the present invention improves the update efficiency of the energy management strategy for the entire bus line by 45.86%. The Markov state transition probability matrix of the drive demand power for the first station section is as Figure 6 (a) shown, and the Markov state transition probability matrix of the drive demand power for the second station section is as Figure 6 (b) shown; through calculation using formula (22), the difference rate between the two matrices is 0.40. The comparison of the hydrogen consumption of different energy management strategies is as Figure 7 shown, where the "static" strategy represents the unupdated energy management strategy; through calculation, the energy management strategy of the present invention reduces the hydrogen consumption by 6.19% compared to the unupdated energy management strategy, and the optimization effect is obvious. These results verify the effectiveness of the optimal vehicle speed planning method in the present invention.
[0157] It should be understood that the sequence numbers of the steps in the embodiments of the present invention do not imply the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0158] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An online update method for the intelligent energy management strategy of a connected fuel cell bus, characterized in that: Specifically, it includes the following steps: Step 1: In the vehicle networking environment, obtain the multi-source traffic information of the connected fuel cell bus in real time and upload it to the cloud computing platform; Regarding the global optimal vehicle speed planning problem, set the spatio-temporal coupling constraint conditions in the traffic information; define the state variables including vehicle speed and driving time, and the vehicle acceleration as the action variable respectively; by discretizing the driving distance of each station interval according to a specific step size, transform the global planning into a finite-step optimization, and design the state transition equation and the vehicle speed planning objective; on this basis, design a dynamic programming algorithm to solve the global optimal vehicle speed of the future bus station interval in the cloud platform; Step 2: Individually fine-tune the target energy management strategy neural network based on deep reinforcement learning for any station interval, with the future optimal vehicle speed solved by the above dynamic programming as the input and the fine-tuned target energy management strategy as the output; After fine-tuning this station interval, use the network parameters to initialize the neural network parameters when fine-tuning the next station interval to achieve knowledge transfer and strategy reuse; Step 3: After using Step 2 to fine-tune the target energy management strategy of the current station interval, establish a Markov state transition probability matrix of the drive demand power corresponding to the historical vehicle speed and the future vehicle speed, and quantitatively update the target energy management strategy based on the difference rate quantization index of the transition probability matrix of adjacent stations; Step 4: Download and deploy the updated target energy management strategy from the cloud platform to the vehicle-mounted controller, so that the vehicle plans the optimal vehicle speed based on the real-time multi-source traffic information, and finally realizes the optimization of power distribution during the driving process.
2. The method according to claim 1, wherein: The multi-source traffic information obtained in Step 1 includes at least: signal lamp position, signal lamp timing and phase, bus stop position, number of passengers, road surface gradient, road speed limit, driving distance and driving time; Since the spatial constraint is only related to the vehicle position, the time constraint is transformed into a spatial constraint, so that the global optimal planning problem is defined as an optimization problem in the spatial domain. The state variables are defined as the vehicle speed v(s) and the driving time t(s), and the action variable is the vehicle acceleration a(s), that is, the state space x(s) = [v(s), t(s)] T , the action space u(s) = a(s), where s represents the system state; Discretize the driving distance of each station interval according to the step size Δs, so that the global planning is transformed into a finite-step optimization. When the discretization step size Δs is short enough, the vehicle is regarded as performing uniform acceleration and deceleration motion within each discrete step. Therefore, the state transition equation of the optimization problem is defined as: In the formula, v(k) and v(k + 1) respectively represent the vehicle speeds at the kth (k = 0, 1, …, N - 1) and k + 1th space steps, and t(k) and t(k + 1) respectively represent the passing times at the kth and k + 1th space steps; Define the vehicle speed planning objective as: solve an optimal speed trajectory to minimize the total energy consumption required by the vehicle wheel end drive, and the corresponding cost function is: where P m represents the electric power of the drive motor, m represents the mass of the whole vehicle, g represents the acceleration due to gravity, f represents the rolling resistance coefficient, represents the road surface gradient, C d represents the air resistance coefficient, A represents the frontal area, δ represents the rotating mass conversion coefficient, η m represents the working efficiency of the motor, κ characterizes the working state of the motor, κ = -1 in the driving state, and κ = 1 in the braking state; Set various constraint conditions, including: Speed constraint: v(s0) = 0, v(s f ) = 0 0 ≤ v(k) ≤ v limit 0 ≤ v(k) ≤ v max where s0 represents the starting point of each station section, and s f represents the end point of each station section, and v limit represents the maximum speed of the speed-limited section, and v max represents the maximum designed speed of the fuel cell bus; Driving time constraint: t(s0) = 0, t(s f ) = t f where t f represents the driving duration of each station section; Signal lamp safety constraint: Wherein, i represents the signal lamp number, i = 1, 2, …, n; t represents the absolute running time of the vehicle; T0 represents the offset time of the signal lamp, that is, the time that the signal lamp has run within its cycle when the vehicle departs; T, T r and T g respectively represent the cycle duration, red light phase duration and green light phase duration of the signal lamp; Acceleration constraint: -a max ≤a(k)≤a max where a max represents the maximum acceleration of the vehicle; The design of the dynamic programming algorithm includes: Use π = {a0, a1, …, a N-1} to represent any solution to the global optimal vehicle speed planning problem, represent the feasible sets of state variables and action variables. At the initial state x(0) = x0, the cost function with the control strategy π applied is redefined as: where, g N (x N ) represents the terminal cost, g k (x k , u k ) represents the cost after applying the action u k at the state x k ; Based on the Bellman optimality principle, solve the minimum cost function at the terminal k = N: Solve the minimum cost function at the intermediate steps k = N - 1, N - 2, …, 0: After performing the above reverse calculation, given the initial state value, a complete optimal control sequence is obtained through forward optimization The global optimal vehicle speed trajectory can be obtained 3. The method according to claim 2, wherein: The process of fine-tuning the energy management strategy for adjacent station intervals in Step 2 includes: (1) Policy pre-training: Train the initial energy management policy for each station section M represents the total number of station sections on the bus line; (2) Policy parameter migration: Initialize the shallow neural network parameters of the target energy management strategy for the next station section using the shallow neural network parameters of the initial energy management strategy, and randomly initialize the deep neural network parameters of the target energy management strategy; (3) Strategy fine-tuning: Using the planned future vehicle speed as the operating condition input, fine-tune the target energy management strategy for the next site that has been initialized in the cloud platform until stable convergence; represent the target energy management strategy for each site interval as Represent the fine-tuning process as 4. The method according to claim 3, wherein: The same deep reinforcement learning algorithm is used for the pre-training and fine-tuning processes of the target energy management strategy, and the loss function of its value network is as follows: where r t , s t and a t represent rewards, states, and actions respectively, γ ∈ (0, 1) represents the discount factor, α ∈ (0, 1) represents the temperature factor, denotes the experience pool, represents the value network with parameter θ i , represents the target value network with parameter , π represents the policy network, π φ represents the policy network with parameter φ; The loss function of the policy network is: Specifically, define the state space as: s t = [v, a, P fc , SOC] where v, a, P fc , and SOC represent speed, acceleration, fuel cell power, and state of charge of the battery, respectively; Define the action space as: a t = ΔP fc , ΔP fc ∈ [-5 kW, 5 kW] where ΔP fc represents the power change rate of the fuel cell; And define the reward function as: In the formula, represents the hydrogen consumption rate, and SOC tar represents the target SOC value, and ψ is the weight factor.
5. The method according to claim 4, wherein: The update process of the target energy management strategy corresponding to the entire bus line includes: Initial energy management strategy for the first stop interval It is pre-trained using any operating condition data applicable to fuel cell buses before the vehicle departs; from the first stop interval to the penultimate stop interval, the fuel cell bus obtains the traffic information of the next stop interval at departure and plans the optimal speed trajectory, and then fine-tunes the target energy management strategy for the next stop interval; and so on, making the target energy management strategy for each stop interval serve as the initial energy management strategy for the next stop interval By continuously fine-tuning the target energy management strategy for the next stop interval until reaching the terminal of the bus line; and in the last stop interval, only use the target energy management strategy of the penultimate stop interval to perform real-time energy management, without solving the optimal speed trajectory and updating network parameters anymore; thus, the complete update process of the energy management strategy on the entire bus line is expressed as:
6. The method according to claim 5, characterized in that: Step 3 specifically includes the following steps: First, establish the Markov state transition probability matrix of the driving demand power using the maximum likelihood estimation method: where p xy (λ) represents the one-step transition probability that the driving demand power transfers from state x at time λ to state y at time λ + 1, represents the number of times of transferring from state x to state y, and z represents the total number of times of transferring from state x; After that, use the JS divergence value δ to quantify the matrix difference rate; finally, use the moving average method to quantitatively update the policy network parameters: φ←δ·φ0+(1-δ)φ In the formula, φ0 represents the original parameters of the policy network.
7. The method according to claim 6, characterized in that: In Step 4, use the optimal vehicle speed planning result of each station section as the recommended vehicle speed of the fuel cell bus; when the fuel cell bus departs each time, download the fine-tuned policy network parameters from the cloud platform and deploy them to the vehicle-mounted controller, and optimize the power distribution in real time during driving.