Hybrid power supply electric public transportation system charging scheduling optimization method based on reinforcement learning
Through a multi-objective Markov decision model based on reinforcement learning and a multi-agent deep reinforcement learning algorithm, the charging scheduling of the electric bus system is optimized, which solves the problem of irrational resource allocation under the hybrid power supply mode and achieves efficient resource utilization and low carbon emissions.
Patent Information
- Application Number
- CN202511293141.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-11
AI Technical Summary
The existing electric bus system faces problems such as resource redundancy and waste, high energy consumption and high operating costs caused by irrational resource allocation under the hybrid power supply mode. Especially in the case of uncertainty in photovoltaic power generation and limited energy storage, it is difficult to meet real-time charging needs.
A multi-objective Markov decision model based on reinforcement learning is adopted, combined with a multi-agent deep reinforcement learning algorithm, and a multi-task head architecture is designed. By training the reinforcement learning network model, the charging scheduling strategy is optimized, considering different tasks under sufficient and insufficient sunlight, meeting real-time charging needs and minimizing operating costs and carbon emissions.
It realizes the direct mapping of multi-source observation data to charging control instructions under different environmental conditions, meets the real-time charging scheduling needs, reduces resource waste and energy consumption, reduces operating costs, and optimizes carbon emissions.
Smart Images

Figure CN120806572A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of urban public transport operation management, and particularly relates to a charging scheduling optimization method for a hybrid power supply electric bus system. BACKGROUND
[0002] In order to reduce carbon emissions and operating costs of urban public transport systems, countries around the world are vigorously promoting the electrification of public transport vehicles. However, the electric energy consumed by electric buses is derived from the power grid, and under the power structure dominated by thermal power generation, the carbon emissions generated by electric buses due to the consumption of electric energy cannot be ignored. In order to further reduce the carbon emissions of electric bus systems, a photovoltaic- energy storage-grid (hereinafter referred to as light-storage-grid) hybrid power supply mode has emerged. Among them, photovoltaic power generation equipment converts solar energy into electric energy, energy storage equipment stores the electric energy generated by photovoltaic power generation to avoid the intermittency and volatility of photovoltaic power generation, and provides relatively stable electric energy supply for buses; and the power grid provides electric energy supplement for buses in the case of insufficient photovoltaic power generation. This new power supply mode not only reduces the dependence of electric buses on traditional power grids, but also avoids the impact of concentrated charging on power grids, which is of great significance to the sustainable development of urban transportation.
[0003] However, this multi-source hybrid power supply mode has intensified the complexity of bus charging scheduling, making the public transport system face new challenges while dealing with original problems such as the unfixed time point of return charging due to uncertain travel time and uneven distribution of charging demand. For example, the amount of photovoltaic power generation will fluctuate greatly with changes in time and weather conditions, making the clean electric energy available to vehicles have significant uncertainty. Although energy storage can buffer photovoltaic fluctuations to some extent, its capacity is limited, and if the charging strategy is not appropriate, it may still be unable to meet the real-time charging demand of buses or lead to idle photovoltaic power resources. In this case, charging scheduling methods based on a single power supply and ignoring the uncertainty of photovoltaic power generation are difficult to fully respond to dynamic changes in the external environment, and the generated bus charging scheme often cannot achieve optimal results in practice. Therefore, the present application considers the uncertainty brought by multi-energy coupling and proposes a charging scheduling method based on reinforcement learning to solve the charging scheduling problem of the "light-storage-grid" hybrid power supply electric bus system. SUMMARY
[0004] The purpose of the present application is to solve the problem of unreasonable resource allocation in existing electric bus systems, which leads to resource redundancy and waste, high operating energy consumption, and high operating costs, and to propose a hybrid power supply electric bus system charging scheduling optimization method based on reinforcement learning.
[0005] The specific process of the hybrid power supply electric bus system charging scheduling optimization method based on reinforcement learning is as follows:
[0006] Step 1, define state space;
[0007] Step 2, define action space;
[0008] Step 3, define reward function;
[0009] Step 4, based on state space, action space, reward function, establish multi-objective Markov decision model;
[0010] Step 5, based on multi-objective Markov decision model, construct reinforcement learning network model;
[0011] Step 6, obtain data set;
[0012] Step 7, based on the data set, train the reinforcement learning network model, obtain the trained reinforcement learning network model, and save the parameters;
[0013] Step 8, based on the trained reinforcement learning network model, obtain the charging scheme of the actual optimization period.
[0014] The beneficial effects of the present application are:
[0015] The present application models the bus charging scheduling problem under the "light-storage-grid" hybrid power supply mode as a multi-objective Markov decision model; considering the feature distribution deviation and strategy conflict problem caused by environmental difference, the charging scheduling task is divided into regular task in sufficient light and special task in insufficient light, and a multi-agent deep reinforcement learning solving algorithm with multi-task head architecture is designed; based on the actual bus system operation data and environmental data, charging scheduling strategies under various target preferences are trained.
[0016] The reinforcement learning charging scheduling optimization method proposed in the present application can directly map multi-source observation data (such as real-time light radiation, temperature, energy storage battery state of charge, vehicle arrival condition, etc.) to charging control instructions under different environmental conditions according to user preferences, minimize operating cost and carbon emissions, meet real-time charging scheduling demand, and solve the problems of resource redundancy waste, high operating cost and high operating energy consumption caused by unreasonable resource allocation in existing electric bus systems. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The flowchart of the present application. DETAILED DESCRIPTION
[0018] DETAILED DESCRIPTION
[0019] Step 1, define state space;
[0020] Step 2, define action space;
[0021] Step 3, define reward function;
[0022] Step 4, based on state space, action space, reward function, establish multi-objective Markov decision model;
[0023] Step 5, based on multi-objective Markov decision model, construct reinforcement learning network model;
[0024] Step 6, obtain data set;
[0025] Step 7, based on data set, train reinforcement learning network model, obtain trained reinforcement learning network model, and save parameters;
[0026] Step 8, based on trained reinforcement learning network model, obtain charging scheme of actual optimization period.
[0027] Specific implementation method two: the difference between this implementation method and the specific implementation method one is that: in the step 1, the state space is defined; the specific process is:
[0028] Step 1.1, let L represent the number of bus routes; N represents the number of public transport vehicles in the public transport system;
[0029] represents the average travel energy consumption of bus n on the line, ;
[0030] K represents the number of charging piles, represents the charging pile power;
[0031] A represents the photovoltaic layout area, represents the rated power of photovoltaic components;
[0032] represents the battery capacity of the energy storage device; represents the battery capacity of the bus;
[0033] unit: kWh; unit: kW; A unit: m 2 ; unit: kW; unit: kWh; unit: kWh;
[0034] Step 1.2, take as the starting time of the optimization period, as the end time of the optimization period, divide the optimization period into T steps, and each step lasts 1 minute;
[0035] Step 1.3,
[0036] Let denote whether bus n is on a trip at step t, and let if so, otherwise take the value 0;
[0037] Let denote whether bus n will depart from the depot at step t, and let if so, otherwise take the value 0;
[0038] Let denote the state of charge (SOC) of bus n at step t, in %;
[0039] Let denote the charging duration of bus n from the time of the last charge to the start of step t, and let if the bus has finished charging before the start of step t; in min;
[0040] denote the remaining time until the next departure of bus n (counted from the start of step t), in min;
[0041] Step 1.4, let denote the SOC of the energy storage device at step t, in %;
[0042] SOC denotes the state of charge of the battery;
[0043] denote the charging duration of the energy storage device at step t, in min;
[0044] Step 1.5, let denote the solar irradiance at step t, , in kW / m 2 ;
[0045] denote the ambient temperature at step t, in °C;
[0046] denote the power generated by the photovoltaic module at step t, in kW;
[0047] , (see equation (7)), denote the supply price of electricity from photovoltaics, energy storage, and the grid, respectively, 、 、 All units are in yuan / kWh;
[0048] Step 1.6, the bus and the energy storage device are regarded as independent agents;
[0049] Let denote the state of bus n at step t, The composition of
[0050] Let denote the state of the energy storage device at step t, The composition of
[0051] Let denote the environmental state of the bus system at step t, The composition of
[0052] (1)
[0053] (2)
[0054] (3)
[0055] Step 1.7, the state of the bus system at step t is represented by the states of all agents and the environmental state , as shown in equation (4);
[0056] (4)
[0057] In the formula, denotes vector splicing;
[0058] denotes the state of bus 1 at step t;
[0059] denotes the state of bus 2 at step t;
[0060] denotes the state of bus at step t;
[0061] All agents are a bus and an energy storage device;
[0062] Step 1.8, the state space is composed of the states of the bus system at T steps , as shown in equation (5);
[0063] (5)
[0064] wherein, denotes the bus system state of step 1, denotes the bus system state of step 2, denotes the bus system state of step .
[0065] The other steps and parameters are the same as in the first embodiment.
[0066] The third embodiment is different from the first or second embodiment in that the action space is defined in step 2. The specific process is as follows:
[0067] Step 2.1, use to denote the actual action (corresponding to the charging mode) of bus n at step t, since the power of the bus in the bus system is derived from photovoltaic, energy storage and power grid, there are 0, 1, 2, 3 values, respectively, indicating that bus n does not charge at step t, uses photovoltaic to charge, uses energy storage to charge, and uses power grid to charge;
[0068] Step 2.2, use to denote the actual action (corresponding to the charging mode) of the energy storage device at step t, considering that the power of the energy storage is derived from photovoltaic and power grid, there are 0, 1, 2 values, respectively, indicating that the energy storage device does not charge at step t, uses photovoltaic to charge, and uses power grid to charge;
[0069] Step 2.3, use the actual action of all agents to jointly represent the actual action set at step t (corresponding to the charging scheme of the bus system at step t), as shown in equation (6):
[0070] (6)
[0071] wherein, denotes the actual action of bus 1 at step t, denotes the actual action of bus 2 at step t, denotes the actual action of bus at step t.
[0072] The other steps and parameters are the same as in the first or second embodiment.
[0073] The fourth embodiment is different from one of the first to third embodiments in that the reward function is defined in step 3. The specific process is as follows:
[0074] Step 3.1, define the calculation method of the unit price of energy storage charging , is equal to the ratio of total charging cost to total charging amount at the end of step t-1, as shown in equation (7):
[0075] (7)
[0076] In the formula,
[0077] is the charging efficiency of the charging pile, and the recommended value is 0.95;
[0078] is an indicator function, which is equal to 1 only when the condition in the parentheses is true , otherwise it is equal to 0; for example, when , is equal to 1, otherwise it is equal to 0;
[0079] is the step is the charging power when charging the energy storage device, wherein is the minimum charging power for charging the energy storage device, and the recommended value is 40, in units of kWh;
[0080] represents the actual action of the energy storage device in step ;
[0081] represents the supply price of the power grid;
[0082] Step 3.2, use to represent the total charging cost of the public transportation system in step t, is equal to the sum of the charging cost of all buses in the public transportation system in step t and the charging cost of the energy storage device, as shown in equation (8):
[0083] (8)
[0084] is represented as:
[0085] (9)
[0086] is represented as:
[0087] (10)
[0088] In the formula, is the charging power when charging the energy storage device in step t;
[0089] in Yuan; in Yuan; in Yuan;
[0090] Step 3.3, the CO2 emission of the photovoltaic power in the production process of the bus system in the step t is represented as: the CO2 emission of the photovoltaic power in the production process of the bus system in the step t is represented as: the CO2 emission of the photovoltaic power in the production process of the bus system in the step t is represented as: the CO2 emission of the photovoltaic power in the production process of the bus system in the step t is represented as:
[0091] (11)
[0092] is represented as:
[0093] (12)
[0094] is represented as:
[0095] (13)
[0096] wherein, , are the photovoltaic carbon emission factor and the energy storage carbon emission factor, respectively;
[0097] , in kgCO2 / kWh; in kgCO2; in kgCO2; in kgCO2;
[0098] Step 3.4, a constraint penalty term is designed for the actual operation constraints of the bus system;
[0099] The specific process is as follows:
[0100] Step 3.4.1, a charging location constraint penalty term is constructed to limit the charging of the bus to the bus station only, the calculation method is shown in formula (14):
[0101] (14)
[0102] wherein, is the recommended action obtained by the bus n in the step t, and the obtaining method is shown in step 5.5;
[0103] Step 3.4.2, a battery state of charge constraint penalty term is constructed , , , respectively limit the SOC of the bus battery and the energy storage battery to vary between the lower and upper SOC limits, , The calculation method is shown in equations (15) and (16):
[0104] (15)
[0105] (16)
[0106] wherein, and are the lower and upper SOC limits of the bus battery, respectively, in %;
[0107] and are the lower and upper SOC limits of the energy storage battery, respectively, in %;
[0108] The intermediate variable is used to represent ;
[0109] Step 3.4.3, constructing a charging time constraint penalty term to limit the single charging time of the bus to be no less than the minimum charging time of the bus , in min, The calculation method is shown in equation (17):
[0110] (17)
[0111] wherein, represents the minimum charging time of the bus;
[0112] Step 3.4.4, constructing a charging bus number constraint penalty term to limit the number of buses charging at the bus station at the same time to be no more than the number of charging piles, The calculation method is shown in equation (18):
[0113] (18)
[0114] is compared with ;
[0115] Step 3.4.5, constructing a charging power constraint penalty term to limit the photovoltaic power used by the bus to be no more than the generated power, The calculation method is shown in equation (19):
[0116] (19)
[0117] judge After adding, multiply After Compare;
[0118] Step 3.5, use represents the overall situation of constraint violation at time step t, The calculation method of is shown in formula (20):
[0119] (20)
[0120] Step 3.6: The construction of the reward function needs to correspond to the design of the optimization target. The optimization target of this invention is to minimize the daily charging cost of the bus system and minimize CO2 emissions; therefore, a two-dimensional reward function is constructed. , as shown in formula (21), Incentive items based on charging costs and CO2 emission incentives composition, as well as The calculation method is shown in formulas (22) and (23):
[0121] (twenty one)
[0122] (twenty two)
[0123] (twenty three).
[0124] The other steps and parameters are the same as those in the first to third embodiments.
[0125] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that a multi-objective Markov decision model is established in step 4; the specific process is as follows:
[0126] Step 4.1, use represents the discount factor, and a recommended value is 0.99. In addition, since the reinforcement learning charging scheduling algorithm involved in the present invention can directly approximate the optimal strategy through interaction with the environment, the state transition probability is not explicitly defined.
[0127] Step 4.2: Based on the above content, the bus charging scheduling problem under the "solar-storage-grid" hybrid energy supply mode is described as consisting of the state space (obtained in step 1), the action space (obtained in step 2), the reward function (obtained in step 3), and the discount factor A multi-objective (two objectives, corresponding to minimizing the daily charging cost of the bus transit system and minimizing the CO2 emission) Markov decision model is constructed.
[0128] The other steps and parameters are the same as those in Embodiments One to Four.
[0129] Embodiment Six: Different from Embodiments One to Five, in this embodiment, the reinforcement learning network model is constructed based on a multi-objective Markov decision model, and the specific process is as follows:
[0130] Step 5.1: Selecting daily average solar radiation as the division condition of the charging scheduling task, the charging scheduling task is divided into a regular task in sufficient light and a special task in insufficient light.
[0131] The sufficient light refers to the daily average power generation of photovoltaics being greater than .
[0132] The insufficient light refers to the daily average power generation of photovoltaics being less than or equal to .
[0133] Step 5.2: Building a shared network for multiple tasks, the shared network is used to extract common features in the input information, and the shared network is composed of three layers of MLP, GRU and MLP.
[0134] The input information of the shared network is a complex vector composed of seven kinds of information of the state of bus n at step t (agent state, see Step 1), the environment state of the bus transit system at step t (environment state, see Step 1), the actual action of bus n at step t (historical action, see Step 2), the overall situation of constraint violation (see Step 3), the date (the year, month and day information corresponding to the current charging scheduling task), the step and the target preference vector (two-dimensional vector, the two elements are in the interval
[0135] (24)
[0136] (25)
[0137] (26)
[0138] where, is an intermediate variable of the shared network;
[0139] denotes the step size is the feature vector of the corresponding gated recurrent unit (GRU) output;
[0140] denotes vector concatenation; MLP denotes a multi-layer perceptron network;
[0141] denotes the step size is the feature vector of the corresponding gated recurrent unit (GRU) output; is a zero vector; GRU denotes a gated recurrent unit;
[0142] denotes the output result of the shared network;
[0143] The multi-layer perceptron network MLP includes an input layer, a hidden layer, and an output layer, and the number of neurons in the hidden layer is denoted by B, which is recommended to be 128; formulas (27)-(29) show the information flow in the MLP network.
[0144] (27)
[0145] (28)
[0146] (29)
[0147] The simplified form of the MLP network is shown in formula (30):
[0148] (30)
[0149] where, is the feature vector input to the MLP;
[0150] and is an intermediate variable in the MLP;
[0151] , , are the weight matrices of the input layer, the hidden layer, and the output layer of the MLP, respectively;
[0152] , are the bias terms of the input layer, the hidden layer, and the output layer of the MLP, respectively (all weight matrices and bias terms are automatically updated during network training);
[0153] is the output result of the MLP;
[0154] represents an activation function;
[0155] Step 5.3, build a specific task-oriented parallel task head network for optimization target 1 and optimization target 2; the specific process is:
[0156] The specific task-oriented parallel task head network includes two structurally identical parallel task head networks, corresponding to the two optimization targets of the model, and can generate local action function values (which can be understood as long-term cumulative rewards) in the corresponding task scene according to the extracted common features;
[0157] The two structurally identical parallel task heads are a regular task head and a special task head; the structure is a multi-layer perceptron network MLP;
[0158] Formula (31) shows the information flow of bus n in the two structurally identical parallel task head networks for optimization target 1:
[0159] (31)
[0160] Formula (32) shows the information flow of bus n in the two structurally identical parallel task head networks for optimization target 2:
[0161] (32)
[0162] In the formula,
[0163] represents the output result of the task head network, and represents the local action Q vector of bus n under optimization target 1 at step t, which contains the local action function values that bus n can obtain by taking four kinds of actions at step t;
[0164] Optimization target 1 is the total charging cost of the bus system at step t ;
[0165] is the task type; is a regular task;
[0166] represents the output result of the task head network, and represents the local action Q vector of bus n under optimization target 2 at step t, which contains the local action function values that bus n can obtain by taking four kinds of actions at step t;
[0167] Optimization target 2 is the CO2 emission of the power used by the bus system in the production process at step t ;
[0168] Step 5.4, set a weighted greedy strategy to select a recommended action based on the local action Q vector;
[0169] The specific process is as follows:
[0170] Taking the bus agent n as an example, after obtaining the local action Q vector, the agent first calculates the comprehensive local action Q vector by using formula (33) for weighted calculation ; is expressed as:
[0171] (33)
[0172] In the formula, and are two elements in the target preference vector ;
[0173] Subsequently, the bus n will select the element with the maximum value in , and the action corresponding to the element with the maximum value is the recommended action of the bus n at step t (taking values 0, 1, 2, 3, corresponding to different charging modes); as shown in formula (34):
[0174] (34)
[0175] In the formula, is the action corresponding to the element with the maximum value ;
[0176] Step 5.5, calculate the total number of record violations of the recommended action selected in step 5.4 ; and correct the recommended action that does not meet the constraint condition through the action correction strategy to obtain the actual action set;
[0177] The specific process is as follows:
[0178] Step 5.5.1, obtain the set of buses whose action values at step t are not 0 (the bus whose action value is not 0 is the charging bus, which is determined in step 5.4);
[0179] Step 5.5.2, calculate the charging priority of all vehicles in the set based on formula (35):
[0180] (35)
[0181] In the formula, represents the charging priority of the bus n at step t;
[0182] Step 5.5.3, for the charging location constraint, set the action of the bus in the driving state to 0, and set the charging duration of the bus in the driving state to 0;
[0183] Step 5.5.4, for the battery state of charge constraint, set the action of the bus whose battery SOC reaches the upper limit of SOC to 0; set the action of the bus whose battery SOC is less than to 1, and adjust the charging priority of the bus whose battery SOC is less than to the maximum value , which is recommended to be 5000;
[0184] Step 5.5.5, for the charging time constraint, set the action of the bus whose charging duration is but the action is 0 to 1, and adjust the charging priority of the bus whose charging duration is to ;
[0185] Step 5.5.6, for the charging bus number constraint, when the number of buses charging simultaneously is greater than the number of charging piles, allocate charging piles to buses according to the charging priority from high to low until the charging pile resources are exhausted, and set the action of the bus not allocated to the charging pile to 0;
[0186] If there are buses with the same charging priority, randomly select a bus and allocate a charging pile (such as the bus whose priority is reached at this time, this priority has 5 buses, and each is allocated a charging pile if the number of charging piles is sufficient, and a charging pile is randomly selected and allocated to a bus if the number of charging piles is insufficient);
[0187] Step 5.5.7, consider the photovoltaic power generation situation and the storage SOC to process the charging power constraint;
[0188] Step 5.5.8, set the recommended action of all agents to the actual action to determine the actual action set;
[0189] Step 5.6, build a parallel hybrid network for multiple tasks, which also has two parallel networks with the same structure to correspond to the two optimization objectives of the model; the specific process is as follows:
[0190] The hybrid network is composed of a weight parameter network and a value function hybrid network;
[0191] Among them, the weight parameter network is composed of four super networks; the value function hybrid network is composed of two linear networks;
[0192] The input of the four networks in the weight parameter network is the bus system state at step t , and the output corresponds to the weights of the two linear networks respectively 、 ( ), and deviation 、 ; Among them, except is calculated using a single-layer linear network, and the remaining parameters are obtained by fitting a double-layer linear network combined with the ReLU function; formulas (36)-(39) show the information flow in the weight parameter network;
[0193] (36)
[0194] (37)
[0195] (38)
[0196] (39)
[0197] Where,
[0198] 、 is the weight matrix of hypernetwork 1; 、 is the weight matrix of hypernetwork 2; is the weight matrix of hypernetwork 3; 、 is the weight matrix of hypernetwork 4;
[0199] 、 is the bias term of hypernetwork 1; 、 is the bias term of hypernetwork 2; is the bias term of hypernetwork 3; 、 is the bias term of hypernetwork 4;
[0200] Indicates the target, Indicates target 1, Indicates target 2;
[0201] Input information of the value function hybrid network It is composed of the local action function values corresponding to the actual actions of each agent. Taking bus agent 1 as an example, agent 1 will move from and Extract the corresponding value and place it in and The output of the value function hybrid network is the joint action Q value under each target of step size t ;
[0202] The information flow in the value function hybrid network is expressed as follows in combination with the output result of the weight parameter network:
[0203] (40)
[0204] (41)
[0205] wherein, is an intermediate variable of the value function hybrid network;
[0206] is an absolute value operation, and the non-negative weight makes the local action value function network and the joint value function network be able to be optimized in the same direction during the training process;
[0207] represents summing up the elements of a vector.
[0208] The combination of the shared network, the parallel task head network, the weighted greedy strategy, the action correction strategy and the parallel hybrid network is a multi-task head multi-agent deep reinforcement learning algorithm (referred to as M3RL) framework.
[0209] The other steps and parameters are the same as those in the first to fifth embodiments.
[0210] The seventh embodiment is different from one of the first to sixth embodiments in that the photovoltaic power generation condition and the energy storage SOC are considered at the same time in the step 5.5.7 to process the charging power constraint, and the specific process is as follows:
[0211] Step 5.5.7.1, obtaining a set of vehicles using photovoltaic charging at the step t (step 5.4 determines), the set of vehicles using energy storage charging is defined as (step 5.4 determines), step 5.5.7.2 is executed;
[0212] Step 5.5.7.2, based on formula (42), the total charging capacity of the buses using photovoltaic charging within the step t is calculated (step 5.5.7.1), step 5.5.7.3 is executed;
[0213] (42)
[0214] Step 5.5.7.3, if , a bus is randomly selected in , the recommended action of the randomly selected bus is replaced with energy storage charging, and (step 5.5.7.2) is reacquired, and step 5.5.7.4 is executed; otherwise, step 5.5.7.6 is entered;
[0215] Step 5.5.7.4, calculate the total charging amount of the bus using the energy storage device in the step t based on formula (43) ; if , replace the recommended action of the randomly selected bus in step 5.5.7.3 with charging using the power grid, and execute step 5.5.7.5; otherwise, execute step 5.5.7.6;
[0216] (43)
[0217] Step 5.5.7.5, return to step 5.6.7.1 (at this time, the data of step 5.6.7.1 is also updated);
[0218] Step 5.5.7.6, calculate based on formula (43) ; if , randomly select a bus in , replace the recommended action of the randomly selected bus with charging using the power grid, and execute step 5.5.7.7; otherwise, go to step 5.5.7.8;
[0219] Step 5.5.7.7, reacquire , and return to step 5.5.7.6;
[0220] Step 5.5.7.8, calculate the remaining photovoltaic power based on formula (44) ; if and , set the action of the energy storage device to 1, and set the charging power ; otherwise, set the action of the energy storage device to zero;
[0221] (44).
[0222] The other steps and parameters are the same as those in specific embodiments one to six.
[0223] Specific embodiment eight: this embodiment is different from one of specific embodiments one to seven in that the data set is obtained in step 6; the specific process is as follows:
[0224] Step 6.1,
[0225] Collect (such as continuously collecting for one year) the state of buses at the step t, the state of the energy storage device at the step t, and the environmental state of the bus system at the step t;
[0226] Step 6.2, use the collected data as the training set;
[0227] Step 6.3: Set up a sampling method that combines sequential sampling and random sampling; specifically:
[0228] When the number of training times is less than 50, data is input in the order of spring, summer, autumn, and winter to help the agent learn seasonal patterns. When the number of training times is greater than or equal to 50, random input is used to enhance the generalization ability of the reinforcement learning network model in step 5.
[0229] The other steps and parameters are the same as those in the first to seventh embodiments.
[0230] Specific embodiment 9: This embodiment differs from any one of specific embodiments 1 to 8 in that: in step 7, a reinforcement learning network model is trained based on the data set to obtain a trained reinforcement learning network model and save the parameters; the specific process is:
[0231] Step 7.1: Set two target preference subspaces according to the number of optimization targets (2), which are:
[0232] ;
[0233] ;
[0234] is the sampling interval of the target preference subspace, The recommended value is 0.1;
[0235] Step 7.2: Construct two reinforcement learning network models as described in step 5 as the evaluation network and target network respectively;
[0236] Set the two lengths to Experience pools are used to store experience data for regular tasks and special tasks respectively. ; The recommended value is 5000;
[0237] in, Indicates step length Action; Indicates step length Action; Indicates step length the status of the public transportation system; Indicates step length the status of the public transportation system; Indicates step length rewards;
[0238] Step 7.3
[0239] Set the maximum number of iterations (Recommended value 2000), minimum update iteration number (Recommended value 50), batch sample data volume (Recommended value 32), learning rate (Recommended value );
[0240] Set initial value, increment, final value of special experience sampling ratio of general task , , (Recommended value 0.1, 0.01, 0.2);
[0241] Set initial value, increment, final value of special experience sampling ratio of special task , , (Recommended value 0.2, 0.05, 0.8);
[0242] Randomly initialize evaluation network parameters (including weights and bias terms), and copy the evaluation network parameters to the target network;
[0243] Step 7.4, sample from the training set, activate the task head network of the general task or the special task according to the sampling data type, and obtain the current task date; randomly obtain a set of preference vectors from the target preference subspace;
[0244] Step 7.5, each agent observes the environment and returns the corresponding state and action information, obtains the recommended action based on the shared network, the task-specific parallel task head network for optimization target 1 and optimization target 2, and the weighted greedy strategy, and calculates the total constraint violation of the recommended action;
[0245] Step 7.6, each agent uses the action correction strategy to correct the recommended action that does not meet the constraint condition to obtain an actual action set, and the actual action set is returned to the environment for execution, and the environment feedbacks a reward vector and moves to the next state;
[0246] Step 7.7, the action of step , the action of step , the bus system state of step , the bus system state of step , and the reward of step are stored in the experience pool corresponding to the general task or the special task as experience;
[0247] Step 7.8, when the number of experiences in the experience pool exceeds the threshold (recommended value 100), extract sample experience data from the experience pool to update the reinforcement learning network model parameters described in step 5;
[0248] When the number of iterations is less than When , the sample experience data only comes from the regular task experience pool;
[0249] When the number of iterations is greater than or equal to When, for routine tasks, the sample experience data will be 、 The proportion is randomly obtained from the regular task experience pool and the special task pool;
[0250] When the number of iterations is greater than or equal to When targeting special tasks, sample experience data will be 、 The proportion is randomly obtained from the regular task experience pool and the special task pool;
[0251] Every In the iteration, for conventional tasks, the reinforcement learning network model will Increase the sampling ratio of special experience in regular tasks and special tasks (the sample experience data will be based on 、 The ratio is randomly obtained from the regular task experience pool and the special task pool) until the preset maximum value is reached;
[0252] Every In the iteration, for special tasks, the reinforcement learning network model will Increase the sampling ratio of special experience in regular tasks and special tasks (the sample experience data will be based on 、 The ratio is randomly obtained from the regular task experience pool and the special task pool) until the preset maximum value is reached;
[0253] It is a constant, and the recommended value is 10;
[0254] Step 7.9: Based on the sample empirical data, calculate the loss function according to formulas (45)-(47). ; expressed as:
[0255] (45)
[0256] (46)
[0257] (47)
[0258] Where,
[0259] It is The loss value of each target under the current batch of sample experience;
[0260] is the sample experience The joint action Q value under the first target in the evaluation network is evaluated The corresponding action is the actual action set recorded in the experience pool;
[0261] is the evaluation network parameter
[0262] is the sample experience The joint action Q value under the first target in the evaluation network is evaluated The corresponding action is the actual action set recorded in the experience pool;
[0263] is the target network parameter
[0264] is the sample experience The reward value of the first target in the evaluation network is evaluated The corresponding action is the actual action set recorded in the experience pool;
[0265] is the sample experience The joint action Q value under the first target in the evaluation network is evaluated The corresponding action is the actual action set recorded in the experience pool;
[0266] Step 7.10, based on The gradient is calculated in the way of back propagation, and the Adam optimizer is used to update the parameters in the reinforcement learning solution network model described in step 5;
[0267] Step 7.11, remove the dominated experience in the experience pool;
[0268] Dominated means that under the same local observation condition, the current solution is not better than another solution on all targets, and is strictly worse than another solution on at least one target function;
[0269] Step 7.12, increase the sampling probability of the experience sparse preference subspace (2 optimization targets, such as more biased to the first target, i.e. The data points corresponding to the weight of optimization target 1 are 20, which are more biased to the second target, i.e. The data points corresponding to the weight of optimization target 1 are 80, which are more biased to the second target, i.e.
[0270] Step 7.13, every time step 7.4-7.12 The evaluation network parameters are copied to the target network;
[0271] is a constant, and a value of 10 is recommended;
[0272] Step 7.14, repeat steps 7.4 to step 7.13 until the number of iterations reaches the preset maximum number of iterations, obtain the trained reinforcement learning solution network model, and save the parameters.
[0273] The verification set (in the data taken continuously for one year, it will be divided into two independent parts. One part is set as the training set to train the model parameters. Another part is the data that the model has not seen and has not been trained. However, since the source is consistent, the overall distribution is consistent with the training set, so this part of the data can verify the model effect) is used to evaluate the network performance. If the generated charging scheduling scheme can reduce the operation cost and carbon emission while meeting the real-time charging scheduling demand of electric buses, the final parameters of the shared network and the parallel task head network in the evaluation network are output and saved.
[0274] The other steps and parameters are the same as those in the first to eighth embodiments.
[0275] The tenth embodiment is different from one of the first to ninth embodiments in that the charging scheme of the actual optimization period is obtained based on the trained reinforcement learning network model in step 8. The specific process is as follows:
[0276] Step 8.1, determine the actual optimization period and the user target preference weight, and load the finally saved model parameters before starting optimization;
[0277] Step 8.2, obtain the state information of the bus and the energy storage, and input the state information together with the target preference weight into the shared network of the trained reinforcement learning network model. The shared network outputs the result to the task head network, and obtains the actual action set by combining the weighted greedy strategy and the action correction strategy. The actual action set is converted into a charging scheme for execution;
[0278] Step 8.3, repeat step 8.2 until the end of the optimization period.
[0279] The other steps and parameters are the same as those in the first to ninth embodiments.
[0280] The present application can also have other various embodiments, and those skilled in the art can make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application. However, these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.
Claims
1. A reinforcement learning-based charging scheduling optimization method for a hybrid electric bus system, characterized by: The specific process of the method is: Step 1: Define the state space; Step 2: Define the action space; Step 3: Define the reward function. Step 4: Based on the state space, action space, and reward function, a multi-objective Markov decision model is established; Step 5: Construct a reinforcement learning network model based on the multi-objective Markov decision model; Step 6: Get the dataset; Step 7: Train the reinforcement learning network model based on the data set to obtain the trained reinforcement learning network model and save the parameters; Step 8: Obtain the charging plan for the actual optimization period based on the trained reinforcement learning network model.
2. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 1 is characterized in that: In step 1, the state space is defined; the specific process is: Step 1.1, let L be the number of bus routes; N be the number of buses in the bus system; represents the average travel energy consumption of bus n on the route, ; K represents the number of charging piles, Indicates the charging pile power; A represents the photovoltaic layout area, Indicates the rated power of the photovoltaic module; Indicates the battery capacity of the energy storage device; Indicates the battery capacity of the bus; Step 1.2, is the starting time of the optimization period, To optimize the end time of the optimization period, divide the optimization period into T steps, each step lasting 1 minute; Step 1.3 make Indicates whether bus n is on the journey in step t. If so, let , otherwise the value is 0; make Indicates whether bus n will start from the station in step t. If so, let , otherwise the value is 0; make represents the battery state of charge of bus n at step length t; make represents the charging duration of bus n from the last charging start time to the start time of step t. If the bus has finished charging before the start of step t, then let ; Indicates the remaining time until the next departure of bus n; Step 1.4, use represents the SOC of the energy storage device at step length t; SOC represents the state of charge of the battery; represents the charging duration of the energy storage device at step t; Step 1.5, use represents the illumination radiance with step length t, ; represents the ambient temperature with a step length of t; represents the power generated by the PV module in step t; 、 、 Represent the unit prices of photovoltaic, energy storage, and grid power supply respectively; Step 1.6: Treat the bus and energy storage device as independent intelligent entities; use represents the state of bus n at step t, The composition of is shown in formula (1); use represents the state of the energy storage device at step t, The composition of is shown in formula (2); use represents the environmental state of the bus system at step length t, The composition of is shown in formula (3); (1) (2) (3) Step 1.7: Use the states of all agents and the environment to represent the state of the bus system with a step length of t. , as shown in formula (4); (4) Where, Represents vector concatenation; represents the state of bus 1 at step length t; represents the state of bus 2 at step length t; Indicates bus The state at step length t; All agents are A bus and an energy storage device; Step 1.8: The state space is formed by the bus system states of T steps , as shown in formula (5); (5) in, represents the bus system state with step size 1, represents the bus system state with step size 2, Indicates step length The status of the public transportation system.
3. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 2, characterized in that: In step 2, the action space is defined; the specific process is as follows: Step 2.1, use represents the actual action of bus n in step t, There are four values: 0, 1, 2, and 3, which respectively indicate that bus n is not charged, charged using photovoltaics, charged using energy storage, and charged using the grid in step t; Step 2.2, use represents the actual action of the energy storage device at step t, There are three values: 0, 1, and 2, which respectively indicate that the energy storage device is not charged, charged using photovoltaic power, and charged using the grid at step t; Step 2.3: Use the actual actions of all agents to represent the actual action set of step length t , as shown in formula (6): (6) in, represents the actual action of bus 1 in step t, represents the actual action of bus 2 in step t, Indicates bus The actual action at step size t.
4. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 3 is characterized in that: In step 3, the reward function is defined; the specific process is: Step 3.1: Define the energy storage charging unit price The calculation method of It is equal to the ratio of the total charging cost to the total charging capacity at the end of step t-1, as shown in formula (7): (7) Where, is the charging efficiency of the charging pile; Is an indicator function, only when the conditions in the brackets are met The value is 1, otherwise the value is 0; is the step length Charging power when charging the energy storage device, ,in is the minimum charging power for charging the energy storage device; Indicates that the energy storage device is in step length the actual action; Indicates the power supply unit price of the power grid; Step 3.2, use represents the total charging cost of the bus system at step t, Equal to the charging cost of all buses in the bus system with step length t Charging costs with energy storage devices The sum of , as shown in formula (8): (8) Expressed as: (9) Expressed as: (10) Where, is the charging power when the step length t is the energy storage device charging; Step 3.3, use represents the CO2 emissions from the electricity used by the public transportation system in the production process at step t, Equal to the CO2 emissions from photovoltaic power generation in the bus system with a step length of t CO2 emissions from grid electricity production The sum of , as shown in formula (11): (11) Expressed as: (12) Expressed as: (13) Where, 、 They are the photovoltaic carbon emission factor and the energy storage carbon emission factor; Step 3.4: Design constraint penalty items based on the actual operation constraints of the public transportation system; The specific process is: Step 3.4.1: Construct the charging location constraint penalty term , The calculation method of is shown in formula (14): (14) Where, is the recommended action obtained by bus n at step t; Step 3.4.2: Construct battery state of charge constraint penalty term 、 , 、 The calculation method is shown in formulas (15) and (16): (15) (16) Where, and They are the lower and upper limits of the SOC of the bus battery; and They are the lower and upper limits of the SOC of the energy storage battery; Using intermediate variables express ; Step 3.4.3: Construct charging time constraint penalty , The calculation method of is shown in formula (17): (17) Where, Indicates the minimum charging time of the bus; Step 3.4.4: Construct a penalty term for the number of charging buses , The calculation method of is shown in formula (18): (18) Step 3.4.5: Construct charging power constraint penalty term , The calculation method of is shown in formula (19): (19) Step 3.5, use represents the overall situation of constraint violation at time step t, The calculation method of is shown in formula (20): (20) Step 3.6: Construct a two-dimensional reward function , as shown in formula (21), Incentive items based on charging costs and CO2 emission incentives composition, as well as The calculation method is shown in formulas (22) and (23): (21) (22) (23)。 5. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 4 is characterized in that: In step 4, a multi-objective Markov decision model is established; the specific process is: Step 4.1, use represents the discount factor; Step 4.2: From the state space, action space, reward function and discount factor Constitute a multi-objective Markov decision model.
6. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 5, characterized in that: In step 5, a reinforcement learning network model is constructed based on a multi-objective Markov decision model; the specific process is as follows: Step 5.1, divide the charging scheduling task into a regular task when there is sufficient sunlight and a special task when there is insufficient sunlight; Sufficient sunlight refers to the average daily photovoltaic power generation power greater than ; Insufficient sunlight refers to the average daily photovoltaic power generation power being less than or equal to ; Step 5.2: Build a shared network for multiple tasks. The shared network is used to extract common features from the input information. The shared network consists of three layers: MLP, GRU, and MLP. The input information of the shared network is the state of bus n at step t , the environmental state of the bus system at step length t ,the bus In step length Actual action , Overall situation of constraint violations ,date , step length and the target preference vector The complex vector is composed of seven types of information; the output is a vector containing the common features of the input information; as shown in formulas (24)-(26): (24) (25) (26) Where, is the intermediate variable of the shared network; Indicates step length The feature vector output by the corresponding gated recurrent unit GRU; Represents vector concatenation; MLP represents multi-layer perceptron network; Indicates step length The feature vector output by the corresponding gated recurrent unit GRU; is a zero vector; GRU stands for gated recurrent unit; Represents the output result of the shared network; The multilayer perceptron network MLP includes a three-layer network structure of input layer, hidden layer and output layer, and B represents the number of neurons in the hidden layer; formulas (27)-(29) show the information flow in the MLP network; (27) (28) (29) The simplified form of the MLP network is shown in formula (30): (30) Where, is the feature vector input to the MLP; and It is the intermediate variable in MLP; 、 、 are the weight matrices of the input layer, hidden layer, and output layer of MLP respectively; 、 are the bias terms of the input layer, hidden layer, and output layer of MLP respectively; is the output of MLP; express Activation function; Step 5.3: Build a parallel task head network for specific tasks for optimization goals 1 and 2. The specific process is as follows: The parallel task head network for specific tasks includes two parallel task head networks with the same structure; The two parallel task heads with the same structure are the regular task head and the special task head; The structure is a multi-layer perceptron network MLP; Formula (31) shows the information flow in two parallel task head networks with the same structure for the optimization goal 1 bus n: (31) Formula (32) shows the information flow in two parallel task head networks with the same structure for the optimization goal 2 bus n: (32) Where, represents the output of the task head network, which represents the local action Q vector of bus n under the optimization goal 1 with a step size of t, and contains the local action function values that can be obtained by bus n taking four actions in step size t; Optimization goal 1 is the total charging cost of the bus system at step t ; is the task type; It is a routine task; represents the output of the task head network, which represents the local action Q vector of bus n under optimization objective 2 with step size t, and contains the local action function values that can be obtained by bus n taking four actions with step size t; Optimization goal 2 is the CO2 emissions of the electricity used by the bus system in the production process at step t ; Step 5.4: Set up a weighted greedy strategy to select a recommended action based on the local action Q vector. The specific process is: First, use formula (33) to calculate the weighted comprehensive local action Q vector ; expressed as: (33) Where, and They are the target preference vectors Two elements in ; Then bus n will choose The element with the largest value in the , the action corresponding to the element with the largest value is the recommended action for bus n at step length t ; As shown in formula (34): (34) Where, It is used to take the action corresponding to the maximum value element ; Step 5.5: Calculate the total number of records that violate the constraints for the recommended actions selected in step 5.
4. ; And correct the suggested actions that do not meet the constraints through the action correction strategy to obtain the actual action set; The specific process is: Step 5.5.
1. Obtain the set of buses whose step length t action value is not 0 ; Step 5.5.2: Calculate the set based on formula (35) Charging priority for all vehicles in: (35) Where, represents the charging priority of bus n at step t; Step 5.5.3: For the charging location constraint, set the action of the bus in the driving state to 0, and set the charging duration of the bus in the driving state to 0; Step 5.5.4: For the battery state of charge constraint, set the action of the bus whose battery SOC reaches the upper limit to 0; set the action of the bus whose battery SOC is less than The action of the bus is set to 1, and the battery SOC is less than The bus charging priority is adjusted to the maximum ; Step 5.5.5: For the charging time constraint, set the charging duration to But the action of the bus with action 0 is set to 1, and all charging durations are in The charging priority of buses is adjusted to ; Step 5.5.6: For the constraint on the number of charging buses, when the number of buses charging simultaneously is greater than the number of charging piles, assign charging piles to the buses according to the charging priority from high to low until the charging pile resources are exhausted, and set the actions of the buses not assigned to charging piles to 0; If there are cases where the charging priority is the same, the bus will be randomly selected and assigned a charging pile; Step 5.5.7: Consider both the photovoltaic power generation and the energy storage SOC to process the charging power constraint; Step 5.5.8: Set the recommended actions of all current agents as actual actions and determine the actual action set; Step 5.6: Build a hybrid network; the specific process is: The hybrid network consists of a weight parameter network and a value function hybrid network; Among them, the weight parameter network consists of four super networks; the value function hybrid network consists of two linear networks; The inputs of the four networks in the weight parameter network are all the bus system states with a step size of t , the outputs correspond to the weights of the two linear networks 、 , and deviation 、 ; Formulas (36)-(39) show the information flow in the weight parameter network; (36) (37) (38) (39) Where, 、 is the weight matrix of hypernetwork 1; 、 is the weight matrix of hypernetwork 2; is the weight matrix of hypernetwork 3; 、 is the weight matrix of hypernetwork 4; 、 is the bias term of hypernetwork 1; 、 is the bias term of hypernetwork 2; is the bias term of hypernetwork 3; 、 is the bias term of hypernetwork 4; Indicates the target, Indicates target 1, Indicates target 2; Input information of the value function hybrid network It is composed of the local action function values corresponding to the actual actions of each agent; the output of the value function hybrid network is the joint action Q value under each target of step size t ; Combined with the output results of the weight parameter network, the information flow in the value function hybrid network can be expressed as follows: (40) (41) Where, It is the intermediate variable of the value function hybrid network; It is an absolute value operation; Represents the sum of vector elements.
7. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 6, characterized in that: In step 5.5.7, the charging power constraint is processed by taking into account both the photovoltaic power generation situation and the energy storage SOC. The specific process is as follows: Step 5.5.7.
1. Obtain the set of vehicles using photovoltaic charging with a step length of t The set of vehicles charged using energy storage is defined as , execute step 5.5.7.2; Step 5.5.7.2: Based on formula (42), calculate the total charging capacity of the bus using photovoltaic charging within step t , execute step 5.5.7.3; (42) Step 5.5.7.3, if , then in Randomly select a bus and replace the recommended action of the randomly selected bus with charging with energy storage to obtain , execute step 5.5.7.4; Otherwise, proceed to step 5.5.7.6; Step 5.5.7.
4. Based on formula (43), calculate the total charge capacity of the bus charged by the energy storage device within the step t ;like , then replace the recommended action of the randomly selected bus in step 5.5.7.3 with charging from the grid and execute step 5.5.7.5; otherwise, execute step 5.5.7.6; (43) Step 5.5.7.5, return to step 5.6.7.1; Step 5.5.7.6: Calculate based on formula (43) ;like , then in Randomly select a bus and replace the recommended action of the randomly selected bus with charging from the grid and execute step 5.5.7.7; otherwise, proceed to step 5.5.7.8; Step 5.5.7.7, Re-acquire , return to step 5.5.7.6; Step 5.5.7.8: Calculate the remaining photovoltaic power based on formula (44) ;like and When the action of the energy storage device is set to 1, the charging power Set to ; Otherwise, the action of the energy storage device is set to zero; (44)。 8. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 7, characterized in that: The data set is obtained in step 6; the specific process is: Step 6.1: Collect the actual bus station information The state of a bus at step t , the state of the energy storage device at step length t , the environmental state of the bus system at step length t ; Step 6.2, use the collected data as the training set; Step 6.3: Set up a sampling method that combines sequential sampling and random sampling; specifically: When the number of training times is less than 50, data is input in the order of spring, summer, autumn and winter; when the number of training times is greater than or equal to 50, data is input randomly.
9. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 8, characterized in that: In step 7, the reinforcement learning network model is trained based on the data set to obtain the trained reinforcement learning network model and save the parameters. The specific process is as follows: Step 7.1: Set two target preference subspaces according to the number of optimization targets, which are: ; ; is the sampling interval of the target preference subspace; Step 7.2: Construct two reinforcement learning network models as described in step 5 as the evaluation network and target network respectively; Set the two lengths to Experience pools are used to store experience data for regular tasks and special tasks respectively. ; in, Indicates step length Action; Indicates step length Action; Indicates step length the status of the public transportation system; Indicates step length the status of the public transportation system; Indicates step length rewards; Step 7.3 Set the maximum number of iterations , minimum number of update iterations , batch sample data volume , learning rate ; Set the initial value, increase, and final value of the special experience sampling ratio for regular tasks 、 、 ; Set the initial value, increase, and final value of the special task special experience sampling ratio 、 、 ; Randomly initialize the evaluation network parameters and copy the evaluation network parameters to the target network; Step 7.4: Sample from the training set, activate the task head network for regular tasks or special tasks according to the sampled data type, and obtain the current task date; randomly obtain a set of preference vectors from the target preference subspace; Step 7.5: Each agent observes the environment and returns the corresponding state and action information. Based on the shared network, the task-specific parallel task head network for optimization goal 1 and optimization goal 2, and the weighted greedy strategy, it obtains the recommended action and calculates the overall constraint violation of the recommended action. Step 7.6: Each agent uses the action correction strategy to correct the recommended actions that do not meet the constraints, obtaining the actual action set. The actual action set is returned to the environment for execution, and the environment feeds back a reward vector and transitions to the next state. Step 7.7: Change the step length Movement, step length Movement, step length The bus system status and step length The status and pace of the public transportation system The rewards are stored as experience in the experience pool of the corresponding regular task or special task; Step 7.8: When the experience in the experience pool exceeds the threshold When , draw from the experience pool The sample experience data is used to update the reinforcement learning network model parameters described in step 5; When the number of iterations is less than When , the sample experience data only comes from the regular task experience pool; When the number of iterations is greater than or equal to When, for routine tasks, the sample experience data will be 、 The proportion is randomly obtained from the regular task experience pool and the special task pool; When the number of iterations is greater than or equal to When targeting special tasks, sample experience data will be 、 The proportion is randomly obtained from the regular task experience pool and the special task pool; Every In the iteration, for conventional tasks, the reinforcement learning network model will Increase the sampling ratio of special experience in regular tasks and special tasks until it reaches the preset maximum value; Every In the iteration, for special tasks, the reinforcement learning network model will Increase the sampling ratio of special experience in regular tasks and special tasks until it reaches the preset maximum value; is a constant; Step 7.9: Based on the sample empirical data, calculate the loss function according to formulas (45)-(47). ; expressed as: (45) (46) (47) Where, It is The loss value of each target under the current batch of sample experience; It is sample experience In the evaluation network The joint action Q value under the target; is to evaluate the network parameters; It is sample experience In the The Q value of the joint action of the target under the target; are the target network parameters; It is sample experience Middle The reward value of each goal; It is sample experience In the target network The joint action Q value under the target; Step 7.10, based on Back-propagation is used to calculate the gradient, and the Adam optimizer is used to update the parameters of the reinforcement learning solution network model described in step 5. Step 7.11: Remove the dominated experience from the experience pool; Step 7.12: Increase the sampling probability of the empirically sparse preference subspace; Step 7.13, repeat steps 7.4-7.12 Second, copy the evaluation network parameters to the target network; is a constant; Step 7.14: Repeat steps 7.4 to 7.13 until the number of iterations reaches the preset maximum number of iterations, obtain the trained reinforcement learning solution network model, and save the parameters.
10. The method for optimizing charging scheduling of a hybrid electric bus system based on reinforcement learning according to claim 9, characterized in that: In step 8, a charging plan for the actual optimized period is obtained based on the trained reinforcement learning network model; The specific process is: Step 8.1: Determine the actual optimization period and user target preference weights; Step 8.2: Obtain the bus and energy storage status information, along with the target preference weights, and input them into the shared network in the trained reinforcement learning network model. The shared network output is input into the task head network, and the actual action set is obtained by combining the weighted greedy strategy and the action correction strategy. The actual action set is converted into a charging plan for execution. Step 8.3: Repeat step 8.2 until the end of the optimization period.
Citation Information
Patent Citations
Shared electric vehicle scheduling method based on deep reinforcement learning
CN114971251A
Electric automobile fleet scheduling method and system based on deep reinforcement learning
CN116822898A
Electric bus dynamic scheduling system and method based on deep reinforcement learning
CN116895144A
Electric vehicle charging planning method and device, computer equipment and storage medium
CN117610763A
New energy automobile charging intelligent scheduling method based on deep reinforcement learning
CN118536726A
Cited By
Ink path optimization control method, system and equipment based on reinforcement learning and medium
CN121697353A
Ink path optimization control method, system, device and medium based on reinforcement learning
CN121697353B