Energy scheduling method and system for railway station
Through deep reinforcement learning methods, the state and action space are constructed and the control action is selected, which solves the problem of unstable scheduling of multi-source power systems in railway stations, and efficient and automated energy scheduling is achieved, reducing costs and carbon emissions.
Patent Information
- Application Number
- CN202510011185.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The prior art cannot efficiently dispatch the multi-source power system of railway stations, resulting in unstable operation.
Deep reinforcement learning (DRL) method is adopted to build state space and action space, and control actions are selected through the main policy network and the main valuation network, and the reward value is calculated based on cost savings, energy utilization efficiency, user satisfaction, equipment wear degree and environmental impact degree to realize dynamic scheduling and management of the system.
It realizes automated and efficient scheduling of multi-source complementary power systems of railway stations, reduces costs, improves user satisfaction, extends equipment service life, and reduces carbon emissions.
Smart Images

Figure CN119941446A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy dispatching, and in particular to an energy dispatching method and system for railway stations. Background Art
[0002] In order to meet the needs of reducing costs and increasing efficiency and reducing carbon emissions, railway stations have transformed from a single energy model that relies entirely on power grid power supply to a diversified energy model that relies on the environment to establish multi-source power supply equipment combined with power supply networks. Low-carbonization in all links faces new requirements and challenges, and a high proportion of renewable energy power generation has become an inevitable trend. Multi-source power supply equipment mainly uses renewable energy for power generation, such as photovoltaic panels, wind turbines, and cogeneration equipment. However, the process of renewable energy power generation is highly dependent on the external natural environment, and there are problems of unstable, unpredictable, and uncontrollable power generation, which seriously affects the stable operation of railway stations. In order to ensure the stable operation of infrastructure such as railway stations, it is necessary to carry out complex scheduling of multi-source power supply equipment, energy storage equipment, load equipment, and external power networks in the power system of railway stations. Therefore, a scheduling solution based on the complementarity of multiple energy sources is urgently needed. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides an energy scheduling method and system for a railway station to eliminate or improve one or more defects existing in the prior art and solve the problem that the prior art cannot efficiently schedule the multi-source power system of the railway station to ensure stable operation.
[0004] One aspect of the present invention provides an energy dispatching method for a railway station, the method being applicable to a multi-source complementary power system, the multi-source complementary power system comprising a plurality of power generation equipment, energy storage equipment and load equipment connected to the grid, the method comprising the following steps:
[0005] Constructing a state space, wherein the state space includes environmental parameters, electricity price state parameters, energy equipment state parameters and user behavior parameters of the target railway station; the environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; the electricity price state parameters include electricity purchase and selling prices; the energy equipment state parameters include current power generation and expected power generation of multiple types of power generation equipment, remaining power and charge and discharge status of energy storage equipment, energy demand of load equipment, equipment health status parameters and carbon emissions; the user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment use preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters;
[0006] Constructing an action space, wherein the action space includes control parameters for multiple types of power generation equipment, energy storage equipment, power trading control parameters, and load equipment control parameters;
[0007] Based on deep reinforcement learning, a main network including a main policy network and a main valuation network is constructed, wherein the main policy network takes the parameters of the state space at the current moment as input and outputs the control action selected in the action space; the reward value is calculated in combination with cost savings, energy utilization efficiency, user satisfaction, equipment wear and tear, and environmental impact, and a value function of the action selected based on the state is established according to the reward value, and the main valuation network fits the value function to obtain a main Q value; an experience replay buffer is constructed to store the data generated by each interaction for random sampling training and updating of the main network; wherein a target network with the same structure as the main policy network and the main valuation network is introduced, and the target Q value obtained by fitting the value function of the target network and the main Q value are used to calculate the mean square error as the loss function, and the parameters of the main network are updated by minimizing the loss function, and the parameters of the main network are periodically copied to the target network;
[0008] The control actions selected by the main network are executed to realize energy dispatch of the target railway station, and the main network is continuously updated and optimized.
[0009] In some embodiments, the external natural environment parameters include meteorological information, geological information, and time information of the area where the target railway station is located; the meteorological information includes light intensity, wind speed, temperature, and humidity; the time information includes date, season, and current time;
[0010] The power generation equipment includes: photovoltaic panels, wind turbines and cogeneration subsystems;
[0011] The equipment usage preference parameters are quantified by marking the load that the user chooses to run, the frequency of use of the load equipment, and the load operation time;
[0012] The user satisfaction feedback parameters include the satisfaction score fed back by the user through a preset link, the quantified adjustment opinion, the use frequency of the load device and the recommendation index of the load device;
[0013] The power generation equipment control parameters include start and stop control parameters and operating power control parameters of the power generation equipment;
[0014] The energy storage device control parameters include scheduling parameters and charge and discharge power control parameters for multiple energy storage devices;
[0015] The power purchase and sale control parameters include sale scheduling parameters for the output power of the power generation equipment and the power stored in the energy storage equipment, as well as external power purchase control parameters;
[0016] The load device control parameters include start / stop control parameters and operation status control parameters for each load device;
[0017] The estimated power generation is predicted by an ANN neural network.
[0018] In some embodiments, the main strategy network and the main valuation network adopt a fully connected network, a convolutional neural network or a recurrent neural network, and introduce a ReLU or sigmoid activation function.
[0019] In some embodiments, the reward value is calculated by combining cost savings, energy efficiency, user satisfaction, equipment wear and tear, and environmental impact, including:
[0020] The purchase cost of the multi-source complementary power system is calculated as follows:
[0021] R cost_saving =-(Price purchased ×Power purchased );
[0022] Among them, R cost_saving Indicates the purchase cost, Price purchased Indicates the purchase price of electricity, Power purchased Indicates the purchased electricity;
[0023] The sales revenue of the multi-source complementary power system is calculated as follows:
[0024] R sales_gain =Price sold ×Power sold ;
[0025] Among them, R sales_gain represents the sales revenue, Price sold Indicates the sales price of electricity, Power sold Indicates the amount of electricity sold;
[0026] The energy utilization rate of the multi-source complementary power system at each moment is obtained, and the change in energy utilization rate at the current moment is calculated, and the calculation formula is:
[0027] R efficiency =Efficiency current -Efficiency previous ;
[0028] Among them, R efficiency Indicates the change in energy utilization rate, Efficiency current Indicates the current energy utilization rate, Efficiency previous Indicates the energy utilization rate at the last moment;
[0029] The health status of the multi-source complementary power system at each moment is obtained to characterize the degree of wear of the equipment, and the change in the health status of the system at the current moment is calculated. The calculation formula is:
[0030] R health =Health current -Health previous ;
[0031] Among them, R health Indicates the change in the health status of the system. current Indicates the health status of the multi-source complementary power system at the current moment, Health previous Indicates the health status of the multi-source complementary power system at the last moment; the health status is calibrated using the remaining service life or failure rate of each device;
[0032] The reward value is calculated as follows:
[0033] R t =ω 1 (R cost_saving +R sales_gain )+ω 2 ·R efficiency +ω 3 ·R satisfaction +ω 4 ·R halth +ω 5 ·R environmental ;
[0034] R environmental =-Emissions current ;
[0035] Among them, R t represents the reward value at time t, R satisfaction is the quantitative satisfaction score, R environmental is the environmental impact value, Emissions current is the carbon emissions of the multi-source complementary power system, ω 1 ,ω 2 ,ω 3 ,ω 4 and ω 5 is the weight coefficient.
[0036] In some embodiments, the main valuation network fits the cost function to obtain the main Q value, which is expressed as:
[0037] Q(s,a,θ)=E[R t +γmax a′ Q(s′,a′,θ)|s,a];
[0038] Where Q(s,a,θ) represents the main Q value, E represents the expectation, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ is the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ is the parameter of the main valuation network;
[0039] The target Q value obtained by fitting the value function to the target network is expressed as:
[0040] y=R t +γmax a′ Q target (s′, a′, θ′);
[0041] Where y represents the target Q value, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ is the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ′ is the parameter of the target network.
[0042] In some embodiments, the loss function is expressed as:
[0043] L(θ)=E[(yQ(s,a,θ)) 2 ]
[0044] Among them, L(θ) represents the loss function, E represents expectation, y represents the target Q value, and Q(s, a, θ) represents the main Q value.
[0045] In some embodiments, the main strategy network and the main valuation network are pre-trained based on a historical state space and a historical action space constructed based on historical data, and are migrated to the multi-source complementary power system for operation and continuous updating.
[0046] On the other hand, the present invention also provides a multi-source complementary power system suitable for a railway station, the system comprising:
[0047] A multi-source power generation subsystem, including a photovoltaic panel, a wind turbine and a cogeneration subsystem, wherein the multi-source power generation subsystem is connected to an external power grid;
[0048] An energy storage device subsystem, connected to the multi-source power generation subsystem and the external power grid;
[0049] A plurality of load devices, connecting the multi-source power generation subsystem, the energy storage device subsystem and the external power grid;
[0050] The equipment management subsystem is used to execute the above-mentioned energy scheduling method for railway stations, select and execute control actions according to environmental parameters, electricity price status parameters, energy equipment status parameters and user behavior parameters to achieve energy scheduling of the target railway station; the environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; the electricity price status parameters include electricity purchase and selling prices; the energy equipment status parameters include the current power generation and expected power generation of various types of power generation equipment, the remaining power and charging and discharging status of energy storage equipment, the energy demand of load equipment, equipment health status parameters and carbon emissions; the user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment use preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters.
[0051] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when the computer program / instruction is executed by a processor.
[0052] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0053] The beneficial effects of the present invention are at least:
[0054] The energy dispatching method and system for railway stations described in the present invention are based on the multi-source complementary power supply system constructed by the railway station, and the environmental parameters, electricity price state parameters, energy equipment state parameters and user behavior parameters are constructed as a state space, and the control parameters of multiple types of power generation equipment, energy storage equipment control parameters, power trading control parameters and load equipment control parameters are constructed as an action space. The control action is selected based on deep reinforcement learning DRL to realize the dynamic dispatching and management scheme of various power supply equipment, energy storage equipment and load equipment in the system. The cost saving amount, energy utilization efficiency, user satisfaction, equipment wear degree and environmental impact degree are introduced to calculate the reward value, so as to ensure that the cost can be reduced and the efficiency can be increased in the control optimization process, improve user satisfaction, reduce equipment loss and carbon emissions, and realize automatic, efficient and stable energy dispatching.
[0055] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.
[0056] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:
[0058] Figure 1 The figure is a flow chart of an energy dispatching method for a railway station according to an embodiment of the present invention.
[0059] Figure 2 This is a graph showing changes in photovoltaic output, wind energy output, charging power and discharge power during the execution of an energy scheduling method for a railway station according to another embodiment of the present invention.
[0060] Figure 3 This is a diagram of battery energy ratio changes during the execution of an energy scheduling method for a railway station according to another embodiment of the present invention.
[0061] Figure 4 This is a diagram of changes in user power demand during the execution of the energy dispatching method for railway stations described in another embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0063] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0064] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0065] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0066] The present invention adopts the application of Deep Reinforcement Learning (DRL) in the Multi-Energy Complementary System (MECS) of the railway station intelligent building energy management system, mainly through the interaction between the learning environment and the agent to optimize the system's operation strategy to achieve higher energy utilization efficiency, lower cost and better reliability.
[0067] The present invention uses DRL in intelligent scheduling to realize the intelligent scheduling of energy storage systems, ensuring the rational allocation and use of energy in different time periods. In a multi-energy complementary system, there are multiple energy sources (such as solar energy, wind energy, natural gas, etc.) and equipment (such as energy storage devices, cogeneration systems, etc.). The dynamic scheduling and optimization of these energy sources and equipment improve the overall performance of the system. To adapt to changes in environmental conditions (such as weather changes) and user needs, manual or static optimization is applied. In DRL, dynamic optimization is achieved by automatically adjusting strategies through learning the laws of environmental changes, and then used to realize intelligent control of buildings or power grids. By training the DRL model, it can automatically adjust the operating status of various equipment in the building according to the real-time energy supply and demand situation, so as to achieve the purpose of both meeting comfort requirements and saving energy.
[0068] One aspect of the present invention provides an energy dispatching method for a railway station, the method being applicable to a multi-source complementary power system, the multi-source complementary power system comprising a plurality of power generation equipment, energy storage equipment and load equipment interconnected to the grid, such as Figure 1 As shown, the method includes the following steps S101 to S104:
[0069] Step S101: construct a state space, which includes environmental parameters, electricity price state parameters, energy equipment state parameters and user behavior parameters of the target railway station; environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; electricity price state parameters include electricity purchase and selling prices; energy equipment state parameters include current power generation and expected power generation of various types of power generation equipment, remaining power and charge and discharge status of energy storage equipment, energy demand of load equipment, equipment health status parameters and carbon emissions; user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment use preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters.
[0070] Step S102: constructing an action space, where the action space includes control parameters for multiple types of power generation equipment, energy storage equipment, power trading control parameters, and load equipment control parameters.
[0071] Step S103: Based on deep reinforcement learning, a main network including a main policy network and a main valuation network is constructed, wherein the main policy network takes the parameters of the state space at the current moment as input and outputs the control action selected in the action space; the reward value is calculated in combination with the cost savings, energy utilization efficiency, user satisfaction, equipment wear and environmental impact, and a value function based on state-selected actions is established according to the reward value, and the main valuation network fits the value function to obtain a main Q value; an experience replay buffer is constructed to store the data generated by each interaction for random sampling training and updating of the main network; wherein a target network with the same structure as the main policy network and the main valuation network is introduced, and the target Q value and the main Q value obtained by fitting the value function of the target network are used to calculate the mean square error as the loss function, and the loss function is minimized to update the parameters of the main network, and the parameters of the main network are periodically copied to the target network.
[0072] Step S104: Execute the control action selected by the main network to realize the energy dispatch of the target railway station, and continuously update and optimize the main network.
[0073] The railway station targeted by the method described in steps S101 to S104 is a multi-source complementary power system that introduces a variety of energy sources for power supply, relying on light energy, wind energy, geothermal energy and other energy sources in the natural resource environment to generate electricity and assist in energy supply. However, there is a problem of unstable energy supply for electricity generated by natural resources. In order to achieve stable power supply, the power generation equipment of the railway station must be incorporated into the national power grid. At the same time, energy storage equipment is deployed in the station to work together in power generation, power storage, power dispatching and load control to achieve efficient, stable, environmentally friendly and low-carbon control goals. Relying on deep reinforcement learning, a mapping relationship between environmental state and system-wide control parameters is established. Environmental modeling involves defining the state space, action space and reward function of the environment, which are the basis for the DRL algorithm to learn and optimize. Configuring a suitable environmental model based on the application scenario can capture the dynamic characteristics of environmental changes and support agents to learn effective strategies by interacting with it.
[0074] The state space S is a set of variables used to describe the current state of the environment. For example, the current time, weather conditions, the status of various energy devices (such as the power generation of photovoltaic panels, the remaining power of energy storage systems, etc.), electricity prices, etc. The action space A is used for all possible actions that the agent can take. For example, the charging and discharging operations of energy storage systems, the startup / shutdown of renewable energy equipment, load management, etc. The reward function R is used to define the reward or penalty after the agent takes a certain action. The design of the reward function should be set according to the specific optimization goals, such as energy saving, cost saving, user satisfaction, etc.
[0075] In step S101, based on the multi-source complementary power system of the railway station, it is necessary to construct a state space for multiple types of data at the same time, including the external environment that affects the power output of power generation equipment and the energy consumption of loads, such as light that affects the production efficiency of photovoltaic power stations, and temperature determines the power consumption of air conditioners in railway stations. Due to its instability, power generation equipment requires complex energy scheduling, such as storing energy during periods of high power generation efficiency, selling electricity to the power grid, and scheduling reserved electricity or buying electricity from the power grid during periods of low power generation efficiency. In the power scheduling link, the electricity price state parameters affect the decision-making of power sales and purchases. The energy equipment state parameters reflect the operating status of power generation equipment, energy storage equipment, and load equipment, which is one of the bases for scheduling decisions. User behavior parameters reflect the specific experience and needs of users during use, which is also one of the bases for scheduling decisions.
[0076] In some embodiments, the external natural environment parameters include weather information, geological information, and time information of the area where the target railway station is located; weather information includes light intensity, wind speed, temperature, and humidity; time information includes date, season, and current time. Various parameter information is marked according to preset dimensional units, for example, the light intensity unit is watts per square meter (W / m 2 ), wind speed is in meters per second (m / s), and temperature is in degrees Celsius (℃).
[0077] Power generation equipment includes: photovoltaic panels, wind turbines and cogeneration subsystems.
[0078] Energy equipment status includes the current or expected power generation of photovoltaic panels and wind turbines, the current operating status (on / off), power generation, thermal energy output of the cogeneration subsystem, etc. The current remaining power, charge and discharge status of the energy storage device, etc. The current energy demand of the load equipment, such as the energy consumption of air conditioning, lighting and other equipment.
[0079] User behavior includes two parts: user demand and user feedback. User demand usually includes the user's basic requirements for energy use, such as power demand, comfort requirements, etc. User feedback is the user's response to the actual use of the product or service, which can be positive (such as satisfaction, appreciation) or negative (such as dissatisfaction, complaints). These parameters are quantified by referring to.
[0080] Quantitative methods include, for example, 1) Satisfaction score: the user's satisfaction score for the product or service. 2) Specific opinion: the user's detailed opinion or suggestion on the product or service, selecting from multiple preset items. 3) Frequency of use: the frequency with which the user uses the product or service. 4) Willingness to recommend: the degree to which the user is willing to recommend the product or service to others.
[0081] The equipment usage preference parameters are quantified by marking the loads that the user chooses to run, the frequency of load equipment use, and the load operation duration.
[0082] The user satisfaction feedback parameters include the satisfaction score fed back by the user through the preset link, the quantitative adjustment opinions, the frequency of use of the load device and the recommendation index of the load device.
[0083] The estimated power generation is predicted by the ANN neural network, which specifically includes the following steps S201 to S206:
[0084] Step S201: Data preparation and preprocessing, first collect historical power generation equipment operation data, including input (such as weather conditions, temperature, humidity and other environmental factors) and output (i.e. power generation). Clean the data, process missing values and outliers, and perform normalization or standardization to meet the training requirements of the neural network.
[0085] Step S202: Model configuration, select a neural network structure suitable for time series prediction, such as a multi-layer perceptron (MLP). Configure the number of network layers and the number of neurons in each layer. For example, you can set an input layer corresponding to the number of features, several hidden layers for extracting features, and an output layer corresponding to the predicted power generation. Use ReLU or sigmoid as the activation function, and select a suitable optimizer, such as Adam or SGD.
[0086] Step S203: Train the model and divide the data set into a training set and a test set. Use the training set data to train the neural network through the back propagation algorithm, and adjust the weights and biases to minimize the prediction error. During the training process, cross-validation can be used to avoid overfitting, and the performance on the validation set can be monitored by early stopping. Once the performance no longer improves, stop training.
[0087] Step S204: Model evaluation and tuning, evaluate the performance of the model on the test set, and use indicators such as mean square error (MSE) or mean absolute percentage error (MAPE) to measure the prediction accuracy. Adjust the network structure, learning rate and other hyperparameters based on the evaluation results, and repeat the training and evaluation process until satisfactory prediction performance is achieved.
[0088] Step S205: Deployment and use, deploy the trained model to the actual power generation prediction system. Collect new input data in real time, and obtain the predicted power generation through network forward propagation. Make resource planning and scheduling decisions based on the prediction results.
[0089] Step S206: Continuously monitor and update, and regularly retrain the model using newly collected data to adapt to environmental changes and new data patterns to ensure that the prediction accuracy remains stable over time.
[0090] In step S102, the action space is the control actions and parameter settings for all controllable power generation equipment, energy storage equipment, and load equipment in the multi-source complementary power system. The action space defines all possible operations that the agent can perform, which are directly related to energy production, storage, consumption, and interaction with the external energy market. In order to achieve effective energy management and optimization, each dimension in the action space is processed in depth, and how they are combined together to achieve the optimal operation of the system as a whole is considered.
[0091] The control parameters of the power generation equipment include the start and stop control parameters and the operating power control parameters of the power generation equipment. For renewable energy equipment, such as photovoltaic power generation (PV) and wind power generation (Wind), their output is greatly affected by natural conditions. Therefore, the control strategy of the equipment should be able to be adjusted according to weather forecasts, real-time weather conditions and other system status information to achieve maximum energy utilization efficiency. For example, adjust the tilt angle of the photovoltaic panel to 30 degrees and adjust the angle of the wind turbine blades to 45 degrees.
[0092] Energy storage device control parameters include scheduling parameters for multiple energy storage devices and charging and discharging power control parameters. The control of the energy storage system (BESS) is crucial to smoothing the imbalance between supply and demand. The control of charging and discharging power not only affects the immediate performance of the system, but also determines the life and cost-effectiveness of the battery. Therefore, the intelligent agent must consider factors such as the state of the battery (such as SOC-State of Charge), the current electricity price, and the predicted renewable energy output to make decisions. For example, instruct the energy storage system to charge at a power of 2.5kW and instruct the energy storage system to discharge at a power of 1.5kW.
[0093] The power trading control parameters include the sale scheduling parameters for the output power of the power generation equipment and the power stored in the energy storage equipment, as well as the external power procurement control parameters. External energy procurement involves the purchase and sale of power with the power grid. The power trading control can coordinate the interaction between the produced power and the power grid to achieve cost reduction and efficiency improvement. When the local energy is insufficient, the power grid is purchased; conversely, when there is a surplus, the excess power can be sold to the grid. This strategy helps to reduce the overall operating cost and helps to balance the load of the power grid. Alternatively, in order to increase benefits and reduce costs, electricity can be sold during high-price periods, and electricity can be purchased and stored during low-price periods. For example, 2.0 kWh of electricity is purchased from the power grid at 9 a.m. to 11 a.m., and 1.5 kWh of electricity is sold to the power grid at 1 p.m. to 2 p.m.
[0094] The load equipment control parameters include the start and stop control parameters and operating status control parameters of each load equipment. Load management involves the control of various electrical equipment in the building, with the aim of ensuring the efficient use of energy while meeting the user's comfort requirements. This can be achieved by intelligently scheduling the equipment operation time, for example, running high-energy consumption equipment when electricity prices are low, or reducing the use of grid power when solar energy production is high. More specific operations include adjusting the air conditioning set temperature to 22 degrees Celsius and setting the brightness of the lighting equipment to 50%.
[0095] Furthermore, in the case of a combined heat and power system, the combined heat and power system (CHP) can provide both electricity and heat, which is very important for improving energy efficiency. The control strategy should take into account the balance of electricity and heat demand, as well as any possible energy storage solutions. For example, the power output of the CHP system is adjusted to 5.0kW and the heat supply is adjusted to 3.0kW.
[0096] In step S103, after the state space and action space are constructed in steps S101 and S102, a mapping between the state and the action is established based on deep reinforcement learning.
[0097] Deep reinforcement learning (DRL) is an artificial intelligence technology that combines deep learning and reinforcement learning. It solves decision-making problems in complex environments through end-to-end perception and control systems. In DRL, the agent interacts with the environment and maximizes the cumulative reward through trial and error learning.
[0098] During the perception process, at each moment, the agent obtains high-dimensional observation data from the environment and uses deep learning methods (such as convolutional neural networks (CNN), recurrent neural networks (RNN), etc.) to process these data and extract useful state feature representations.
[0099] In the decision-making process, based on the extracted state features, the intelligent body will evaluate the value function of each possible action and select the current optimal action to execute through a certain strategy (such as ε-greedy strategy, Softmax strategy, etc.).
[0100] In the execution and feedback process, the environment responds to the agent's actions, generates new states and reward signals, and feeds this information back to the agent. The agent adjusts its strategy based on the feedback to better adapt to the environment.
[0101] The execution process includes: initialization, setting the initial state of the agent, the policy network and the value network (if a value function-based method is used), and related hyperparameters. In the loop iteration, at each time step, the agent performs the following steps: perceive the state of the environment, select actions using the policy network, execute actions and observe the feedback from the environment (including new states and rewards), store the experience (state, action, reward, new state) in the experience replay pool, and sample a batch of experience from the experience replay pool to update the policy network and value network (if a value function-based method is used). Repeat the above steps until the termination condition is reached (such as reaching the maximum time step, the cumulative reward reaches the threshold, etc.).
[0102] During the training process, the agent continuously optimizes its policy network to improve its ability to obtain higher rewards in future environments. This is usually achieved through optimization algorithms such as gradient descent, with the goal of minimizing the difference between the action output by the policy network and the optimal action (if a policy gradient-based method is used), or maximizing the cumulative reward (if a value function-based method is used). The present invention adopts the former.
[0103] Specifically, the main strategy network and the main valuation network adopt a fully connected network, a convolutional neural network or a recurrent neural network, and introduce a ReLU or sigmoid activation function.
[0104] The reward function R defines the immediate feedback that the agent receives after taking an action. The design of the reward function should be able to guide the agent to learn a strategy that meets the expected goal. In a multi-energy complementary system, the reward function may include: cost savings, calculating rewards based on energy procurement costs and sales revenue. Energy efficiency, giving rewards or penalties based on energy utilization. User satisfaction, adjusting rewards based on user feedback. Equipment health status, giving rewards or penalties based on the health status of the equipment to encourage preventive maintenance. Environmental impact, calculating rewards based on carbon emissions to encourage low-carbon operation.
[0105] In some embodiments, the reward value is calculated by combining cost savings, energy efficiency, user satisfaction, equipment wear and tear, and environmental impact, including steps S301 to S305:
[0106] Step S301: Calculate the procurement cost of the multi-source complementary power system, the calculation formula is:
[0107] R cost_saving =-(Price purchased ×Power purchased );
[0108] Among them, R cost_saving Indicates the purchase cost, Price purchased Indicates the purchase price of electricity, Power purchasedIndicates the purchased electricity.
[0109] Step S302: Calculate the sales revenue of the multi-source complementary power system, the calculation formula is:
[0110] R sales_gain =Price sold ×Power sold ;
[0111] Among them, R sales_gain Indicates sales revenue, Price sold Indicates the sales price of electricity, Power sold Indicates the amount of electricity sold;
[0112] Step S303: Obtain the energy utilization rate of the multi-source complementary power system at each moment, and calculate the change in energy utilization rate at the current moment. The calculation formula is:
[0113] R efficiency =Efficiency current -Efficiency previous ;
[0114] Among them, R efficiency Indicates the change in energy utilization rate, Efficiency current Indicates the current energy utilization rate, Efficiency previous Indicates the energy utilization rate at the last moment;
[0115] Step S304: Obtain the health status of the multi-source complementary power system at each moment to characterize the degree of equipment wear, and calculate the change in the health status of the system at the current moment. The calculation formula is:
[0116] R health =Health current -Health previous ;
[0117] Among them, R health Indicates the change in the health status of the system. current Indicates the health status of the multi-source complementary power system at the current moment, Health previous Indicates the health status of the multi-source complementary power system at the previous moment; the health status is calibrated using the remaining service life or failure rate of each device.
[0118] Step S305: Calculate the reward value, the calculation formula is:
[0119] R t =ω 1 (R cost_saving +R sales_gain )+ω 2 ·Refficiency +ω 3 ·R satisfaction +ω 4 ·R halth +ω 5 ·R environmental ;
[0120] R environmental =-Emissions current ;
[0121] Among them, R t represents the reward value at time t, R satisfaction is the quantitative satisfaction score, R environmental Emissions is the environmental impact value. current is the carbon emissions of the multi-source complementary power system, ω 1 ,ω 2 ,ω 3 ,ω 4 and ω 5 is the weight coefficient.
[0122] The purpose of steps S301 and S302 is to control cost savings, which means reducing the cost of energy procurement by optimizing energy management and trading strategies, and possibly increasing revenue by selling excess energy. Based on the two-part calculation, reducing procurement costs, each time the agent reduces the cost of energy purchased from the grid by adjusting the energy storage system, renewable energy equipment, or external energy procurement strategy, a positive reward can be obtained. Increasing sales revenue, if the agent can sell excess energy to the grid or other consumers at a suitable price, a positive reward can be obtained.
[0123] Step S303 is to control the energy utilization rate. The energy utilization efficiency reflects the degree of energy loss during production, storage, transmission and use. High efficiency means less energy waste. The reward function is designed based on energy utilization to encourage the agent to improve the overall efficiency of energy use.
[0124] Step S304 is to control the health of the equipment, which directly affects the reliability and operating cost of the system. The reward function can be designed based on the health of the equipment to encourage the agent to take actions to extend the life of the equipment and prevent failures. Whenever the agent takes actions to improve the health of the equipment (such as regular maintenance, replacement of worn parts, etc.), a positive reward is given.
[0125] In step S305, user satisfaction is further introduced. User satisfaction is a subjective indicator that depends on the user's perception of the quality of the service. In a multi-energy complementary system, this may involve the user's satisfaction with the availability and quality of energy. The reward function can be adjusted based on user feedback to encourage the agent to provide a better service experience. The reward is adjusted based on the user's feedback score, with positive feedback corresponding to positive rewards and negative feedback corresponding to negative rewards.
[0126] In some embodiments, the main valuation network fits the cost function to obtain the main Q value, which is expressed as:
[0127] Q(s,a,θ)=E[R t +γmax a′ Q(s′,a′,θ)|s,a];
[0128] Among them, Q(s,a,θ) represents the main Q value, E represents the expectation, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ represents the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ is the parameter of the main valuation network;
[0129] The target Q value obtained by fitting the value function of the target network is expressed as:
[0130] y=R t +γmax a′ Q target (s′, a′, θ′);
[0131] Among them, y represents the target Q value, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ is the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ′ is the parameter of the target network.
[0132] In some embodiments, the optimization goal is to minimize the difference between the action output by the main policy network and the optimal action, and the loss function is expressed as:
[0133] L(θ)=E[(yQ(s,a,θ)) 2 ];
[0134] Among them, L(θ) represents the loss function, E represents the expectation, y represents the target Q value, and Q(s,a,θ) represents the main Q value.
[0135] In some embodiments, the main strategy network and the main valuation network are pre-trained based on the historical state space and historical action space constructed based on historical data, and are migrated to the multi-source complementary power system for operation and continuous updating.
[0136] On the other hand, the present invention also provides a multi-source complementary power system suitable for a railway station, the system comprising:
[0137] The multi-source power generation subsystem includes a photovoltaic panel, a wind turbine generator and a combined heat and power subsystem, and the multi-source power generation subsystem is connected to an external power grid.
[0138] The energy storage device subsystem is connected to the multi-source power generation subsystem and the external power grid.
[0139] Multiple load devices connect multiple source generation subsystems, energy storage device subsystems and external power grids.
[0140] The equipment management subsystem is used to execute the energy dispatching method for railway stations described in the above steps S101 to S104, and select and execute control actions according to environmental parameters, electricity price status parameters, energy equipment status parameters and user behavior parameters to achieve energy dispatching of the target railway station; environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; electricity price status parameters include electricity purchase and selling prices; energy equipment status parameters include current power generation and expected power generation of various types of power generation equipment, remaining power and charge and discharge status of energy storage equipment, energy demand of load equipment, equipment health status parameters and carbon emissions; user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment use preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters.
[0141] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when the computer program / instruction is executed by a processor.
[0142] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0143] The present invention is described below in conjunction with a specific embodiment:
[0144] This embodiment provides an energy dispatching method for a railway station, which is implemented for a railway station. The railway station has the following facilities: photovoltaic panels with a total installed capacity of 20kW. Wind turbines with a total installed capacity of 10kW. Energy storage system (BESS) with a total capacity of 50kWh. And functional load equipment including lighting, air conditioning, monitoring, gates, etc.
[0145] By combining the state space, action space, reward function and DQN algorithm formula, a complete deep reinforcement learning framework is constructed to optimize the operation strategy of the multi-energy complementary system. The framework can automatically learn and adjust the strategy to maximize energy efficiency and minimize costs. As the training progresses, the agent will gradually learn to take the best actions in different states, thereby improving the performance of the entire system.
[0146] The following is a detailed description of the specific process:
[0147] 1. Data Collection
[0148] Collect various data required for the operation of the multi-energy complementary system to provide a basis for environmental construction. The data types include the following:
[0149] Historical electricity consumption data, including past electricity consumption, peak electricity consumption, electricity consumption patterns and other information.
[0150] Weather forecast data, including meteorological parameters such as light intensity, wind speed, temperature, etc., are crucial for predicting the output of renewable energy.
[0151] Electricity price data, real-time electricity price information, is very important for optimizing energy procurement strategies.
[0152] Equipment status data, remaining power of energy storage system, output power of photovoltaic panels, output power of wind turbines, etc.
[0153] User behavior data, including information on user electricity usage habits, preferences, etc., helps optimize the user experience.
[0154] Data sources include sensor data (such as photovoltaic panels, energy storage systems, etc.), weather station data, real-time electricity price information provided by power companies, and user feedback data.
[0155] 2. Constructing state space, action space and reward function
[0156] Based on the collected data, a simulation environment is built to train and test the DRL model. The construction steps include:
[0157] Define the state space S, the state space of the environment, including weather conditions, time information, electricity prices, equipment status, etc. The state space includes the environmental parameters of the target railway station, electricity price status parameters, energy equipment status parameters and user behavior parameters; environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; electricity price status parameters include electricity purchase and selling prices; energy equipment status parameters include the current power generation and expected power generation of various types of power generation equipment, the remaining power and charge and discharge status of energy storage equipment, the energy demand of load equipment, equipment health status parameters and carbon emissions; user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment usage preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters. Define the state space S = {s t},s t The state vector represented here at time t contains information in multiple dimensions, such as weather conditions, current time, real-time electricity prices, output power of photovoltaic panels, output power of wind turbines, remaining power of energy storage systems, etc.
[0158] The action space A defines all possible actions that the agent can take, such as charging and discharging the energy storage system, adjusting the angle of the photovoltaic panel, etc. t}, where a t Represents the action vector taken by the agent at time t, which may include charging and discharging power of the energy storage system, adjusting loads (such as air conditioning temperature, lighting brightness, etc.), etc.
[0159] Reward function R, defines the reward function, which is the immediate feedback obtained by the agent after taking a certain action, and is used to guide the agent to learn the optimal strategy.
[0160] In multi-energy complementary systems, it is very common to define a reward function to evaluate the actions taken by an agent (such as a control system) in a specific state. This reward function usually aims to maximize the overall benefit of the system while minimizing adverse effects.
[0161] Calculate the procurement cost of the multi-source complementary power system, the calculation formula is:
[0162] R cost_saving =-(Price purchased ×Power purchased );
[0163] Among them, R cost_saving Indicates the purchase cost, Price purchased Indicates the purchase price of electricity, Power purchased Indicates the purchased electricity.
[0164] Calculate the sales revenue of the multi-source complementary power system as follows:
[0165] Rsales_gain =Price sold ×Power sold ;
[0166] Among them, R sales_gain Indicates sales revenue, Price sold Indicates the sales price of electricity, Power sold Indicates the amount of electricity sold;
[0167] Obtain the energy utilization rate of the multi-source complementary power system at each moment, and calculate the change in energy utilization rate at the current moment. The calculation formula is:
[0168] R efficiency =Efficiency current -Efficiency previous ;
[0169] Among them, R efficiency Indicates the change in energy utilization rate, Efficiency current Indicates the current energy utilization rate, Efficiency previous Indicates the energy utilization rate at the last moment;
[0170] The health status of the multi-source complementary power system at each moment is obtained to characterize the degree of equipment wear, and the change in the system health status at the current moment is calculated. The calculation formula is:
[0171] R health =Health current -Health previous ;
[0172] Among them, R health Indicates the change in the health status of the system. current Indicates the health status of the multi-source complementary power system at the current moment, Health previous Indicates the health status of the multi-source complementary power system at the previous moment; the health status is calibrated using the remaining service life or failure rate of each device.
[0173] Calculate the reward value, the calculation formula is:
[0174] R t =ω 1 (R cost_saving +R sales_gain )+ω 2 ·R efficiency +ω 3 ·R satisfaction +ω 4 ·R halth +ω 5 ·R environmental ;
[0175] R environmental =-Emissions current ;
[0176] Among them, R t represents the reward value at time t, R satisfaction is the quantitative satisfaction score, R environmental Emissions is the environmental impact value. current is the carbon emissions of the multi-source complementary power system, ω 1 ,ω 2 ,ω 3 ,ω 4 and ω 5 is the weight coefficient.
[0177] 3. Network Design
[0178] A deep neural network is designed for learning to approximate the Q-function or policy function.
[0179] Network architecture: Choose a suitable network architecture, such as fully connected network, convolutional neural network (CNN), recurrent neural network (RNN), etc.
[0180] Input layer: The input layer is defined according to the state space to receive the environment state information.
[0181] Output layer: The output layer is defined based on the action space and outputs the actions that the agent can take.
[0182] Hidden layers: Design an appropriate number of hidden layers and nodes to capture the complex relationship between states and actions.
[0183] Activation function: Select a suitable activation function, such as ReLU, sigmoid, etc.
[0184] Specifically, based on DQN, a deep neural network is used to approximate the Q function, and two techniques, Experience Replay Buffer and Target Network, are introduced to improve the stability and efficiency of learning.
[0185] Building the Experience Replay Buffer
[0186] After each interaction, (s t ,a t ,R t ,s t+1 ) is stored in the experience replay buffer D.
[0187] D←(s t ,a t ,R t ,st+1 );
[0188] Then a batch of samples (s, a, R, s′) are randomly selected to update the network.
[0189] The main valuation network fits the value function to obtain the main Q value, which is expressed as:
[0190] Q(s,a,θ)=E[R t +γmax a′ Q(s′,a′,θ)|s,a];
[0191] Among them, Q(s,a,θ) represents the main Q value, E represents the expectation, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ represents the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ is the parameter of the main valuation network;
[0192] Construct the target network Q target The parameter θ target Periodically receive data from the main network Q main The parameter θ main Copied.
[0193] θ target ←θ main ;
[0194] The target network is used to calculate the target Q value:
[0195] y=R t +γmax a′ Q target (s′, a′, θ′);
[0196] Among them, y represents the target Q value, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ is the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ′ is the parameter of the target network.
[0197] Network update, using Mean Squared Error (MSE) as the loss function:
[0198] L(θ)=E[(yQ(s,a,θ)) 2 ];
[0199] Among them, L(θ) represents the loss function, E represents the expectation, y represents the target Q value, and Q(s,a,θ) represents the main Q value.
[0200] Minimize the loss function using the gradient descent method:
[0201]
[0202] Here, α is the learning rate.
[0203] 4. Training process
[0204] By interacting with the environment, the DRL model is trained to learn the optimal strategy. The training steps are as follows:
[0205] Initialization: Set the initial state S 0 , initialize the deep neural network Q main and Q target Parameters.
[0206] Interaction: At each time step t, the agent t and the current policy π(a|s t )Select action a t .
[0207] Execute: Execute action a t , observe the new state s t+1 and reward R t .
[0208] Storage experience: t ,a t ,R t ,s t+1 ) is stored in the experience replay buffer D.
[0209] Random sampling: Randomly draw a batch of samples (s, a, R, s′) from the experience replay buffer D.
[0210] Calculate target Q value: Use the target network to calculate the target Q value y.
[0211] Update network: Update the parameters of the main network by gradient descent.
[0212] Copy parameters: Periodically copy the parameters of the master network to the target network.
[0213] 5. Evaluate and adjust
[0214] Evaluate the performance of the agent and adjust the strategy or network structure based on the evaluation results.
[0215] Set evaluation indicators based on specific optimization goals, such as cost savings, energy efficiency, user satisfaction, etc. Regularly test the performance of the agent in the simulation environment and record key indicators. Adjust the reward function, network parameters, etc. based on the evaluation results to optimize the performance of the agent.
[0216] 6. Deployment and Monitoring
[0217] Deploy the trained model to the actual system, make online decisions, and continuously monitor its performance.
[0218] Deployment steps: Deploy the trained model to the actual operating environment to ensure that the model can interact with the actual system. According to the current environmental status, the agent selects the optimal action in real time to control the operation of the multi-energy complementary system. Continuously monitor the actual performance of the agent and record key indicators such as cost and efficiency. According to the actual operating data, adjust the model parameters regularly to cope with environmental changes.
[0219] Simulation experiment data:
[0220] Based on the previously constructed data set, evaluate the performance of the system after adopting the deep reinforcement learning (DRL) optimization strategy. Construct the state space, weather conditions (clear, cloudy, overcast, rainy), light intensity (watts / square meter), wind speed (meters / second), temperature (degrees Celsius), current time (hours), real-time electricity price (yuan / kWh), photovoltaic panel output power (kilowatts), wind turbine output power (kilowatts), energy storage system remaining power (kilowatt-hours), user power demand (kilowatts), equipment health index (between 0 and 1). Construct the action space, energy storage system charging power (kilowatts), energy storage system discharging power (kilowatts), adjust the load (such as air conditioning temperature, lighting brightness, etc.), and maintain the equipment (whether to maintain).
[0221] like Figure 2 , 3 As shown in Figure 4, during the daytime (08:00 to 16:00), the photovoltaic output power gradually increases, reaches the maximum value, and then gradually decreases. At night (00:00 to 07:00), the photovoltaic output is 0. The wind power output is high at night and in the morning (00:00 to 07:00), reaching 2kW. During the day (08:00 to 16:00), the wind power output is low and remains at 1kW. Charging is carried out during the period of low electricity prices (00:00 to 07:00), and the charging power is 2kW. Discharging is carried out during the period of high electricity prices (17:00 to 23:00), and the discharge power is adjusted according to user demand. The battery is charged at night, and the power gradually increases. The power gradually decreases during the day, but it always remains at a high level. User demand remains constant throughout the day at 2kW.
[0222] Through the optimization strategy, the system is able to store electricity when the electricity price is low and use the stored electricity when the electricity price is high, thereby reducing the overall cost. Before optimization, the system directly purchased electricity when the electricity price was high without considering the use of energy storage systems. After optimization, the system will charge when the electricity price is low and discharge when the electricity price is high, while adjusting the load to reduce the overall cost.
[0223] Through optimization, the system reduces power waste and improves energy efficiency. Before optimization, the system directly uses all the power generated, and the excess is wasted. After optimization, the system will try to store excess power and use it when needed.
[0224] Define the indicators of user satisfaction: Power supply stability (Supply Stability): Whether the power supply is stable and uninterrupted. Cost of electricity (Cost Efficiency): Whether the electricity cost is effectively controlled. Indoor comfort (Comfort Level): Whether the indoor temperature is moderate and the lighting is appropriate. System response speed (Response Speed): Whether the system responds to changes in user needs in a timely manner.
[0225] Each indicator is assigned a weight and given a scoring range (e.g., 0 to 1). The scores of each indicator are calculated based on the situation, and then the overall user satisfaction is calculated.
[0226] Before optimization, power supply stability (Supply Stability): The table shows that the power supply is stable throughout the day, so the score is 0.95. Electricity cost (Cost Efficiency): According to the electricity price and the charging and discharging of the energy storage system, users use energy storage electricity instead of purchasing it directly from the grid when the electricity price is high, which can be considered to have good cost control. The electricity cost is properly controlled throughout the day, with a score of 0.85. Indoor comfort (Comfort Level): According to the adjustment of the load (such as adjusting the air conditioning temperature), the indoor temperature is moderate throughout the day, with a score of 0.95. The table shows that the air conditioning was adjusted at night, so the score is 0.95. System response speed (Response Speed): The system can respond to changes in the user's electricity demand in a timely manner, with a score of 0.9. The system responds promptly, with a score of 0.9. Perform the following calculations:
[0227] User opinion 前 =(Supply Stability×ω 1 )+(Cost Efficiency×ω 2 )+(ComfortLevel×ω 3 )+(Response Speed×ω 4 )
[0228] Among them, ω1, ω2, ω3, and ω4 are the weights of each indicator, assuming that they are all 0.25 (i.e., equal weights), so:
[0229] Customer satisfaction 前=(0.95×0.25)+(0.85×0.25)+(0.95×0.25)+(0.9×0.25)=0.9125
[0230] After optimization, the response speed of the system has been improved, the electricity cost has been better controlled, and other indicators remain unchanged. Supply Stability: Remains unchanged before and after optimization, with a score of 0.95. Cost Efficiency: Cost is better controlled after optimization, with a score of 0.90. Comfort Level: Remains unchanged before and after optimization, with a score of 0.95. Response Speed: The response speed is faster after optimization, with a score of 0.95. Perform the following calculations:
[0231] Customer satisfaction 后 =(Supply Stability×ω 1 )+(Cost Efficiency×ω 2 )+(ComfortLevel×ω 3 )+(Response Speed×ω 4 )
[0232] Among them, ω1, ω2, ω3, and ω4 are the weights of each indicator, assuming that they are all 0.25 (i.e., equal weights), so:
[0233] User satisfaction 后 =(0.95×0.25)+(0.90×0.25)+(0.95×0.25)+(0.95×0.25)=0.9375
[0234] The user satisfaction before optimization was 0.9125, and the user satisfaction after optimization was 0.9375. After optimization, the overall user satisfaction with the system has improved, especially in terms of electricity cost control and system response speed.
[0235] The wear of the equipment is calculated, and the defined indicators include: Charge / Discharge Cycles: The number of charge / discharge cycles of the energy storage system throughout the day. Operating Hours: The operating hours of the energy storage system throughout the day. Maintenance Frequency: The number of maintenance times of the energy storage system throughout the day. Load Adjustment Times: The number of times the load is adjusted throughout the day. A weight is set for each indicator, and a score range (for example, 0 to 1) is given. Then the scores of each indicator are calculated based on the actual situation, and finally the total equipment wear is calculated comprehensively.
[0236] Before optimization, Charge / Discharge Cycles: According to the table data, there are 10 charge / discharge operations throughout the day (from t=0 to t=16, each charge / discharge is counted as one operation). The number of charge / discharge operations throughout the day is small, with a score of 0.3 (0 means no charge / discharge, 1 means frequent charge / discharge). Operating Hours: The operating hours throughout the day are long, with a score of 0.7 (0 means no operation, 1 means operation throughout the day). According to the table data, the energy storage system performs charging operations during the day and discharging operations at night, so it works most of the day. Maintenance Frequency: No maintenance is required throughout the day, with a score of 0 (0 means no maintenance, 1 means frequent maintenance). According to the table data, no maintenance operations were performed throughout the day. Load Adjustment Times: According to the table data, the load was adjusted 5 times throughout the day (from t=17 to t=23, each adjustment is counted as one operation). The load is adjusted less frequently throughout the day, with a score of 0.3 (0 means no adjustment, 1 means frequent adjustment). Perform the following calculation:
[0237] Equipment wear 前 =(Charge / Di charge Cycles×ω 1 )+(Operating Hours×ω 2 )+(Maintenance Frequency×ω 3 )+(Load Adjustment Times×ω 4 )
[0238] Among them, ω 1 ,ω 2 ,ω 3 ,ω 4 are the weights of each indicator, and they are assumed to be 0.25 (i.e. equal weight).
[0239] Equipment wear 前 =(0.3×0.25)+(0.7×0.25)+(0×0.25)+(0.3×0.25)=0.325
[0240] After optimization, the system's charge / discharge times are reduced, the load adjustment times are also reduced, and other indicators remain unchanged. Charge / Discharge Cycles: After optimization, the charge / discharge times are reduced, and the score is 0.2. Operating Hours: Remains unchanged before and after optimization, and the score is 0.7. Maintenance Frequency: Remains unchanged before and after optimization, and the score is 0. Load Adjustment Times: Assuming that the load adjustment times are reduced after optimization, the score is 0.2. Perform the following calculations:
[0241] Equipment wear 后 =(Charge / Di charge Cycles×ω 1 ')+(Operating Hours×ω 2 ')+(Maintenance Frequency×ω 3 ')+(Load Adjustment Times×ω 4 ')
[0242] Among them, ω 1 ,ω 2 ,ω 3 ,ω 4 are the weights of each indicator, and they are assumed to be 0.25 (i.e. equal weight).
[0243] Equipment wear 后 =(0.2×0.25)+(0.7×0.25)+(0×0.25)+(0.2×0.25)=0.275
[0244] The wear degree of the equipment before optimization was 0.325, and the wear degree of the equipment after optimization was 0.275. After optimization, the wear degree of the equipment was reduced, especially in terms of the number of charge and discharge times and the number of load adjustments. Such results show that the optimization measures effectively reduce the wear of the equipment and extend the service life of the equipment.
[0245] Charging period: From 0:00 am to 16:00 pm, as wind power generation and solar energy storage are being charged accordingly, with higher light intensity, PV output power increases, the energy storage system is charged during this period, while ensuring that the load uses excess electricity, until charging is basically completed at 4:00 pm.
[0246] Discharge period: from 5 pm to 0 am the next day. During this period, the wind is weak and the light is weakened, the PV output power is reduced, and the energy storage system begins to discharge to meet the nighttime electricity demand.
[0247] When there is sufficient sunlight during the day, the energy storage system performs multiple charging operations to ensure that the battery is fully charged.
[0248] When the light weakened at night, the energy storage system discharged to meet the needs of nighttime users.
[0249] During the day when electricity prices are lower (approximately 8 a.m. to 4 p.m.), the energy storage system mainly performs charging operations, reducing the need to purchase high-priced grid electricity.
[0250] When electricity prices rise at night (approximately 5 pm to 2 am the next morning), the energy storage system begins to discharge, reducing the cost of purchasing electricity at high prices.
[0251] User demand remains constant throughout the day, and the energy storage system ensures the stability of power supply through reasonable scheduling, while also reducing users' electricity bills.
[0252] Photovoltaic panels and wind turbines, as renewable energy power generation devices, fully utilize solar energy to generate electricity during the day and improve the utilization rate of renewable energy.
[0253] By storing excess electricity through energy storage systems and releasing electricity during off-peak hours, reliance on traditional fossil fuel power generation is reduced.
[0254] Throughout the simulation process, the equipment health index remained at a high level (0.95), indicating that the equipment operated well without obvious failures or performance degradation.
[0255] Simulation results show that equipment maintenance is not required within a given time frame, which helps reduce maintenance costs and downtime.
[0256] These results indicate that the DRL framework has obvious advantages in optimizing the operation strategy of multi-energy complementary systems, which helps to achieve efficient, economical and sustainable operation of the systems.
[0257] Corresponding to the above method, the present invention also provides an apparatus / system, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the apparatus / system implements the steps of the method described above.
[0258] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0259] In summary, the energy dispatching method and system for railway stations described in the present invention, based on the multi-source complementary power supply system constructed by the railway station, constructs environmental parameters, electricity price state parameters, energy equipment state parameters and user behavior parameters into a state space, constructs multi-type power generation equipment control parameters, energy storage equipment control parameters, power trading control parameters and load equipment control parameters into an action space, selects control actions based on deep reinforcement learning DRL, and realizes dynamic dispatching and management solutions for various power supply equipment, energy storage equipment and load equipment in the system. The cost savings, energy utilization efficiency, user satisfaction, equipment wear and environmental impact are introduced to calculate the reward value, ensuring that costs can be reduced and efficiency can be increased in the control optimization process, improving user satisfaction, reducing equipment losses and carbon emissions, and realizing automated, efficient and stable energy dispatching.
[0260] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0261] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.
[0262] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0263] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for energy dispatching at a railway station, characterized in that: The method is applicable to a multi-source complementary power system, which includes a plurality of power generation equipment, energy storage equipment and load equipment connected to the grid, and comprises the following steps: Constructing a state space, wherein the state space includes environmental parameters, electricity price state parameters, energy equipment state parameters and user behavior parameters of the target railway station; the environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; the electricity price state parameters include electricity purchase and selling prices; the energy equipment state parameters include current power generation and expected power generation of multiple types of power generation equipment, remaining power and charge and discharge status of energy storage equipment, energy demand of load equipment, equipment health status parameters and carbon emissions; the user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment use preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters; Constructing an action space, wherein the action space includes control parameters for multiple types of power generation equipment, energy storage equipment, power trading control parameters, and load equipment control parameters; Based on deep reinforcement learning, a main network including a main policy network and a main valuation network is constructed, wherein the main policy network takes the parameters of the state space at the current moment as input and outputs the control action selected in the action space; the reward value is calculated in combination with cost savings, energy utilization efficiency, user satisfaction, equipment wear and tear, and environmental impact, and a value function of the action selected based on the state is established according to the reward value, and the main valuation network fits the value function to obtain a main Q value; an experience replay buffer is constructed to store the data generated by each interaction for random sampling training and updating of the main network; wherein a target network with the same structure as the main policy network and the main valuation network is introduced, and the target Q value obtained by fitting the value function of the target network and the main Q value are used to calculate the mean square error as the loss function, and the parameters of the main network are updated by minimizing the loss function, and the parameters of the main network are periodically copied to the target network; The control actions selected by the main network are executed to realize energy dispatch of the target railway station, and the main network is continuously updated and optimized.
2. The energy dispatching method for railway stations according to claim 1, characterized in that: The external natural environment parameters include meteorological information, geological information, and time information of the area where the target railway station is located; the meteorological information includes light intensity, wind speed, temperature, and humidity; the time information includes date, season, and current time; The power generation equipment includes: photovoltaic panels, wind turbines and cogeneration subsystems; The equipment usage preference parameters are quantified by marking the load that the user chooses to run, the frequency of use of the load equipment, and the load operation time; The user satisfaction feedback parameters include the satisfaction score fed back by the user through a preset link, the quantified adjustment opinion, the use frequency of the load device and the recommendation index of the load device; The power generation equipment control parameters include start and stop control parameters and operating power control parameters of the power generation equipment; The energy storage device control parameters include scheduling parameters and charge and discharge power control parameters for multiple energy storage devices; The power purchase and sale control parameters include sale scheduling parameters for the output power of the power generation equipment and the power stored in the energy storage equipment, as well as external power purchase control parameters; The load device control parameters include start / stop control parameters and operation status control parameters for each load device; The estimated power generation is predicted by an ANN neural network.
3. The energy dispatching method for railway stations according to claim 2, characterized in that: The main strategy network and the main valuation network adopt a fully connected network, a convolutional neural network or a recurrent neural network, and introduce a ReLU or sigmoid activation function.
4. The energy dispatching method for railway stations according to claim 3, characterized in that: The bonus value is calculated by combining cost savings, energy efficiency, user satisfaction, equipment wear and tear, and environmental impact, including: The purchase cost of the multi-source complementary power system is calculated as follows: R cost_saving =-(Price purchased ×Power purchased ); Among them, R cost_saving Indicates the purchase cost, Price purchased Indicates the purchase price of electricity, Power purchased Indicates the purchased electricity; The sales revenue of the multi-source complementary power system is calculated as follows: R sales_gain =Price sold ×Power sold ; Among them, R sales_gain represents the sales revenue, Price sold Indicates the sales price of electricity, Power sold Indicates the amount of electricity sold; The energy utilization rate of the multi-source complementary power system at each moment is obtained, and the change in energy utilization rate at the current moment is calculated, and the calculation formula is: R efficiency =Efficiency current -Efficiency previous ; Among them, R efficiency Indicates the change in energy utilization rate, Efficiency current Indicates the current energy utilization rate, Efficiency previous Indicates the energy utilization rate at the last moment; The health status of the multi-source complementary power system at each moment is obtained to characterize the degree of wear of the equipment, and the change in the health status of the system at the current moment is calculated. The calculation formula is: R health =Health current -Health previous ; Among them, R health Indicates the change in the health status of the system. current Indicates the health status of the multi-source complementary power system at the current moment, Health previous Indicates the health status of the multi-source complementary power system at the last moment; the health status is calibrated using the remaining service life or failure rate of each device; The reward value is calculated as follows: R t =ω1(R cost_savin g +R sales_gain )+ω2·R efficiency +ω3·R satisfacti on +ω4·R halth +ω5·R environmen tal ; R environmen tal =-Emissions current ; Among them, R t represents the reward value at time t, R satisfacti on is the quantitative satisfaction score, R environmen tal is the environmental impact value, Emissions current is the carbon emissions of the multi-source complementary power system, and ω1, ω2, ω3, ω4 and ω5 are weight coefficients.
5. The energy dispatching method for railway stations according to claim 4, characterized in that: The main valuation network fits the value function to obtain the main Q value, which is expressed as: Q(s,a,θ)=E[R t +γmax a′ Q(s′,a′,θ)∣s,a]; Where Q(s,a,θ) represents the main Q value, E represents the expectation, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ is the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ is the parameter of the main valuation network; The target Q value obtained by fitting the value function to the target network is expressed as: y=R t +γmax a′ Q target (s′,a′,θ′); Where y represents the target Q value, R t represents the reward value at time t, s represents the state at time t, a represents the action selected at time t, s′ represents the state at time t+1, a′ is the action that can obtain the maximum value function value in state s′, γ represents the attenuation factor; θ′ is the parameter of the target network.
6. The energy dispatching method for railway stations according to claim 5, characterized in that: The expression of the loss function is: L(θ)=E[(yQ(s,a,θ)) 2 ] Among them, L(θ) represents the loss function, E represents expectation, y represents the target Q value, and Q(s, a, θ) represents the main Q value.
7. The energy dispatching method for railway stations according to claim 6, characterized in that: The main strategy network and the main valuation network are pre-trained based on the historical state space and historical action space constructed based on historical data, and are migrated to the multi-source complementary power system for operation and continuous updating.
8. A multi-source complementary power system suitable for railway stations, characterized in that: The system comprises: A multi-source power generation subsystem, including a photovoltaic panel, a wind turbine and a cogeneration subsystem, wherein the multi-source power generation subsystem is connected to an external power grid; An energy storage device subsystem, connected to the multi-source power generation subsystem and the external power grid; A plurality of load devices, connecting the multi-source power generation subsystem, the energy storage device subsystem and the external power grid; An equipment management subsystem, used to execute the energy dispatching method for railway stations as described in any one of claims 1 to 7, select and execute control actions according to environmental parameters, electricity price status parameters, energy equipment status parameters and user behavior parameters to achieve energy dispatching of the target railway station; the environmental parameters include external natural environment parameters and temperature and humidity parameters inside the railway station; the electricity price status parameters include electricity purchase and selling prices; the energy equipment status parameters include current power generation and expected power generation of multiple types of power generation equipment, remaining power and charge and discharge status of energy storage equipment, energy demand of load equipment, equipment health status parameters and carbon emissions; the user behavior parameters include power demand parameters, indoor temperature and humidity demand parameters, equipment usage preference parameters, cost sensitivity parameters, environmental preference parameters and user satisfaction feedback parameters.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Hybrid energy scheduling method for commercial building comprising electric vehicle charging station
CN116468291A
Cited By
Electric energy dispatching method and system based on railway mobile energy storage
CN120300853A