Energy distribution optimization method and device, electronic equipment and storage medium
By employing a two-stage collaborative architecture and a multi-objective optimization method, the problem of balancing supply security and economic efficiency in energy supply has been solved, enabling efficient and reliable energy allocation decisions that can adapt to multiple constraints and uncertainties.
Patent Information
- Application Number
- CN202610070898.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2046-01-20
AI Technical Summary
Existing technologies cannot effectively balance supply security and economic efficiency in multi-regional and multi-cycle energy supply, and are difficult to cope with multiple constraints and uncertainties, leading to overly conservative allocation or delayed response.
A two-stage collaborative architecture is adopted. In the first stage, an allocation sub-model based on a multi-objective optimization function and a set of constraints is constructed to generate baseline allocation parameters. In the second stage, a multi-objective reinforcement learning framework is used for dynamic adjustment. The long short-term memory neural network and the Markov state transition process are combined to simulate uncertainty scenarios and generate dynamically adjusted parameters.
It achieves highly reliable and economical energy distribution in complex environments, shortens decision response time, and enhances the system's dynamic adaptability and resilience in extreme scenarios.
Smart Images

Figure CN121544008A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to the field of energy dispatching technology, and more specifically to an energy allocation optimization method, apparatus, electronic device, and storage medium. Background Technology
[0002] In modern energy supply chain management, energy suppliers face dynamic decision-making challenges across multiple regions, cycles, and objectives, and need to achieve efficient allocation under multiple constraints such as procurement costs, supply stability, demand fluctuations, and storage and transportation security boundaries.
[0003] In related technologies, robust allocation methods based on stochastic programming fail to capture the high-dimensional interconnected chain reactions of demand, supply, and transportation due to the assumption of independent distribution of uncertainty sources. Furthermore, the exponential growth in the number of scenarios leads to the curse of dimensionality, and overly conservative configurations result in additional inventory costs. Dynamic scheduling methods based on single-objective reinforcement learning struggle to adapt to dynamic priority changes in different load areas due to manually weighted multi-objective rewards, and the exploration process lacks physical constraints and suffers from low sample efficiency. Rolling optimization based on model predictive control relies on deterministic predictions, resulting in a lag in response to sudden disturbances. Therefore, there is an urgent need for an energy allocation optimization method that can balance supply security and economic efficiency. Summary of the Invention
[0004] In view of the above problems, this application provides an energy allocation optimization method, apparatus, electronic device and storage medium to improve the development efficiency of industrial intelligent agents and reduce development costs.
[0005] According to the first aspect of this application, an energy allocation optimization method is provided, comprising: acquiring energy procurement cost data and energy supply attribute data of energy suppliers during a target scheduling period, as well as energy demand attribute data and energy constraint attribute data associated with each energy load area, wherein the energy supply attribute data characterizes the energy supply quantity characteristics and supply stability characteristics of each energy load area, the energy demand attribute data characterizes the demand fluctuation characteristics and demand balance characteristics of the energy load area, and the energy constraint attribute data characterizes the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load area; processing the energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data based on an energy allocation prediction model to obtain energy allocation information, wherein the energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraint conditions; and determining a target energy allocation scheme corresponding to each energy load area based on the energy allocation information, wherein the target energy allocation scheme is used for energy allocation between the energy supplier and each energy load area.
[0006] The method provided in this application, by constructing a first-stage allocation sub-model based on a first multi-objective optimization function and a first set of constraints, can systematically integrate multi-dimensional heterogeneous data such as energy procurement costs, supply and stability, demand fluctuations and balance, storage and transportation boundaries and allocation priorities. Within the target scheduling cycle, it achieves synergistic optimization and Pareto trade-offs of key objectives such as cost, reliability, and timeliness. This effectively overcomes the shortcomings of traditional single-objective optimization or deterministic programming methods that cannot take into account multiple constraints and multiple attribute characteristics. It significantly improves the overall robustness and feasibility of energy allocation schemes in complex and uncertain environments, reduces the comprehensive cost of energy allocation, and the generated scheme strictly satisfies storage and transportation boundaries and priority constraints, eliminating the constraint violation risk in related technologies.
[0007] According to an embodiment of this application, the energy allocation prediction model includes a second-stage allocation sub-model. Based on the energy allocation prediction model, energy procurement costs, energy supply attributes, energy demand attributes, and energy constraint attributes are processed to obtain energy allocation information, including: processing energy supply capacity data, energy procurement cost data, and energy constraint attribute data based on the first-stage allocation sub-model to generate benchmark allocation parameters. The benchmark allocation parameters satisfy the energy storage and energy transportation boundaries represented by the energy constraint attribute data; processing energy demand attribute data and energy constraint attribute data based on the second-stage allocation sub-model, and generating dynamic adjustment parameters within a constraint interval formed by the constraint lower limit values determined by the benchmark allocation parameters. The second-stage allocation sub-model is constructed based on a multi-objective reinforcement learning framework, which is trained using vector reward signals and action space boundaries. The vector reward signals represent the comprehensive return of multiple optimization objectives, including total energy cost, demand satisfaction rate, and delivery delay time. The action space boundaries are determined based on the benchmark allocation parameters and are used to limit the parameter range of the dynamic adjustment parameters within the constraint interval corresponding to the benchmark allocation parameters; and generating energy allocation information based on the benchmark allocation parameters and the dynamic adjustment parameters.
[0008] Through a two-stage collaborative architecture, the first stage generates baseline allocation parameters that satisfy the storage and transportation boundary to establish a guaranteed feasible solution. The second stage uses the action space boundary determined by these baseline parameters as a hard constraint and a multi-objective reinforcement learning framework as the engine to dynamically generate and adjust parameters within the constraint range. This achieves decoupling and complementarity between the underlying robustness and the top-level agility, avoiding the infeasibility risk of pure reinforcement learning exploring beyond the boundary, and overcoming the shortcomings of traditional robust optimization being overly conservative and unable to respond in real time. By using vector reward signals to explicitly quantify and comprehensively evaluate multiple objectives such as cost, demand satisfaction rate, and delivery delay, the policy network can learn Pareto trade-offs online. The final output energy allocation information has both high reliability and high economy, and the decision response time is shortened to the minute level, significantly improving the dynamic adaptive capability of complex energy systems in uncertain environments.
[0009] According to embodiments of this application, the generation of benchmark allocation parameters based on energy supply capacity data, energy procurement cost data, and energy constraint attribute data processed by the first-stage allocation sub-model includes: determining the energy storage and transportation boundary characteristics and safety stock threshold constraints of energy load areas based on energy constraint attribute data, wherein the energy storage and transportation boundary characteristics are used to limit the search space of the benchmark allocation parameters; and constructing a set of uncertainty scenarios associated with energy load areas based on energy supply capacity data, energy procurement cost data, and historical operational statistics of each energy load area, wherein the uncertainty scenarios are used to characterize the energy supply and demand balance state, energy transportation cost state, and energy safety reserve state of the corresponding energy load area. The dynamic influence relationship between them is analyzed. For each scenario in the set of uncertain scenarios, the energy allocation regret value of the current candidate allocation scheme is calculated in the search space. The energy allocation regret value indicates the performance gap between the current candidate allocation scheme and the ideal allocation scheme corresponding to the scenario in terms of total energy cost, demand satisfaction gap, and delivery delay time. The current candidate allocation scheme satisfies the safety stock threshold constraint. The conditional risk value is obtained by weighted averaging the energy allocation regret values of the tail extreme scenario subset. The regret value of the tail extreme scenario subset is greater than the preset threshold. With the goal of minimizing the conditional risk value, the Pareto optimal solution set is output as the benchmark allocation parameter through iterative evolution using the third-generation non-dominated sorting genetic algorithm.
[0010] By constructing energy feasible boundaries and safety stock threshold constraints to limit the search space, and generating a set of uncertain scenarios representing the dynamic coupling relationship between supply and demand balance, transportation costs and safety reserves based on historical operational statistics, a regret value mechanism is used to quantify the multi-objective performance gap between candidate solutions and ideal solutions in each scenario. Then, a Conditional Value at Risk (CVaR) weighted average is applied to the tail extreme scenario subset to focus on preventing low-probability high-loss risks. Finally, a Pareto optimal solution set is output through iterative evolution using a genetic algorithm as the benchmark for parameter allocation. This achieves a scientific trade-off between resilience, economy and reliability in extreme scenarios, ensuring that the generated fallback solution covers extreme conditions while avoiding excessive conservatism, and providing a high-confidence safety boundary benchmark for subsequent dynamic adjustments.
[0011] According to embodiments of this application, constructing a set of uncertainty scenarios associated with energy load areas includes: using a prediction model constructed by combining a long short-term memory neural network and a prophetic prediction model to process historical and real-time demand data for the corresponding energy load areas to obtain predicted demand disturbance values. The long short-term memory neural network is used to capture periodic demand trends, while the prophetic prediction model uses a Poisson impulse response function to characterize the sudden impact of extreme weather on demand; a Markov state transition process is used to simulate the state changes of the supply path in each energy load area to obtain supply state transition probabilities, where the state transition probabilities exhibit an exponential decay characteristic as the time interval increases; and a time-varying transportation cost disturbance model is used to process transportation base cost data and fuel price fluctuation data to determine the transportation cost disturbance vector.
[0012] By employing a spatiotemporal hybrid model to capture cyclical demand trends through a long short-term memory network and supplementing it with a Poisson impulse response function to quantitatively characterize the instantaneous impact of sudden events such as extreme weather, a Markov state transition process is used to simulate the interruption-recovery state changes of each supply path, and an exponential decay factor is introduced to avoid overly pessimistic long-term risk estimation. Combined with a time-varying transportation cost disturbance model, fuel price fluctuations are superimposed with the impact of sudden events, thereby generating a set of uncertain scenarios with multidimensional correlation between demand, supply, and transportation. This improves the accuracy of characterizing the chain disturbance propagation mechanism in complex energy systems and reduces scenario prediction errors.
[0013] According to embodiments of this application, energy demand attribute data and energy constraint attribute data are processed based on the second-stage allocation sub-model. Within the constraint interval formed by the constraint lower limit value determined by the baseline allocation parameters, dynamic adjustment parameters are generated, including: determining the real-time state vector of each energy load region at the current scheduling time based on the energy demand attribute data and energy constraint attribute data. The real-time state vector includes the regional inventory level, transportation route status, real-time demand deviation, cumulative cost value, and cumulative delay value of each energy load region; calculating a vector reward signal based on the real-time state vector. The vector reward signal includes a cost reward component obtained from the cumulative cost value, a demand satisfaction rate reward component obtained from the real-time demand deviation, and a timeliness component obtained from the cumulative delay value. The algorithm evaluates the multi-objective state value of candidate actions and calculates the advantage function based on the vector reward signal and the action search space. It outputs the candidate action vector that maximizes the vector reward signal. The candidate action vector includes emergency energy procurement quantity, inventory allocation quantity, and transportation route switching decisions. Within the action search space, the algorithm performs feasibility checks on the candidate action vectors, pruning action components that exceed the constraint lower limit to the boundary value, generating dynamically adjusted parameters. This process determines the lower limit of constraints based on the baseline allocation parameters and constructs an action search space. The algorithm also determines the minimum inventory threshold, emergency procurement quantity, and transportation route availability constraints based on the minimum inventory threshold and the transportation route availability constraints.
[0014] By explicitly separating cost, demand satisfaction rate, and timeliness reward components from vector reward signals, a multi-objective proximal policy optimization network is used to evaluate state value and calculate advantage function within the action search space formed by hard constraint limits. Combined with a boundary pruning mechanism, out-of-bounds actions are forcibly constrained to the feasible region. This not only eliminates the security risks in the reinforcement learning exploration process in related technologies, but also achieves minute-level online adaptive decision-making, enabling the policy network to converge with hourly sample updates and improving dynamic response speed.
[0015] According to an embodiment of this application, the method further includes: determining the scenario feature vector of each energy load region at the current scheduling time based on energy demand attribute data and energy constraint attribute data, wherein the scenario feature vector includes the regional inventory level, transportation path status and real-time demand deviation of each energy load region; determining the target scenario feature vector with the highest matching degree in the uncertain scenario set based on the scenario feature vector at the current scheduling time; and determining the reinforcement learning policy parameters corresponding to the target scenario represented by the target scenario feature vector through a pre-trained scenario policy library to initialize a multi-objective proximal policy optimization network, wherein the scenario policy library is used to represent the mapping relationship between scenarios and policy parameters.
[0016] By quickly matching the current state with the most similar uncertain scenario in the historical database using scene feature vectors, and loading the corresponding reinforcement learning policy parameters from the pre-trained scene policy library, the multi-objective proximal policy optimization network is warm-started. This compresses the tens of thousands of exploration iterations required for cold start into a single forward inference, reducing the policy initialization time from hours to milliseconds. This significantly improves the real-time decision-making capability of the multi-objective proximal policy optimization network in emergency situations such as sudden energy outages. At the same time, it enhances the generalization performance of the policy under new conditions by utilizing historical scene experience transfer.
[0017] According to an embodiment of this application, the first multi-objective optimization function represents minimizing the total energy cost, the demand satisfaction gap, and the delivery delay time as the optimization objective. The first set of constraints includes at least one of the following: a supply capacity constraint characterizing the limiting relationship between energy supply and the upper limit of supply capacity; an inventory safety constraint characterizing the balance between inventory level and safety stock threshold; and a transportation feasibility constraint characterizing the limiting relationship between transportation route selection status and route availability.
[0018] The first multi-objective optimization function incorporates total energy cost, demand satisfaction gap, and delivery delay time into the minimization objective system. Combined with the first set of constraints consisting of supply capacity constraints, inventory security constraints, and transportation feasibility constraints, it provides a complete framework for defining optimization objectives and feasible regions for the two-stage allocation sub-model. This ensures that the output target energy allocation scheme not only conforms to physical boundaries and policy priorities, but also achieves explicit quantitative trade-offs among multi-dimensional objectives, supporting energy suppliers to flexibly adjust their decision preferences according to different supply guarantee requirements.
[0019] The second aspect of this application provides an energy allocation optimization device, comprising: an acquisition module for acquiring energy procurement cost data and energy supply attribute data of energy suppliers during a target scheduling period, as well as energy demand attribute data and energy constraint attribute data associated with each energy load area, wherein the energy supply attribute data characterizes the energy supply quantity characteristics and supply stability characteristics of each energy load area, the energy demand attribute data characterizes the demand fluctuation characteristics and demand balance characteristics of the energy load area, and the energy constraint attribute data characterizes the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load area; an energy allocation prediction module for processing the energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data based on an energy allocation prediction model to obtain energy allocation information, wherein the energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraint conditions; and an allocation scheme determination module for determining a target energy allocation scheme corresponding to each energy load area based on the energy allocation information, wherein the target energy allocation scheme is used for energy allocation between the energy supplier and each energy load area.
[0020] The energy distribution optimization device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its principle and beneficial effects are similar, and will not be described again here.
[0021] According to an embodiment of this application, the energy distribution prediction module includes a first generation submodule, a second generation submodule, and a third generation submodule.
[0022] The first generation submodule processes energy supply capacity data, energy procurement cost data, and energy constraint attribute data based on the first-stage allocation submodule to generate baseline allocation parameters. These baseline allocation parameters satisfy the energy storage and transportation boundaries represented by the energy constraint attribute data. The second generation submodule processes energy demand attribute data and energy constraint attribute data based on the second-stage allocation submodule. Within the constraint interval defined by the lower limit values determined by the baseline allocation parameters, it generates dynamic adjustment parameters. The second-stage allocation submodule is constructed based on a multi-objective reinforcement learning framework, which is trained using vector reward signals and action space boundaries. The vector reward signals represent the comprehensive reward of multiple optimization objectives, including total energy cost, demand satisfaction rate, and delivery delay time. The action space boundaries are determined based on the baseline allocation parameters and are used to limit the range of the dynamic adjustment parameters within the constraint interval corresponding to the baseline allocation parameters. Finally, the third generation submodule generates energy allocation information based on the baseline allocation parameters and the dynamic adjustment parameters.
[0023] According to an embodiment of this application, the first generation submodule includes a first determination unit, a scene construction unit, a first calculation unit, a second calculation unit, and an output unit.
[0024] The first determining unit is used to determine the energy storage and transportation boundary characteristics and safety stock threshold constraints of the energy load area based on energy constraint attribute data. The energy storage and transportation boundary characteristics are used to limit the search space of the benchmark allocation parameters. The scenario construction unit is used to construct a set of uncertain scenarios associated with the energy load area based on energy supply capacity data, energy procurement cost data, and historical operation statistics of each energy load area. The uncertain scenarios are used to characterize the dynamic influence relationship between the energy supply and demand balance state, energy transportation cost state, and energy safety reserve state of the corresponding energy load area. The first calculation unit is used to calculate each of the uncertain scenario sets. The first unit calculates the energy allocation regret value of the current candidate allocation scheme within the search space. The energy allocation regret value indicates the performance gap between the current candidate allocation scheme and the ideal allocation scheme corresponding to the scenario in terms of total energy cost, demand fulfillment rate gap, and delivery delay time. The current candidate allocation scheme meets the safety stock threshold constraint. The second calculation unit calculates the conditional risk value by weighting the energy allocation regret values of the tail extreme scenario subset. The regret value of the tail extreme scenario subset is greater than a preset threshold. The output unit is used to output the Pareto optimal solution set as the benchmark allocation parameter by iteratively evolving through the third-generation non-dominated sorting genetic algorithm with the goal of minimizing the conditional risk value.
[0025] According to embodiments of this application, the scenario construction unit is further used to process historical and real-time demand data of the corresponding energy load area using a prediction model constructed by a hybrid long short-term memory neural network and a prophetic prediction model, to obtain predicted demand disturbance values. The long short-term memory neural network is used to capture periodic demand trends, and the prophetic prediction model uses a Poisson impulse response function to characterize the sudden impact of extreme weather on demand. A Markov state transition process is used to simulate the state changes of the supply path in each energy area to obtain the supply state transition probability, wherein the state transition probability exhibits an exponential decay characteristic as the time interval increases. The unit also processes transportation base cost data and fuel price fluctuation data through a time-varying transportation cost disturbance model to determine the transportation cost disturbance vector.
[0026] According to an embodiment of this application, the second generation submodule further includes a second determining unit, a third determining unit, a search unit, and a verification unit.
[0027] The second determining unit is used to determine the real-time state vector of each energy load region at the current scheduling time based on energy demand attribute data and energy constraint attribute data. The real-time state vector includes the regional inventory level, transportation path status, real-time demand deviation, accumulated cost value, and accumulated delay value of each energy region. The third determining unit calculates a vector reward signal based on the real-time state vector. The vector reward signal includes a cost reward component obtained from the accumulated cost value, a demand satisfaction rate reward component obtained from the real-time demand deviation, and a timeliness reward component obtained from the accumulated delay value. The search unit is used to determine the constraint lower limit value based on the benchmark allocation parameters, construct the action search space, and approximately... The lower bounds include minimum inventory threshold, minimum emergency procurement quantity, and transportation route availability constraints; they are used to input real-time state vectors into a multi-objective proximal policy optimization network. This network evaluates the multi-objective state value of candidate actions and calculates the advantage function based on the vector reward signal and the action search space, outputting the candidate action vector that maximizes the vector reward signal. The candidate action vectors include emergency energy procurement quantity, inventory allocation quantity, and transportation route switching decisions; a verification unit is used to perform feasibility verification on the candidate action vectors within the action search space, pruning action components exceeding the constraint lower bounds to the boundary values, and generating dynamically adjusted parameters.
[0028] According to embodiments of this application, the apparatus further includes: a scene determination module, a scene matching module, and a scene strategy parameter mapping module.
[0029] The scenario determination module is used to determine the scenario feature vector of each energy load area at the current scheduling time based on energy demand attribute data and energy constraint attribute data. The scenario feature vector includes the regional inventory level, transportation path status and real-time demand deviation of each energy load area. The scenario matching module is used to determine the target scenario feature vector with the highest matching degree in the uncertain scenario set based on the scenario feature vector at the current scheduling time. The scenario policy parameter mapping module is used to determine the reinforcement learning policy parameters corresponding to the target scenario represented by the target scenario feature vector through a pre-trained scenario policy library, so as to initialize the multi-objective proximal policy optimization network. The scenario policy library is used to represent the mapping relationship between scenarios and policy parameters.
[0030] The energy distribution optimization device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its principle and beneficial effects are similar, and will not be described again here.
[0031] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0032] The electronic device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be repeated here.
[0033] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0034] When the computer-executable instructions in the computer-readable storage medium provided in this application are executed by the processor, they can implement the technical solutions shown in the above method embodiments. The implementation principle and beneficial effects are similar, and will not be repeated here. Attached Figure Description
[0035] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0036] Figure 1 An application scenario diagram of the energy allocation optimization method according to an embodiment of this application is shown.
[0037] Figure 2 A flowchart of an energy allocation optimization method according to an embodiment of this application is shown.
[0038] Figure 3 A flowchart of an energy allocation information generation method according to an embodiment of this application is shown.
[0039] Figure 4 A flowchart of a method for generating dynamically adjusted parameters according to an embodiment of this application is shown.
[0040] Figure 5 A structural block diagram of an energy distribution optimization device according to an embodiment of this application is shown.
[0041] Figure 6 A block diagram of an electronic device suitable for implementing an energy distribution optimization method according to an embodiment of this application is shown. Detailed Implementation
[0042] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0043] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0044] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0045] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0046] In modern energy supply chain management, energy suppliers (such as coal groups and integrated energy companies) need to efficiently allocate energy to various energy load areas (such as thermal power plants, steel plants, and urban heating systems) under multi-regional and multi-cycle conditions. This process involves multiple constraints, including energy procurement costs, supply stability, demand volatility, and storage and transportation safety boundaries, and is a typical multi-objective, strongly constrained, and highly uncertain dynamic decision-making problem. Traditional energy allocation mainly relies on experience-based scheduling or static programming models, and its decision-making process is usually based on deterministic forecasts for single-cycle optimization, lacking the ability to dynamically respond to uncertainties across the entire supply and demand chain. With the deepening of energy market reforms and the frequent occurrence of extreme weather, the daily fluctuation range of demand can reach ±15%, and unplanned shutdowns due to mine accidents, policy adjustments, and transportation disruptions occur 3-5 times per year on the supply side. Static models can no longer meet the modern energy supply guarantee requirements of a reliability of no less than 95% and a decision response time of less than 1 hour.
[0047] Among the related technologies, there are three main types of technical paths:
[0048] 1. Robust assignment method based on stochastic programming
[0049] This method constructs a two-stage model that minimizes expected costs by pre-setting a finite number of discrete scenarios (such as high / medium / low demand scenarios). However, this type of method has the following drawbacks: Scenario simplification leads to distortion; it assumes that each uncertainty source (demand, supply, transportation) is independently distributed, failing to depict the high-dimensional chain reaction of "extreme weather → surge in demand + transportation disruption + soaring costs," resulting in incomplete scenario coverage and unexpected losses in practical applications; high computational complexity, with the number of scenarios increasing exponentially with the dimension of uncertainty; 20 nodes × 5 states generate tens of millions of scenario combinations, posing a dimensionality disaster for real-time decision-making; and rigid conservatism, often over-configuring safety stock to cover worst-case scenarios, increasing inventory costs and resulting in significant economic losses.
[0050] 2. Dynamic scheduling method based on single-objective reinforcement learning
[0051] Deep reinforcement learning can be applied to energy dispatch, such as using dual deep Q-network algorithms for real-time path selection. However, these algorithms suffer from several drawbacks: scalar reward problem, requiring multiple objectives such as cost, reliability, and timeliness to be manually weighted into a single reward, with weighting relying on subjective experience and failing to reflect dynamic priority changes in different load areas; lack of safety constraints, as the reinforcement learning exploration process may violate physical constraints, leading to infeasible strategies and extremely high deployment risks; and low sample efficiency, requiring millions of trial-and-error samples to converge, while actual energy dispatch data contains fewer than a thousand effective samples per year, resulting in severe overfitting and insufficient generalization ability.
[0052] 3. Rolling Optimization Method Based on Model Predictive Control
[0053] While it performs online optimization through rolling time domain, it relies on deterministic prediction models and has a response lag of at least 4-6 hours to sudden disturbances (such as coal mine accidents). Furthermore, it does not explicitly consider multi-objective Pareto trade-offs and cannot meet the flexible switching needs of policy orientations such as "supply guarantee first, economy second".
[0054] Based on the aforementioned technical problems, embodiments of this application provide an energy allocation optimization method. The method includes: acquiring energy procurement cost data and energy supply attribute data of the energy supplier during a target scheduling period, as well as energy demand attribute data and energy constraint attribute data associated with each energy load region. The energy supply attribute data characterizes the energy supply quantity and supply stability characteristics of each energy load region; the energy demand attribute data characterizes the demand fluctuation and demand balance characteristics of the energy load region; and the energy constraint attribute data characterizes the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load region. The method further involves processing the energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data based on an energy allocation prediction model to obtain energy allocation information. The energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraint conditions. Finally, the method involves determining a target energy allocation scheme corresponding to each energy load region based on the energy allocation information. The target energy allocation scheme is used for energy allocation between the energy supplier and each energy load region.
[0055] Figure 1 An application scenario diagram of the energy allocation optimization method according to an embodiment of this application is shown.
[0056] like Figure 1 As shown, application scenario 100 according to this embodiment may include an energy distribution scenario. Network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0057] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0058] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0059] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0060] It should be noted that the energy allocation optimization method provided in this embodiment can generally be executed by server 105. Correspondingly, the energy allocation optimization device provided in this embodiment can generally be located in server 105. The energy allocation optimization method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the energy allocation optimization device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0061] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0062] The following will be based on Figure 1 The described scene, through Figures 2-4 The energy allocation optimization method of the disclosed embodiments is described in detail.
[0063] Figure 2 A flowchart of an energy allocation optimization method according to an embodiment of this application is shown.
[0064] like Figure 2 As shown, the energy allocation optimization method of this embodiment includes operations S210 to S240, which can be executed by a server or other computing device.
[0065] In operation S210, energy suppliers acquire energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data associated with each energy load region during the target scheduling period.
[0066] According to embodiments of this application, energy supply attribute data characterizes the energy supply quantity and supply stability characteristics of each energy load area, energy demand attribute data characterizes the demand fluctuation characteristics and demand balance characteristics of the energy load area, and energy constraint attribute data characterizes the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load area.
[0067] In one example, the energy supplier, such as an energy group's dispatch center, first acquires three types of core data for the target dispatch period: energy procurement cost data, energy supply attribute data, and energy demand attribute data, as well as energy constraint attribute data. The energy procurement cost data includes: long-term contract energy purchase price, market spot price fluctuation range, transportation rate benchmark, storage cost parameters, and penalty cost coefficients for delayed delivery. For coal energy, this data corresponds to pithead price, ocean freight, railway freight rates, and demurrage penalties. The energy supply attribute data characterizes the supply capacity and stability characteristics of each energy load region. Specifically, supply volume characteristics include the upper limit of long-term contract volumes for each energy load region, the maximum procurement volume in the spot market, and the available emergency reserve allocation; supply stability characteristics are characterized by quantitative indicators such as historical interruption frequency, recovery time, and supplier credit rating. The demand attribute data includes historical demand time series, seasonal fluctuation patterns, real-time demand deviations, and the impulse impact function of sudden events such as extreme weather for each load region. Constraint attribute data defines the boundary characteristics of energy storage and transportation, such as warehouse capacity limits, maximum throughput capacity of transportation channels, minimum safety stock thresholds, and allocation priority characteristics. For example, residential heating has a higher priority than industrial interruptible loads. This data is collected in real time through IoT sensors, enterprise resource planning systems, and external meteorological / policy data interfaces, and is standardized and preprocessed to form a structured dataset.
[0068] In operation S220, energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data are processed based on the energy allocation prediction model to obtain energy allocation information.
[0069] According to an embodiment of this application, the energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraints.
[0070] According to an embodiment of this application, the first multi-objective optimization function represents minimizing the total energy cost, the demand satisfaction gap, and the delivery delay time as the optimization objective. The first set of constraints includes at least one of the following: a supply capacity constraint characterizing the limiting relationship between energy supply and the upper limit of supply capacity; an inventory safety constraint characterizing the balance between inventory level and safety stock threshold; and a transportation feasibility constraint characterizing the limiting relationship between transportation route selection status and route availability.
[0071] In one example, the first multi-objective optimization function simultaneously minimizes three conflicting objectives: the total energy cost objective, which includes the weighted sum of procurement costs, transportation costs, inventory holding costs, and stockout penalty costs; the demand fulfillment gap objective, which includes quantifying unmet demand to reflect supply reliability; and the delivery delay time objective, which includes the weighted sum of cumulative delayed delivery times to reflect timeliness. The first set of constraints includes: supply capacity constraints, ensuring that decision variables do not exceed the maximum supply of each path; inventory safety constraints, ensuring that ending inventory does not fall below the minimum safety stock threshold; and transportation feasibility constraints, ensuring that the selected transportation path is available during the target time period. The first-stage allocation sub-model, acting as a robust pre-decision engine, primarily generates a baseline allocation parameter at the start of the target scheduling cycle. The first-stage allocation sub-model receives data on energy supply capacity, procurement costs, and constraint attributes. It then uses a min-max regret value optimization framework to perform a globally robust solution to the set of uncertain scenarios. With the objective of minimizing the Value at Risk (CVaR) of the tail extreme scenarios, the NSGA-III algorithm iteratively outputs a Pareto optimal solution set. The final determined baseline allocation parameters include the minimum procurement quantity for each energy load region, safety stock threshold, and emergency path configuration. These constitute inviolable hard boundary constraints for subsequent online decision-making, ensuring the reliability and feasibility of energy supply under any foreseeable extreme conditions. The optimization process of the first-stage allocation sub-model employs the min-max regret value framework; the specific process can be found in operations S311 to S314, and will not be elaborated further here.
[0072] In operation S230, the target energy allocation scheme corresponding to the energy load area is determined based on the energy allocation information.
[0073] According to embodiments of this application, the target energy allocation scheme is used for energy allocation between energy suppliers and various energy load areas.
[0074] In one example, energy allocation is performed based on the final executable target energy allocation plan output by the energy allocation forecasting model. This plan explicitly specifies the baseline and real-time additional purchase quantities for each energy load area; the minimum pre-stock inventory and the currently recommended allocation quantity for each energy load area; the primary transportation route and emergency switching decisions. After the plan is confirmed by dispatchers, the Energy Management System (EMS) automatically issues purchase orders, allocation instructions, and transportation plans, achieving precise matching and allocation between suppliers and load areas.
[0075] It should be noted that the energy allocation optimization method provided in this embodiment is based on a two-stage robust adaptive collaborative approach, applicable to cross-regional, multi-cycle intelligent allocation of various bulk energy commodities such as coal, natural gas, crude oil, and electricity. This method constructs a unified energy allocation prediction model, abstracting the physical characteristics and common decision-making logic of different energy types into a standardized data processing flow, achieving end-to-end automated decision-making from data acquisition to solution generation. In other words, the framework of this embodiment is energy type independent. For coal energy, supply attribute data could include, for example, corresponding pithead capacity, railway transportation capacity, and port throughput capacity; for pipeline natural gas, supply attribute data could include, for example, corresponding pipeline pressure limits, distribution station flow regulation range, and gas storage facility operating capacity. Regarding demand attribute data, the impact of extreme weather pulses on coal is reflected in a surge in daily coal consumption during cold waves, while for natural gas, it is reflected in a surge in daily gas consumption due to temperature drops. It should be understood that, regardless of the type of energy, the technical logic of spatiotemporal fusion model, min-max regret value optimization, multi-objective reinforcement learning and hard constraint boundary pruning is completely universal. It can be quickly adapted simply by adjusting the parameter dimension of the constraint set and the weight configuration of the objective function according to the physical characteristics of the energy.
[0076] The method provided in this application, by constructing a first-stage allocation sub-model based on a first multi-objective optimization function and a first set of constraints, can systematically integrate multi-dimensional heterogeneous data such as energy procurement costs, supply and stability, demand fluctuations and balance, storage and transportation boundaries and allocation priorities. Within the target scheduling cycle, it achieves synergistic optimization and Pareto trade-offs of key objectives such as cost, reliability, and timeliness. This effectively overcomes the shortcomings of traditional single-objective optimization or deterministic programming methods that cannot take into account multiple constraints and multiple attribute characteristics. It significantly improves the overall robustness and feasibility of energy allocation schemes in complex and uncertain environments, reduces the comprehensive cost of energy allocation, and the generated scheme strictly satisfies storage and transportation boundaries and priority constraints, eliminating the constraint violation risk in related technologies.
[0077] According to embodiments of this application, constructing a set of uncertainty scenarios associated with energy load areas includes: using a prediction model constructed by combining a long short-term memory neural network and a prophetic prediction model to process historical and real-time demand data for the corresponding energy load areas to obtain predicted demand disturbance values. The long short-term memory neural network is used to capture periodic demand trends, while the prophetic prediction model uses a Poisson impulse response function to characterize the sudden impact of extreme weather on demand; a Markov state transition process is used to simulate the state changes of the supply path in each energy load area to obtain supply state transition probabilities, where the state transition probabilities exhibit an exponential decay characteristic as the time interval increases; and a time-varying transportation cost disturbance model is used to process transportation base cost data and fuel price fluctuation data to determine the transportation cost disturbance vector.
[0078] In one example, this embodiment details the process of constructing an uncertain scenario set. Through a hybrid modeling approach combining Long Short-Term Memory (LSTM), Prophet prediction models, and Markov state transition processes, a high-dimensional scenario vector representing the multi-dimensional dynamic coupling relationship between demand, supply, and transportation is generated, providing input for the first-stage robust pre-decision-making. Uncertainty Scenario Set It is constructed collaboratively from the following three dimensions:
[0079] (a) Demand disturbance dimension: LSTM-Prophet hybrid prediction model
[0080] The system utilizes a prediction model constructed by combining a long short-term memory neural network and a prophetic prediction model to process historical and real-time demand data from various energy load regions, thereby obtaining predicted demand disturbance values. The mixed output It simultaneously incorporates trends, cycles, and impulses, accurately depicting the multi-scale uncertainty of demand. Specifically, as shown in formula (1):
[0081] (1)
[0082] in, The demand baseline trend output by the Long Short-Term Memory Neural Network reflects the periodic autoregressive characteristics of historical data. 't' represents the current time point or time step. In this embodiment, it can represent a specific day, hour, or other time granularity, depending on the temporal resolution of the dataset and the needs of the prediction task. 'n' is an integer representing the length of historical data used to predict the current time point 't', specifying how many past time points are considered when making the current prediction. For example, if n=7, when predicting the demand on day 't', the model will consider historical demand data from day 't-1' to day 't-7'. The Prophet (seasonal term) explicitly models annual and quarterly seasonality using Fourier series. Let Poisson be the impulse response function. Let be the impact intensity coefficient of demand on the k-th type of extreme weather event (such as cold wave, typhoon). This represents the increment of the Poisson process, characterizing the number of events occurring within tiny time intervals. When a weather warning is triggered... Demand forecasts can jump by 15%-30% instantaneously, enabling "event-driven" demand disturbance modeling.
[0083] (ii) Supply Disruption Dimension: Markov State Transition Process
[0084] Markov state transition processes are used to simulate the state changes of the supply path in each energy load region, and the supply state transition probabilities are obtained. This is used to characterize the risk of supply disruptions caused by mining accidents, policy adjustments, etc. The state space of each supply path is defined as {0: normal supply, 1: minor disruption, 2: complete disruption}. The state transition probability matrix... This was obtained from historical interruption frequency statistics, among which Let represent the basic transition probability from state i to j. To avoid the overestimation of long-term risk due to the "memorylessness" of traditional Markov chains, this application introduces a time decay factor, causing the state transition probability to decay exponentially with increasing time interval, as shown in formula (2):
[0085] (2)
[0086] in, The basic transition probability is obtained from historical statistics, reflecting the inherent reliability of the path; The time decay coefficient, For time intervals; This is the exponential decay factor, which physically means that as time goes on, the system tends to maintain its current state, and the probability of state transition decays exponentially. For example, after a mining accident in a coal mine, the probability of recovery in the short term is low (…). Smaller size, weaker attenuation), and in the long run, the probability of recovery gradually increases. Large, strong attenuation (Dominant), which conforms to the engineering reality that "short-term interruptions are easy to recover from, while long-term interruptions are difficult to heal on their own".
[0087] Calculate independently for each supply path Generate supply state vector , where m is the total number of paths. This vector is related to the demand disturbance value. The correlation reflects the chain reaction of "surge in demand → tight supply → increased probability of state transition".
[0088] (III) Transportation cost disturbance dimension: Time-varying transportation cost model
[0089] The system processes basic transportation cost data and fuel price fluctuation data using a time-varying transportation cost disturbance model to determine the transportation cost disturbance vector. As in formula (3):
[0090] (3)
[0091] in, Let be the transportation cost of path (x, y) at time t. The basic transportation cost from node x to y is determined by the contract freight rate; This indicates the volatility of transportation costs; Simulating fuel price fluctuations (Wiener process) reflects the sensitivity of transportation costs to energy market fluctuations. This represents the random fluctuation portion of transportation costs; For items impacted by sudden events, For the first The unit impact intensity of events such as port congestion and road closures. A Poisson impulse characterizes the randomness of events. Transportation costs are dynamically coupled with demand and supply. When extreme weather causes a surge in demand, the Poisson impulse... This also triggered a transportation disruption. Instantaneous rise; when the supply path state S(t) = 2 (complete interruption), It is set to a maximum value, forcing the optimization algorithm to automatically switch to an alternative path.
[0092] Output the above three dimensions Combining them to form an uncertain scenario vector ,in The inventory disturbance value is calculated from the safety stock threshold and the actual inventory level. A set of m scenario vectors is generated through Monte Carlo sampling, constituting the uncertainty scenario set. Let Ω represent the specific scenarios within the set, where m is the total number of scenarios in the set. Each scenario can represent a specific situation or state, such as different demand levels, supply disruptions, or changes in transportation costs. This is used for min-max optimization in the first stage of robust pre-decision making. Through spatiotemporal fusion modeling, the uncertainties of demand, supply, and transportation are elevated from independent assumptions to a high-dimensional correlation, providing realistic and comprehensive risk inputs for subsequent robust optimization. This is a key foundation for improving decision resilience.
[0093] Figure 3 A flowchart of an energy allocation information generation method according to an embodiment of this application is shown. Figure 3 As shown, this includes operations S310 to S330.
[0094] In operation S310, energy supply capacity data, energy procurement cost data, and energy constraint attribute data are processed based on the first-stage allocation sub-model to generate baseline allocation parameters.
[0095] According to an embodiment of this application, the baseline allocation parameters satisfy the energy storage boundary and energy transport boundary represented by the energy constraint attribute data.
[0096] In one example, the first-stage allocation sub-model acts as an offline computing engine. Its core function is to construct a quasi-allocation parameter covering all uncertainties at the start of the scheduling cycle. It receives energy supply capacity data (maximum recoverable quantity for each supply path, historical interruption frequency, and recovery time distribution), energy procurement cost data (long-term contract benchmark price, spot market price, transportation rates, and penalty cost coefficients), and energy constraint attribute data (warehouse capacity limit, transportation channel capacity, and minimum safety stock threshold). Based on this data, an uncertainty scenario set Ω is first constructed according to the above embodiment. The benchmark allocation parameter is essentially a multi-dimensional decision vector, whose components include the long-term contract procurement quantity for each load area, the market coal procurement limit, the pre-positioned safety stock configuration quantity, and the reserved capacity for emergency transportation paths. This parameter must strictly meet the energy storage and transportation boundaries (i.e., storage quantity does not exceed the storage capacity, and transportation quantity does not exceed the flow limit value) and serves as the lower limit of the hard constraint on the second-stage action space.
[0097] In one example, the first-stage allocation sub-model calculates the performance gap between the current candidate allocation scheme and the ideal allocation scheme for each uncertainty scenario within a preset search space. This gap is quantified through a multi-objective regret value mechanism, covering three dimensions: the increase in total energy cost, the gap in demand satisfaction rate, and the delivery delay time. For details, please refer to operations S311 to S314.
[0098] In operation S320, energy demand attribute data and energy constraint attribute data are processed based on the second-stage allocation sub-model. Within the constraint range formed by the lower limit of the constraint determined by the benchmark allocation parameters, dynamic adjustment parameters are generated.
[0099] According to an embodiment of this application, the energy allocation prediction model includes a second-stage allocation sub-model. The second-stage allocation sub-model is constructed based on a multi-objective reinforcement learning framework. The multi-objective reinforcement learning framework is trained using vector reward signals and action space boundaries. The vector reward signals represent the comprehensive returns of multiple optimization objectives, including total energy cost, demand satisfaction rate, and delivery delay time. The action space boundaries are determined based on benchmark allocation parameters and are used to limit the parameter range of dynamically adjusted parameters to the constraint interval corresponding to the benchmark allocation parameters.
[0100] In one example, the energy allocation prediction model consists of a first-stage allocation sub-model and a second-stage allocation sub-model working together. This hybrid architecture of "pre-decision + online fine-tuning" balances robustness and agility. The second-stage allocation sub-model acts as an online adaptive engine, performing real-time optimization within the hard constraint range defined by the baseline allocation parameter x̅. Built on a multi-objective reinforcement learning framework, the second-stage allocation sub-model, as an online adaptive decision engine, dynamically optimizes operational efficiency in real-time within the safe range defined by the baseline parameters. Based on this framework, the sub-model constructs a state vector using real-time data such as current inventory levels, demand deviations, and transportation route status. It explicitly quantifies the combined rewards of cost, demand fulfillment rate, and delivery delay through vector reward signals, driving the multi-objective proximal policy optimization network to search for the optimal strategy within the action space defined by the constraint lower bounds determined by the baseline allocation parameters. The output dynamically adjusted parameters are strictly limited to the feasible region through a boundary pruning mechanism, enabling agile fine-tuning of the baseline solution. This avoids the out-of-bounds risk of pure reinforcement learning and addresses the overly conservative shortcomings of traditional robust optimization, achieving a dynamic balance between resilience and economy.
[0101] When operating S330, energy allocation information is generated based on baseline allocation parameters and dynamic adjustment parameters.
[0102] In one example, after obtaining the baseline allocation parameters and dynamic adjustment parameters, the system superimposes them with constraints to generate the final executable energy allocation information. Specifically, the procurement execution volume for each load area equals the sum of the minimum procurement volume determined by the baseline parameters and the emergency procurement volume in the dynamic adjustment parameters, with its upper limit constrained by the total monthly contract volume; the inventory target value equals the sum of the minimum safety stock set by the baseline parameters and the dynamic allocation volume, ensuring that the total inventory is always higher than the safety threshold; the transportation plan consists of the main path specified by the baseline parameters and the backup path switching decision triggered by the dynamic adjustment parameters, and the backup path switching needs to pass real-time availability verification. After the generated energy allocation information is confirmed by the dispatcher through a visual interface, it is automatically converted into purchase orders, allocation instructions, and transportation plans, and issued to the execution unit through the energy management system, completing the closed-loop decision-making process from data input to instruction output. This fusion mechanism achieves an organic unity of underlying robustness and top-level agility, enabling the decision results to have both high reliability in extreme scenarios and high economy in daily operations.
[0103] According to an embodiment of this application, the process of generating the baseline allocation parameters specifically includes operations S311 to S314.
[0104] In operation S311, the energy storage and transportation boundary characteristics and safety stock threshold constraints of the energy load area are determined based on the energy constraint attribute data.
[0105] According to embodiments of this application, energy storage and transportation boundary features are used to define the search space for baseline allocation parameters.
[0106] In one example, a feasible region for baseline allocation parameters is constructed using mathematical programming methods based on energy constraint attribute data. Specifically, for each load region r, its ending inventory level... A minimum safety stock threshold constraint must be met, which is dynamically calculated based on historical demand fluctuation statistics.
[0107]
[0108] in, This is the safety margin factor (usually taken as 1.5~2.0). Let r be the standard deviation of historical demand in region r. This allows for lead time for energy replenishment. This constraint ensures that inventory redundancy can cover more than 95% of demand uncertainty.
[0109] For each transport route l, introduce a binary variable for availability. Forced when the route is interrupted due to weather or maintenance Corresponding transportation volume decision variables It must be 0 to constitute a transportation feasibility constraint.
[0110] Supply capacity constraints limit all decision variables to the upper limits of supplier contract volume and physical throughput capacity of the channel, ultimately forming the search space. Let A be a matrix multiplied by a vector x to represent linear inequality constraints. In practical applications, A typically consists of the coefficients of multiple inequality constraints. Represents the matrix-vector multiplication operation; This represents a set of linear inequality constraints that vector x must satisfy, i.e. The result must be less than or equal to vector b. Vector b contains the constant term on the right side of each inequality constraint. This represents a lower bound constraint, ensuring that the solution will not fall below a certain threshold.
[0111] In operation S312, based on energy supply capacity data, energy procurement cost data, and historical operating statistics of each energy load area, a set of uncertainty scenarios associated with the energy load area is constructed.
[0112] According to embodiments of this application, uncertainty scenarios are used to characterize the dynamic influence relationship between the energy supply and demand balance, energy transportation cost, and energy security reserve status of the corresponding energy load area.
[0113] In one example, m high-dimensional related scenes are generated using a spatiotemporal fusion model. Each scenario consists of a three-dimensional dynamic coupling of demand disturbances, supply status, and transportation costs.
[0114] In operation S313, for each scenario in the set of uncertain scenarios, the energy allocation regret value of the current candidate allocation scheme is calculated in the search space.
[0115] According to an embodiment of this application, the energy allocation regret value indicates the performance gap between the current candidate allocation scheme and the ideal allocation scheme corresponding to the scenario in terms of total energy cost, demand fulfillment rate gap, and delivery delay time. The current candidate allocation scheme meets the safety stock threshold constraint. The conditional risk value is obtained by weighted averaging the energy allocation regret values of the tail extreme scenario subset. The regret value of the tail extreme scenario subset is greater than a preset threshold.
[0116] In one example, for each candidate benchmark parameter , To operate on the search space constructed by S311, calculate its position in the scene. Regret value for multi-objective energy allocation:
[0117]
[0118] in, For multi-objective performance vectors, Representation in a given decision and scene The total cost below Representation in a given decision and scene The reliability of the following system In order to make a given decision and scene The delay cost or delay time; For the scene If the ideal solution for the disturbance is known in advance (post-hoc optimal), the regret value quantifies the gap between the current solution and the ideal solution in terms of cost, reliability, and latency.
[0119] Subset of extreme tail scenarios Calculate the conditional value at risk (CVaR): .
[0120] in, To include all regret values The set of scenarios ω that are greater than or equal to a given risk value VaRa(x). Confidence level (For example Conditional value of risk (at 95%); The objective function is the current objective function, which can be cost, reliability, latency, or other objectives that need to be optimized. When calculating CVaR, the focus is on scenarios where the objective function exceeds a certain threshold (usually the high percentile of the distribution). The cumulative distribution function of f(x,ω) Quantiles, i.e., f(x,ω) greater than or equal to The probability is ; Indicates in The expected value is calculated under the following conditions. The optimization objective is set as follows:
[0121]
[0122] in, The risk threshold represents the maximum acceptable risk level, which is dynamically adjusted to meet certain conditions. , For safety margin. Value at condition (Va) represents the conditional risk value at a given confidence level β, which measures the objective function in the worst-case scenario. Expected value; Indicates all possible scenarios The maximum value of conditional risk cannot exceed the threshold. The optimization objective is to find a decision x that maximizes reliability while minimizing cost and delay, and ensures that the worst-case objective function value does not exceed a preset risk threshold τ in all possible scenarios, so that the baseline parameter x has the minimum average regret value in the foreseeable worst 5% scenario, thus ensuring that the solution has resilience under extreme operating conditions.
[0123] In operation S314, with the goal of minimizing conditional risk value, the Pareto optimal solution set is output as the benchmark assignment parameter through iterative evolution using the third-generation non-dominated sorting genetic algorithm.
[0124] In one example, a third-generation non-dominated sorting genetic algorithm is used to solve the above min-CVaR problem. Chromosomes are encoded as candidate baseline parameter vectors. Each gene corresponds to a path purchase quantity or regional inventory configuration. Fitness evaluation is performed in parallel to calculate the CVaR of all individuals in the population under m scenarios. a (x) employs a fast non-dominated sorting hierarchical approach. A set of reference points is uniformly generated in the cost-reliability-delay three-dimensional target space, and the population is guided to evolve uniformly towards the Pareto front using a niche-preserving operator. After a preset number of iterations, e.g., 200 iterations, 5-10 Pareto optimal solutions are output as the baseline allocation parameter x*. Each solution satisfies the lower constraint requirement, forming an inviolable hard boundary for the second stage.
[0125] Figure 4A flowchart of a method for generating dynamically adjusted parameters according to an embodiment of this application is shown. As shown, operation S320 includes operations S321 to S325.
[0126] In operation S321, the real-time state vector of each energy load region at the current scheduling moment is determined based on energy demand attribute data and energy constraint attribute data.
[0127] According to embodiments of this application, the real-time state vector includes the regional inventory level of each energy load area, the transportation route status, the real-time demand deviation, the cumulative cost value, and the cumulative delay value.
[0128] In one example, the system collects the real-time state vectors of each energy load region at the current scheduling time t. :
[0129]
[0130] in, , Let R represent the set of real numbers, and R be the total number of load areas. This represents an R-dimensional column vector of regional inventory levels. Vector elements... The inventory level of region v at time t; Characterizes the availability of transport path e at time t; It represents the dynamic demand for a specific category c at time t, where c is related to the application scenario. For example, in the energy distribution problem, c can represent coal, oil or natural gas. Characterizing the normalized scalar of cumulative cost, The cumulative delay penalty scalar is used to represent this state vector. This state vector is then standardized and input into the policy network to ensure comparability of data with different dimensions.
[0131] In operation S322, a vector reward signal is calculated based on the real-time state vector.
[0132] According to embodiments of this application, the vector reward signal includes a cost reward component obtained from the accumulated cost value, a demand satisfaction rate reward component obtained from the real-time demand deviation, and a timeliness reward component obtained from the accumulated delay value.
[0133] In one example, based on real-time state vectors The system calculates the vector reward signal. Among them, cost incentive portion , The additional procurement, transportation, and inventory costs within the current period Δt are divided by the baseline. Achieve normalization; reward weight for demand satisfaction rate , The actual energy demand met within time period t For the total demand, this formula maps the demand satisfaction rate to the interval [0,1]. When the satisfaction rate is ≥100%, the reward is 1; when the demand is not fully satisfied, the reward decreases linearly according to the actual proportion, directly incentivizing the agent to maximize its supply capacity. (Delayed penalty component) , This represents the amount of delayed deliveries added within time period t. The upper limit of the total periodic delay tolerance is represented by θ, with negative values penalizing delay behavior. θ is a sensitivity coefficient that can be dynamically adjusted according to the scheduling strategy. If there is no new delay in time period t, this component is 0. The three components combine to form a vector reward, and the policy network automatically learns the trade-offs between each objective through the advantage function.
[0134] In operation S323, the lower limit of the constraint is determined based on the baseline allocation parameters, and the action search space is constructed.
[0135] According to embodiments of this application, the lower bound constraints include a minimum inventory threshold, a minimum emergency procurement quantity, and a transportation route availability constraint.
[0136] In one example, action space , This indicates the emergency increase or decrease in procurement volume for the spot market based on the long-term contract procurement volume determined by the benchmark allocation parameters, expressed in tens of thousands of tons (coal). A positive value indicates that emergency procurement is initiated due to unexpected real-time demand or supply disruptions, while a negative value indicates that procurement is postponed to save costs when demand is weak. Its range is strictly limited by the hard constraints of the lower and upper limits of the emergency procurement volume set by the benchmark parameters in the first stage, ensuring that it will not exceed the monthly procurement capacity boundary.
[0137] Let L be an L-dimensional probability distribution vector (L being the total number of transport routes), and each element... ∈[0,1] represents the probability of selecting the l-th transport path for energy allocation at time t. When the path availability indicator in the real-time state vector... When [l]=0 (e.g., when a typhoon causes shipping disruptions), force = 0; For normally available paths, the policy network outputs a probability distribution through softmax, supporting the decision-making of the traffic allocation ratio between the main path and the backup path, and realizing a dynamic trade-off between transportation costs and reliability. Let R be the transfer vector, with elements This represents the net inflow or outflow to region r. The transfer must satisfy the resource conservation constraint. That is, total outflows equal total inflows, and regional inventory safety constraints are met. , The minimum safety stock threshold for region r is the minimum inventory level set to prevent stockouts.
[0138] In operation S324, the real-time state vector is input into the multi-objective proximal policy optimization network. The multi-objective proximal policy optimization network evaluates the multi-objective state value of candidate actions and calculates the advantage function based on the vector reward signal and the action search space, and outputs the candidate action vector that maximizes the vector reward signal.
[0139] According to embodiments of this application, candidate action vectors include emergency energy procurement volume, inventory allocation volume, and transportation route switching decisions.
[0140] In operation S325, the feasibility of candidate action vectors is verified in the action search space. Action components that exceed the constraint lower limit value are clipped to the boundary value, and dynamic adjustment parameters are generated.
[0141] In one example, the real-time status Input a multi-objective proximal policy optimization network, and the policy network outputs an action probability distribution. Value function networks evaluate multi-objective state values. The strategy parameters are updated through a dominance function pruning mechanism, with the core objective parameter being:
[0142]
[0143] in, Let be the expectation operator, representing the expectation over all possible states. Make expectations This represents the importance sampling ratio between the old and new strategies, and its physical meaning is that the current strategy and the old strategy are in the same state. Select action The probability difference is used to measure the importance weight of policy updates; For the dominant function, Let the action value function be... The difference between the two values represents the action value function, evaluating the action. The degree of superiority or inferiority relative to the average level of the state; To reduce the threshold and ensure that the policy update range does not exceed ±20%, we prevent the policy from fluctuating violently due to real-time data noise, and ensure the training stability and online decision smoothness in high-risk scenarios such as coal dispatching.
[0144] To avoid gradient conflicts among multiple objectives and improve training stability, the multi-objective proximal policy optimization network in this application adopts a separate objective attention mechanism. Each objective is encoded independently and then fused. This mechanism utilizes objective-specific attention weights (cost weights W). cost Reliability weight W reli Delay weight W delayThis approach achieves independent encoding of each optimization objective, avoiding the conflict between cost minimization and reliability maximization gradients that cancel each other out in traditional single-network structures. Finally, feature concatenation is performed at the output layer, preserving the independent representation of each objective while allowing the policy network to learn the trade-offs between objectives. This improves the model's convergence speed by over 60% in high-dimensional state spaces such as coal scheduling, and results in a more uniform distribution of the generated Pareto policy set. The final policy output layer is as follows:
[0145]
[0146] Here, ⊕ represents the vector concatenation operation, which connects the low-dimensional attention representations of the three targets into a fusion vector along the feature dimension. These are the encoded vectors for cost, reliability, and delayed delivery, respectively; Ua is the output layer weight matrix, which maps the fused vector to the log probabilities of the action space; the Softmax function generates a normalized action probability distribution. This is used for strategy sampling and execution.
[0147] The multi-objective proximal policy optimization network limits the policy update range through the aforementioned pruning mechanism. When operating the S325, it still needs to optimize the output candidate actions. Hard boundary checks are performed to ensure that the data strictly falls within the feasible region defined by the baseline allocation parameter x*. Specifically, the check process is conducted independently for each component: For the emergency procurement quantity component, the system checks whether it exceeds the emergency procurement upper limit set by the baseline parameter; if it does, it is forcibly reduced to the upper limit. For the inventory transfer component, the system verifies whether the inventory level in each region is still higher than the safety stock threshold after the transfer; if the transfer quantity is too large and causes the inventory in any region to fall below the safety line, it is reduced to the maximum allowable value that meets safety requirements. For the transportation route switching decision component, the system compares the current route availability status in real time; if a route is unavailable due to weather or congestion, the route switching decision is forcibly invalidated, and selection is only allowed among available routes. After all components have completed boundary checks, the pruned action vector is the final dynamically adjusted parameter. This parameter retains the optimal response of the policy network to the real-time state and strictly follows the safety boundary defined by the first-stage pre-decision; it can be directly issued to the energy procurement and logistics execution system, achieving a closed-loop integration of bottom-level resilience and top-level agility.
[0148] According to an embodiment of this application, the method further includes: determining the scenario feature vector of each energy load region at the current scheduling time based on energy demand attribute data and energy constraint attribute data, wherein the scenario feature vector includes the regional inventory level, transportation path status and real-time demand deviation of each energy load region; determining the target scenario feature vector with the highest matching degree in the uncertain scenario set based on the scenario feature vector at the current scheduling time; and determining the reinforcement learning policy parameters corresponding to the target scenario represented by the target scenario feature vector through a pre-trained scenario policy library to initialize a multi-objective proximal policy optimization network, wherein the scenario policy library is used to represent the mapping relationship between scenarios and policy parameters.
[0149] In one example, to achieve millisecond-level warm-start and experience transfer of reinforcement learning policies and significantly improve emergency response capabilities under sudden conditions, a pre-trained scenario policy library stores the mapping relationship between scenarios and policy parameters. This mapping mechanism performs weighted Mahalanobis distance matching between the current real-time scenario and the Pareto optimal baseline parameter solution set generated in the offline stage, directly loading the optimal policy parameters verified in historical similar scenarios. This avoids the tens of thousands of trial-and-error iterations required for cold start in traditional reinforcement learning, enabling the policy network to achieve suboptimal performance in the first decision cycle. Specifically, the scenario policy library stores the mapping relationship between scenarios and policy parameters as shown in the following equation:
[0150]
[0151] in This is the Pareto solution set for the first stage. Let be the scene covariance matrix. The function is a mapping function. Its input is the currently observed real-time scene feature vector ω (such as current inventory, path availability, and demand deviation), and its output is the optimally initialized policy parameters θ. The representation involves traversing the Pareto optimal baseline parameter solution set P output in the first stage to find the solution that minimizes the objective function value. . The representation uses Mahalanobis distance to measure the difference between the current scene ω and the reference scene corresponding to the baseline solution. The similarity between them.
[0152] In one example, at each scheduling decision time t, the system constructs a feature vector for the current scenario based on energy demand attribute data and energy constraint attribute data. This vector is a compact three-dimensional representation, distinct from the complete state vector, and focuses on depicting the core patterns of the scene. This is a regional inventory level vector, reflecting the adequacy of reserves. The transport path state vector identifies the connectivity of the logistics network. This vector represents the real-time demand deviation, indicating the degree of supply-demand imbalance. It is obtained through principal component analysis for dimensionality reduction, retaining over 95% of the scene variation information and reducing the computational complexity of matching. The current scene feature vector is then used... With a pre-generated set of uncertain scenarios The feature vectors of each scene are compared for similarity, using cosine similarity as a metric. After traversing the scene set, the target scene with the highest similarity is selected. Based on a pre-trained scene policy library, the policy parameters corresponding to the target scene are determined. Since the current working conditions may still differ, after executing an action at time t, the system collects actual state transition samples, stores them in an experience replay pool, and fine-tunes the network parameters online. The decision interface visualizes the Pareto front. After the user selects the preference weight w, the second-stage MORL objective is converted into a scalarized reward. It generates adjustment plans in real time and supports manual correction commands. Update action execution:
[0153]
[0154] in, This represents the action actually performed at time t. It is a parameter that intervenes between 0 and 1, used to control the trade-off between the policy network's recommended actions and human intervention.
[0155] By constructing a scenario-policy mapping function, millisecond-level hot-start and experience transfer of reinforcement learning policies were achieved, significantly improving emergency response capabilities under sudden conditions. This mapping mechanism performs weighted Mahalanobis distance matching between the current real-time scenario and the Pareto optimal baseline parameter solution set generated in the offline stage, directly loading the optimal policy parameters verified in similar historical scenarios. This avoids the tens of thousands of trial-and-error iterations required for cold start in traditional reinforcement learning, enabling the policy network to achieve suboptimal performance in the first decision cycle. Simultaneously, since the loaded policy has been validated through offline robust optimization, it satisfies both hard constraint boundary guarantees and event memory to cope with similar disturbances, significantly enhancing the policy's generalization ability and convergence stability under new conditions. Ultimately, this reduces the decision response latency of the energy distribution system from hours to minutes, and improves the supply satisfaction rate by more than 15% in extreme scenarios.
[0156] Based on the above-mentioned energy allocation optimization method, this application also provides an energy allocation optimization device. The following will be combined with... Figure 5 The device is described in detail.
[0157] Figure 5 A schematic block diagram of an energy distribution optimization device according to an embodiment of this application is shown.
[0158] like Figure 5As shown, the energy allocation optimization device 500 of this embodiment includes an acquisition module 510, an energy allocation prediction module 520, and an allocation scheme determination module 530.
[0159] The acquisition module 510 is used to acquire energy procurement cost data, energy supply attribute data, and energy demand attribute data and energy constraint attribute data associated with each energy load area during the target scheduling period. The energy supply attribute data characterizes the energy supply quantity and stability characteristics of each energy load area, the energy demand attribute data characterizes the demand fluctuation and demand balance characteristics of the energy load area, and the energy constraint attribute data characterizes the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load area. In one embodiment, the acquisition module 510 can be used to perform the operation S210 described above, which will not be repeated here.
[0160] The energy allocation prediction module 520 is used to process energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data based on the energy allocation prediction model to obtain energy allocation information. The energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraint conditions. In one embodiment, the energy allocation prediction module 520 can be used to perform the operation S220 described above, which will not be repeated here.
[0161] The allocation scheme determination module 530 is used to determine the target energy allocation scheme corresponding to the energy load area based on the energy allocation information. The target energy allocation scheme is used for energy allocation between the energy supplier and each energy load area. In one embodiment, the allocation scheme determination module 530 can be used to perform the operation S230 described above, which will not be repeated here.
[0162] According to an embodiment of this application, the energy distribution prediction module 520 includes a first generation submodule, a second generation submodule, and a third generation submodule.
[0163] The first generation submodule processes energy supply capacity data, energy procurement cost data, and energy constraint attribute data based on the first-stage allocation submodule to generate baseline allocation parameters. These baseline allocation parameters satisfy the energy storage and transportation boundaries represented by the energy constraint attribute data. The second generation submodule processes energy demand attribute data and energy constraint attribute data based on the second-stage allocation submodule. Within the constraint interval defined by the lower limit values determined by the baseline allocation parameters, it generates dynamic adjustment parameters. The second-stage allocation submodule is constructed based on a multi-objective reinforcement learning framework, which is trained using vector reward signals and action space boundaries. The vector reward signals represent the comprehensive reward of multiple optimization objectives, including total energy cost, demand satisfaction rate, and delivery delay time. The action space boundaries are determined based on the baseline allocation parameters and are used to limit the range of the dynamic adjustment parameters within the constraint interval corresponding to the baseline allocation parameters. Finally, the third generation submodule generates energy allocation information based on the baseline allocation parameters and the dynamic adjustment parameters.
[0164] According to an embodiment of this application, the first generation submodule includes a first determination unit, a scene construction unit, a first calculation unit, a second calculation unit, and an output unit.
[0165] The first determining unit is used to determine the energy storage and transportation boundary characteristics and safety stock threshold constraints of the energy load area based on energy constraint attribute data. The energy storage and transportation boundary characteristics are used to limit the search space of the benchmark allocation parameters. The scenario construction unit is used to construct a set of uncertain scenarios associated with the energy load area based on energy supply capacity data, energy procurement cost data, and historical operation statistics of each energy load area. The uncertain scenarios are used to characterize the dynamic influence relationship between the energy supply and demand balance state, energy transportation cost state, and energy safety reserve state of the corresponding energy load area. The first calculation unit is used to calculate each of the uncertain scenario sets. The first unit calculates the energy allocation regret value of the current candidate allocation scheme within the search space. The energy allocation regret value indicates the performance gap between the current candidate allocation scheme and the ideal allocation scheme corresponding to the scenario in terms of total energy cost, demand fulfillment rate gap, and delivery delay time. The current candidate allocation scheme meets the safety stock threshold constraint. The second calculation unit calculates the conditional risk value by weighting the energy allocation regret values of the tail extreme scenario subset. The regret value of the tail extreme scenario subset is greater than a preset threshold. The output unit is used to output the Pareto optimal solution set as the benchmark allocation parameter by iteratively evolving through the third-generation non-dominated sorting genetic algorithm with the goal of minimizing the conditional risk value.
[0166] According to embodiments of this application, the scenario construction unit is further configured to use a prediction model constructed by combining a long short-term memory neural network and a prophetic prediction model to process historical and real-time demand data for the corresponding energy load area, thereby obtaining predicted demand disturbance values. The long short-term memory neural network is used to capture periodic demand trends, while the prophetic prediction model characterizes the sudden impact of extreme weather on demand using a Poisson impulse response function. A Markov state transition process is used to simulate the state changes of the supply path in each energy load area, obtaining the supply state transition probability, where the state transition probability exhibits an exponential decay characteristic as the time interval increases. Finally, a time-varying transportation cost disturbance model is used to process transportation base cost data and fuel price fluctuation data to determine the transportation cost disturbance vector.
[0167] According to an embodiment of this application, the second generation submodule further includes a second determining unit, a third determining unit, a search unit, and a verification unit.
[0168] The second determining unit is used to determine the real-time state vector of each energy load region at the current scheduling time based on energy demand attribute data and energy constraint attribute data. The real-time state vector includes the regional inventory level, transportation path status, real-time demand deviation, cumulative cost value, and cumulative delay value of each energy load region. The third determining unit calculates a vector reward signal based on the real-time state vector. The vector reward signal includes a cost reward component obtained from the cumulative cost value, a demand satisfaction rate reward component obtained from the real-time demand deviation, and a timeliness reward component obtained from the cumulative delay value. The search unit is used to determine the constraint lower limit value based on the benchmark allocation parameters and construct the action search space. The lower bound constraints include minimum inventory threshold, minimum emergency procurement quantity, and transportation route availability constraints. A multi-objective proximal policy optimization network is used to input the real-time state vector into the network. Based on the vector reward signal and the action search space, the network evaluates the multi-objective state value of candidate actions and calculates the advantage function, outputting the candidate action vector that maximizes the vector reward signal. These candidate action vectors include emergency energy procurement quantity, inventory allocation quantity, and transportation route switching decisions. A verification unit is used to perform feasibility verification on the candidate action vectors within the action search space, pruning action components exceeding the constraint lower bound boundaries to the boundary values, and generating dynamically adjusted parameters.
[0169] According to embodiments of this application, the apparatus further includes: a scene determination module, a scene matching module, and a scene strategy parameter mapping module.
[0170] The scenario determination module is used to determine the scenario feature vector of each energy load area at the current scheduling time based on energy demand attribute data and energy constraint attribute data. The scenario feature vector includes the regional inventory level, transportation path status and real-time demand deviation of each energy load area. The scenario matching module is used to determine the target scenario feature vector with the highest matching degree in the uncertain scenario set based on the scenario feature vector at the current scheduling time. The scenario policy parameter mapping module is used to determine the reinforcement learning policy parameters corresponding to the target scenario represented by the target scenario feature vector through a pre-trained scenario policy library, so as to initialize the multi-objective proximal policy optimization network. The scenario policy library is used to represent the mapping relationship between scenarios and policy parameters.
[0171] According to embodiments of this application, any plurality of modules in the acquisition module 510, energy allocation prediction module 520, and allocation scheme determination module 530 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the acquisition module 510, energy allocation prediction module 520, and allocation scheme determination module 530 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 510, energy allocation prediction module 520, and allocation scheme determination module 530 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0172] Figure 6 A block diagram of an electronic device suitable for implementing an energy distribution optimization method according to an embodiment of this application is shown schematically.
[0173] like Figure 6As shown, an electronic device 600 according to an embodiment of this application includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0174] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 602 and / or RAM 603. It should be noted that programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.
[0175] According to embodiments of this application, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.
[0176] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0177] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.
[0178] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the energy distribution optimization method provided in the embodiments of this application.
[0179] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0180] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 609, and / or installed from the removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0181] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0182] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0183] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0184] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0185] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. An energy allocation optimization method, characterized in that, The method includes: The system acquires energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data associated with each energy load region during the target scheduling cycle. The energy supply attribute data characterizes the energy supply quantity and supply stability characteristics of each energy load region, the energy demand attribute data characterizes the demand fluctuation and demand balance characteristics of the energy load region, and the energy constraint attribute data characterizes the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load region. Energy allocation prediction model processes the energy procurement cost data, energy supply attribute data, energy demand attribute data, and energy constraint attribute data to obtain energy allocation information. The energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraint conditions. Based on the energy allocation information, a target energy allocation scheme corresponding to the energy load area is determined. The target energy allocation scheme is used for energy allocation between the energy supplier and each energy load area.
2. The method according to claim 1, characterized in that, The energy allocation prediction model includes a second-stage allocation sub-model. Based on the energy allocation prediction model, it processes the energy procurement cost, energy supply attributes, energy demand attributes, and energy constraint attributes to obtain energy allocation information including: Based on the first-stage allocation sub-model, the energy supply capacity data, the energy procurement cost data, and the energy constraint attribute data are processed to generate benchmark allocation parameters. The benchmark allocation parameters satisfy the energy storage capacity boundary and the energy transportation capacity boundary represented by the energy constraint attribute data. Based on the second-stage allocation sub-model, the energy demand attribute data and the energy constraint attribute data are processed. Within the constraint interval formed by the constraint lower limit determined by the baseline allocation parameters, dynamic adjustment parameters are generated. The second-stage allocation sub-model is constructed based on a multi-objective reinforcement learning framework, which is trained using vector reward signals and action space boundaries. The vector reward signals represent the comprehensive reward of multiple optimization objectives, including total energy cost, demand satisfaction rate, and delivery delay time. The action space boundaries are determined based on the baseline allocation parameters and are used to limit the parameter range of the dynamic adjustment parameters within the constraint interval corresponding to the baseline allocation parameters. The energy allocation information is generated based on the baseline allocation parameters and the dynamic adjustment parameters.
3. The method according to claim 2, characterized in that, Based on the first-stage allocation sub-model, the energy supply capacity data, the energy procurement cost data, and the energy constraint attribute data are processed to generate baseline allocation parameters, including: The energy storage and transportation boundary characteristics and safety stock threshold constraints of the energy load area are determined based on the energy constraint attribute data. The energy storage and transportation boundary characteristics are used to limit the search space of the benchmark allocation parameters. Based on the energy supply capacity data, the energy procurement cost data, and the historical operating statistics of each energy load region, a set of uncertainty scenarios associated with the energy load region is constructed. The uncertainty scenarios are used to characterize the dynamic influence relationship between the energy supply and demand balance, energy transportation cost, and energy security reserve status of the corresponding energy load region. For each scenario in the set of uncertain scenarios, the energy allocation regret value of the current candidate allocation scheme is calculated in the search space. The energy allocation regret value indicates the performance gap between the current candidate allocation scheme and the ideal allocation scheme corresponding to the scenario in terms of total energy cost, demand satisfaction gap and delivery delay time. The current candidate allocation scheme satisfies the safety stock threshold constraint. The conditional risk value is obtained by weighted averaging the energy allocation regret values of the tail extreme scenario subset, where the regret value of the tail extreme scenario subset is greater than a preset threshold. With the goal of minimizing the conditional risk value, the Pareto optimal solution set is output as the baseline assignment parameter through iterative evolution using a third-generation non-dominated sorting genetic algorithm.
4. The method according to claim 3, characterized in that, Constructing a set of uncertainty scenarios associated with the energy load region includes: A prediction model constructed by combining a long short-term memory neural network and a prophetic prediction model processes historical and real-time demand data for a corresponding energy load region to obtain predicted demand disturbance values. The long short-term memory neural network is used to capture periodic demand trends, and the prophetic prediction model uses a Poisson impulse response function to characterize the sudden impact of extreme weather on demand. A Markov state transition process is used to simulate the state changes of the supply path in each energy load region, yielding the supply state transition probability, wherein the state transition probability exhibits an exponential decay characteristic with increasing time interval; and The transportation cost disturbance vector is determined by processing basic transportation cost data and fuel price fluctuation data using a time-varying transportation cost disturbance model.
5. The method according to claim 2, characterized in that, Based on the second-stage allocation sub-model, the energy demand attribute data and the energy constraint attribute data are processed. Within the constraint interval formed by the constraint lower limit value determined by the baseline allocation parameters, dynamic adjustment parameters are generated, including: Based on the energy demand attribute data and the energy constraint attribute data, the real-time state vector of each energy load region at the current scheduling time is determined. The real-time state vector includes the regional inventory level, transportation route status, real-time demand deviation, cumulative cost value and cumulative delay value of each energy load region. A vector reward signal is calculated based on the real-time state vector. The vector reward signal includes a cost reward component obtained from the accumulated cost value, a demand satisfaction rate reward component obtained from the real-time demand deviation, and a timeliness reward component obtained from the accumulated delay value. Based on the baseline allocation parameters, a lower limit of constraints is determined, and an action search space is constructed. The lower limit of constraints includes a minimum inventory threshold, an emergency procurement quantity lower limit, and a transportation route availability constraint. The real-time state vector is input into a multi-objective proximal policy optimization network. Based on the vector reward signal and the action search space, the multi-objective proximal policy optimization network evaluates the multi-objective state value of candidate actions and calculates the advantage function, and outputs a candidate action vector that maximizes the vector reward signal. The candidate action vector includes emergency energy procurement quantity, inventory allocation quantity and transportation route switching decision. Within the action search space, the feasibility of the candidate action vectors is verified, and action components exceeding the constraint lower limit value are clipped to the boundary value to generate dynamic adjustment parameters.
6. The method according to claim 5, characterized in that, The method further includes: Based on the energy demand attribute data and the energy constraint attribute data, a scenario feature vector is determined for each energy load region at the current scheduling time. The scenario feature vector includes the regional inventory level, transportation path status, and real-time demand deviation of each energy load region. Based on the scene feature vector at the current scheduling moment, determine the target scene feature vector with the highest matching degree in the set of uncertain scenes; The reinforcement learning policy parameters corresponding to the target scene represented by the target scene feature vector are determined by a pre-trained scene policy library to initialize a multi-objective proximal policy optimization network. The scene policy library is used to represent the mapping relationship between the scene and the policy parameters.
7. The method according to any one of claims 1 to 6, characterized in that, The first multi-objective optimization function represents minimizing the total energy cost, the demand fulfillment gap, and the delivery delay time as the optimization objectives, and the first set of constraints includes at least one of the following: Supply capacity constraints characterize the limiting relationship between energy supply and the upper limit of supply capacity; inventory safety constraints characterize the balance between inventory level and safety stock threshold; and transportation feasibility constraints characterize the limiting relationship between transportation route selection status and route availability.
8. An energy distribution optimization device, characterized in that, The device includes: The acquisition module is used to acquire energy procurement cost data, energy supply attribute data, energy demand attribute data and energy constraint attribute data associated with each energy load area during the target scheduling period. The energy supply attribute data represents the energy supply quantity characteristics and supply stability characteristics of each energy load area, the energy demand attribute data represents the demand fluctuation characteristics and demand balance characteristics of the energy load area, and the energy constraint attribute data represents the energy storage and transportation boundary characteristics and allocation priority characteristics of the energy load area. The energy allocation prediction module is used to process the energy procurement cost data, the energy supply attribute data, the energy demand attribute data, and the energy constraint attribute data based on the energy allocation prediction model to obtain energy allocation information. The energy allocation prediction model includes a first-stage allocation sub-model constructed based on a multi-objective optimization function and a set of constraint conditions. The allocation scheme determination module is used to determine the target energy allocation scheme corresponding to the energy load area based on the energy allocation information. The target energy allocation scheme is used for energy allocation between the energy supplier and each energy load area.
9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Flexible optimization scheduling method considering wind-solar and load uncertainty for comprehensive energy system
CN111815025A
Park integrated energy system planning and operation optimization method and terminal
CN115936220A
Comprehensive energy system dynamic scheduling method based on multi-agent deep reinforcement learning
CN119904057A
Energy storage configuration optimization method
CN120930872A
Large model energy consumption optimization method and device, computer equipment, readable storage medium and program product
CN121116587A
Cited By
User-side-oriented multi-target decision execution method and system
CN122022400A