Power grid emergency material intelligent supply chain dynamic allocation and priority ranking method based on reinforcement learning

By constructing a digital twin environment and reinforcement learning intelligent agents, the problems of flexibility and global optimization in the traditional power grid emergency material allocation method when facing complex disasters have been solved. The material allocation plan can be generated and optimized in real time, dynamically and intelligently, thereby improving the efficiency of emergency response and resource utilization.

CN121562891APending Publication Date: 2026-02-24STATE GRID LIAONING ELECTRIC POWER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511661409.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Traditional power grid emergency material allocation methods are ill-suited to adapting to the dynamic evolution of disaster situations when faced with large-scale, complex and ever-changing disaster scenarios. They lack flexibility and overall optimization capabilities, resulting in unreasonable resource allocation and an inability to quickly respond to the most urgent needs.

Method used

A digital twin environment is constructed, and a reinforcement learning agent is used for resource allocation decision-making. Through multi-source data fusion and dynamic assessment of urgency, the resource allocation plan can be generated and optimized in real time, dynamically, and intelligently. The reinforcement learning agent is used for training and decision-making, and the allocation strategy is optimized by combining multi-objective reward functions and collaborative decision-making mechanisms.

Benefits of technology

It significantly improves the efficiency and effectiveness of emergency response, ensures that limited resources are prioritized for key nodes, maximizes resource utilization efficiency, avoids resource misallocation, has the ability to learn and adapt to complex environments, and enhances the intelligence level of emergency management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562891A_ABST
    Figure CN121562891A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid emergency material intelligent supply chain dynamic allocation and priority ranking method based on reinforcement learning, and relates to the technical field of power grid emergency management and intelligent supply chains, in particular to a power grid emergency material intelligent supply chain dynamic allocation and priority ranking method based on reinforcement learning. The method comprises the following steps: constructing a digital twin environment model integrating geographic information, real-time traffic, inventory state and fault prediction; a reinforcement learning framework is defined, the state space covers inventory, in-transit materials, demand urgency and traffic states, and the action space comprises a material allocation instruction and task priority ranking, reward function comprehensive consideration shortage relief, delivery time efficiency and cost; and a strategy network capable of outputting an optimal dynamic allocation strategy is obtained by training the intelligent agent. According to the invention, intelligent dynamic distribution and priority optimization of emergency materials are realized, and the emergency response efficiency and the resource utilization rate of a power grid are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of power grid emergency management and smart supply chain technology, specifically a method for dynamic allocation and priority ranking of smart supply chains for power grid emergency materials based on reinforcement learning. Background Technology

[0002] As a critical national infrastructure, the safe and stable operation of the power grid is crucial to the national economy, people's livelihood, and national security. However, natural disasters (such as typhoons, snowstorms, and earthquakes) or sudden equipment failures often impact the power grid, leading to widespread power outages. In such emergency scenarios, the rapid and accurate allocation of necessary emergency supplies (such as transformers, circuit breakers, and cables) from warehouses to damaged power grid points is a key link in quickly restoring power supply, minimizing losses, and is essential for ensuring the normal operation of society.

[0003] Traditional power grid emergency material allocation largely relies on pre-established static plans and the experience-based decisions of dispatchers. This model reveals many limitations when dealing with large-scale, complex, and ever-changing disaster situations: on the one hand, static plans are difficult to adapt to the dynamic evolution of disasters (such as secondary disasters, traffic disruptions, sudden changes in demand), lacking flexibility; on the other hand, human decision-making is highly dependent on personal experience, and in the face of information overload and multi-source data (inventory, road conditions, multi-point and multi-type demands), it is difficult to quickly perform global optimization calculations, which can easily lead to unreasonable resource allocation, such as critical materials not being prioritized for the most urgent disaster-stricken areas, or delays caused by poor transportation route selection.

[0004] In recent years, with the development of IoT, big data, and AI technologies, the concept of smart supply chains has been introduced into the field of emergency management. Some studies have attempted to use optimization algorithms (such as linear programming and genetic algorithms) to solve resource allocation schemes, but these methods usually require complete and accurate mathematical models, have limited adaptability to complex dynamic environments, struggle to handle real-time changing uncertain information, and cannot learn and continuously improve through historical experience.

[0005] Therefore, there is an urgent need for an intelligent allocation method that can deeply integrate real-time data, dynamically perceive environmental changes, and possess autonomous learning and decision-making capabilities. Reinforcement learning, as an important branch of artificial intelligence, is well-suited for solving sequential decision-making problems due to its ability to learn optimal decision-making strategies through continuous interaction between the agent and the environment. Combining reinforcement learning with digital twin technology, which can accurately map the physical world, to build a data-driven, online-learning, and dynamically optimized intelligent supply chain decision-making system represents a promising development direction for improving the efficiency and intelligence level of power grid emergency material allocation. Summary of the Invention

[0006] The purpose of this invention is to provide a method for dynamic allocation and priority ranking of smart supply chains for power grid emergency materials based on reinforcement learning. By constructing a digital twin environment and utilizing reinforcement learning agents to learn and make decisions within it, the method can achieve real-time, dynamic, and intelligent generation and optimization of material allocation plans, significantly improving the efficiency and effectiveness of emergency response.

[0007] To achieve the above objectives, this invention employs the following technical solution: a dynamic allocation and priority ranking method for the intelligent supply chain of power grid emergency supplies based on reinforcement learning. Dynamic allocation and priority ranking of the power grid emergency supplies supply chain is a crucial aspect of power system disaster prevention and mitigation. This method constructs a digital twin environment model, providing a highly simulated virtual battlefield for decision-making. The core of this model lies in its multi-source data fusion capability. It is not simply a display of geographical information, but rather a deep coupling of static warehouse networks, dynamic traffic flow, fluctuating inventory levels, and power grid fault prediction based on multiple factors such as weather and equipment status. For example, after a typhoon warning is issued, the model can simulate the impact of wind and rainfall along different paths, predict potential tower collapses and line breaks, and dynamically deduce changes in road capacity leading to these points. The reinforcement learning agent is trained in this environment. Each decision it makes, such as "allocating 5 transformers from warehouse A to fault point X," triggers a complete logistics simulation process in the twin environment, including loading, route planning, in-transit transportation, and delivery and installation. This closed-loop simulation enables agents to perceive the chain reactions resulting from decisions, such as the impact of a warehouse's depletion of supplies on subsequent rescue efforts, or the delays caused by traffic congestion when choosing a certain route. This allows them to learn how to make globally optimal decisions under complex constraints, rather than responding to individual needs in isolation.

[0008] Furthermore, the dynamic assessment of the urgency level of material needs is the cornerstone of this method's precise allocation. Its dynamism lies in the fact that it is not a one-time static judgment, but a continuously updated process as the disaster evolves. The assessment model comprehensively considers factors across multiple dimensions: equipment failures at high voltage levels have a wide impact and are typically assigned a higher basic urgency; the criticality of equipment in the power grid topology, such as equipment located in key substations or important transmission lines, whose failures may lead to grid disconnection or large-scale power outages, further increases the urgency; the importance of affected users, such as lifeline users like hospitals, command centers, and transportation hubs, have the highest priority for power restoration; in addition, the estimated duration of the power outage is also a key variable—the longer the expected outage, the greater the impact on social operations and people's lives, and the higher the urgency. These factors are comprehensively calculated through a weighted model, and the weight coefficients can be set and adjusted based on historical emergency cases and expert knowledge. For example, in extreme events involving livelihood security, the weight of "affected users" may be increased. The model runs in real time. When the estimated repair time of a certain fault point is extended due to waiting for materials, or when its impact range expands, its urgency score will increase dynamically, which may trigger the system to reassess and allocate priorities to ensure that limited resources are always invested in the most critical parts.

[0009] Furthermore, the "priority ranking" in the action space is a dedicated optimization module, the challenge of which lies in how to scientifically prioritize a large number of concurrent allocation tasks. This ranking network receives a feature vector for each task, which includes not only the urgency score of the task's objective point but also the nature of the task itself, such as the type and quantity of required materials, the distance to the originating warehouse, and currently available transportation tools. The goal of the ranking network is to generate a task execution sequence that satisfies multiple constraints: the primary objective is to ensure that tasks with high urgency are prioritized for rapid acquisition of materials and transportation capacity; however, at the same time, a "fairness" or "anti-starvation" mechanism must be introduced to prevent tasks with temporarily lower urgency but equally important status from being indefinitely delayed due to continuous resource hoarding, thus preventing secondary disasters from escalating into major incidents. For example, while prioritizing the protection of core hub substations, resources also need to be allocated to a number of distribution transformer failures affecting residential areas to prevent prolonged, large-scale power outages that could cause social problems. The sorting network optimizes algorithms to prioritize high-urgency tasks while rationally arranging the execution order of low-priority tasks. It may also break down large tasks and coordinate across warehouses to maximize the overall emergency response efficiency of the system and ensure the orderly and efficient operation of rescue operations.

[0010] Furthermore, the reward function acts as a "command stick" guiding the behavior of the reinforcement learning agent, and its multi-objective weighted design directly determines the tendency of the learning strategy. In its design, it clearly distinguishes the value of different outcomes. Successfully delivering supplies to disaster-stricken areas with high urgency results in significant positive rewards, with the reward value increasing as urgency increases. This strongly incentivizes the agent to prioritize meeting the most urgent needs. Conversely, for situations where improper decision-making leads to delayed delivery or poor inventory planning renders the task impossible, high negative rewards are imposed, making the agent acutely aware of the cost of mistakes. Simultaneously, the reward function also considers economic objectives, appropriately penalizing the total cost of the transportation route to prevent the agent from pursuing speed at the expense of cost, guiding it to find the optimal balance between "emergency relief" and "economic conservation." This design ensures that the learned strategy is neither recklessly aggressive regardless of cost nor excessively conservative and inefficient, but rather a rational and intelligent decision-making process that balances emergency effectiveness and operational costs under extreme pressure.

[0011] Furthermore, collaborative decision-making mechanisms are an important supplement to scenarios involving cross-regional and multi-warehouse joint supply under large-scale complex disasters. When the backbone solution output by the intelligent agent requires collaboration among multiple warehouses, direct execution may face challenges at the distributed execution level, such as delays in updating inventory information in each warehouse and conflicts in local transportation capacity scheduling. The virtual "collaborative decision-making module" simulates the consultation and coordination process between warehouse management units in reality. This module, based on the intelligent agent's macro-level solution and combined with the real-time local information of each warehouse (such as accurate inventory balance, loading and unloading capacity, available vehicles, etc.), performs feasibility verification and fine-tuning. For example, the intelligent agent may instruct warehouse A to supply 10 pieces of equipment to a certain point, but the collaborative module finds that warehouse A only has 8 pieces available for immediate dispatch. Therefore, it will coordinate with nearby warehouse B to supplement the supply of 2 more pieces and adjust the transportation route to form the optimal combination. The fine-tuned solution is more realistic, improving the feasibility and efficiency at the execution level. More importantly, this fine-tuning process and its results will be fed back to the agent as new learning samples, enabling it to gradually learn how to generate more collaborative and easily distributed initial schemes in subsequent training, thereby achieving co-evolution between the agent and the execution environment.

[0012] Furthermore, the choice of algorithm and network structure for reinforcement learning agents directly impacts their perception and decision-making capabilities. Algorithms such as proximal policy optimization or deep deterministic policy gradient are suitable for handling continuous state and action space problems in this approach. The agent's network structure is specifically designed to integrate multimodal information. Convolutional neural network branches excel at processing data with spatial topological structures, such as power grid geographical distribution maps and location relationship maps between warehouses and fault points, from which spatial features can be extracted. Recurrent neural network branches are used to process sequential time-dependent information, such as changes in material demand over a series of consecutive time steps, the movement trajectory of materials in transit, and the temporal evolution of traffic conditions, thereby understanding the dynamic development trend of disasters. By combining the spatial feature extraction capabilities of CNNs with the time-series modeling capabilities of RNNs, the agent can simultaneously grasp "what happened where" and "how the situation will develop," thus achieving a profound understanding of the environmental state in the spatiotemporal dimension, laying a solid foundation for making forward-looking dynamic allocation and prioritization decisions.

[0013] Furthermore, the system achieves closed-loop intelligent management through a clear three-layer architecture. The environmental perception and modeling layer is the foundation of the system. It is responsible for accessing and integrating heterogeneous data from multiple sources such as SCADA, GIS, traffic management departments, and warehouse WMS, performing cleaning, alignment, and correlation to build and drive the real-time updating of the digital twin environment model, providing a unified, accurate, and timely data foundation for upper-level decision-making. The reinforcement learning intelligent decision-making layer is the "brain" of the system, equipped with a policy network trained extensively both offline and online. This layer receives real-time state snapshots from the perception layer and uses the policy network for calculations, outputting specific material allocation instructions and task priority ranking lists. The solution execution and feedback layer is the "hands and feet" of the system. It translates the instructions from the decision-making layer into actual operations, issuing them to the warehouse management system for picking and outbound processing, and to the logistics system for vehicle and route scheduling. Simultaneously, this layer closely tracks the actual effects of instruction execution, such as whether materials are delivered on time and the effectiveness after installation, and sends this feedback data (including success, delay, and failure information) back to the decision-making layer. Decision-makers use this real-world feedback data to fine-tune and adaptively optimize the strategy model online, enabling it to adapt to the ever-changing external environment and continuously improve decision-making capabilities.

[0014] Furthermore, the simulation and contingency plan evaluation modules significantly enhance the system's foresight and contingency plan preparation capabilities. Before an actual disaster occurs, operations and maintenance personnel can act as "directors," inputting different disaster scenario parameters, such as typhoon landfall point, earthquake magnitude and epicenter, and hail coverage. Based on these parameters, the system rapidly simulates the extent of damage to the power grid, the resulting material demands, and the impact on the transportation network in a digital twin environment. Subsequently, it drives a reinforcement learning agent to conduct emergency response simulations in this simulated disaster environment, generating various dynamic allocation and priority-based contingency plans. The system performs multi-dimensional quantitative evaluations of the expected effects of each plan, including but not limited to: average material shortage time, highest urgency demand fulfillment rate, total transportation cost, and power restoration time for critical users. By comparing the evaluation indicators of different plans, operators can clearly understand the advantages and disadvantages and applicable conditions of various strategies, enabling them to quickly select or integrate the most suitable plan when a real disaster occurs, or even directly activate the optimal plan. This provides powerful, data-driven, forward-looking decision support for practical application, transforming passive response into proactive preparedness.

[0015] Furthermore, the system's layered distributed architecture deployment scheme fully considers the comprehensive requirements for reliability, efficiency, and low latency in emergency response. Deploying the intelligent decision-making layer as a centralized "cloud brain" leverages the powerful computing capabilities of the cloud for complex model training and strategy optimization, ensuring the up-to-dateness and efficiency of core algorithms. Deploying the environmental perception and modeling layer, and the scheme execution and feedback layer at the edge (such as a local power grid emergency command center), close to the data source and execution terminal, offers significant advantages: the edge can quickly collect and process local real-time data (such as warehouse inventory video recognition and local traffic flow), reducing latency and bandwidth pressure from data uploads to the cloud; simultaneously, the edge can quickly respond to and execute instructions issued by the decision-making layer, greatly reducing instruction transmission and execution latency. The cloud and edge synchronize critical data and transmit instructions through a secure and encrypted network channel, ensuring data integrity and confidentiality while forming a collaborative model of "centralized optimization in the cloud and rapid execution at the edge." This architecture guarantees the global optimality of core decisions while meeting the critical low-latency requirements of emergency response, enhancing the system's practicality and robustness.

[0016] Furthermore, this claim clarifies that the method can be implemented in the form of a computer program and executed by a general-purpose or dedicated processor. This means that the method is not merely a theoretical framework, but rather a practical and reusable software implementation. When the computer program containing the steps of this method is loaded and executed by the processor, the processor will sequentially complete the construction and updating of the digital twin environment model, the definition of the core elements of the reinforcement learning framework, the training of the agent, or the invocation of the trained policy network, according to the program instructions, ultimately realizing the dynamic allocation and priority ranking functions as described in any one of claims 1 to 6. This implementation method allows the innovative approach to be easily integrated into existing power grid emergency command information systems as a core module for intelligent decision support, improving the intelligence level and response efficiency of the entire power grid emergency management system. The program implementation is also easy to deploy and adapt to different power grid companies, possessing significant promotional value.

[0017] This invention provides a method for dynamic allocation and prioritization of smart supply chains for power grid emergency materials based on reinforcement learning, which has the following beneficial effects: 1. By constructing a digital twin environment integrating multi-source real-time data, this method enables high-fidelity simulation and extrapolation of the entire resource allocation process in virtual space. This allows decision-makers to quickly generate and evaluate multiple dynamic allocation plans when actual disasters occur, significantly shortening the response time from disaster analysis to decision-making. Simultaneously, the environment incorporates power grid fault prediction information, making resource allocation plans forward-looking, enabling anticipation of expected needs and resource preparation, effectively enhancing the initiative in responding to emergencies.

[0018] This method designs a multi-objective reward function that comprehensively considers the degree of relief of material shortages, the timeliness of delivery of critical materials, and transportation costs, guiding the system to automatically optimize the balance between meeting the most urgent needs and controlling overall operating costs. In particular, by dynamically assessing the urgency of needs at each disaster-stricken point and prioritizing them accordingly, it ensures that limited emergency resources (such as special equipment and transportation capacity) are preferentially allocated to critical nodes that have the greatest impact on the safe and stable operation of the power grid and the restoration of power supply to important users. This maximizes the efficiency of emergency resource utilization and avoids resource misallocation caused by egalitarianism or empiricism.

[0019] The collaborative decision-making mechanism and hierarchical distributed architecture incorporated in this method ensure the feasibility and efficiency of the allocation plan in complex real-world environments. The virtual "collaborative decision-making module" can simulate cross-regional and multi-warehouse collaboration, fine-tuning the main plan to meet localized constraints and enhancing the adaptability of the plan's implementation. Simultaneously, online fine-tuning is achieved through feedback from actual performance data after plan execution, enabling the system to continuously learn and adapt to dynamic changes in transportation networks, inventory status, and other factors, forming a closed loop of "decision-execution-feedback-optimization." This enhances the resilience and reliability of the entire supply chain system in the face of uncertainty and sudden disruptions. Attached Figure Description

[0020] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0021] Figure 1 This is an overview of the method and a flowchart of the core steps of the present invention; Figure 2 This is a flowchart of the dynamic material demand urgency assessment model of the present invention; Figure 3 This is a flowchart illustrating the workflow of the priority ranking network of the present invention. Figure 4 This is a flowchart illustrating the calculation logic of the multi-objective reward function of this invention. Figure 5 This is a flowchart illustrating the system architecture and collaborative decision-making process of the present invention. Detailed Implementation

[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses consistent with some aspects of this disclosure as detailed in the appended claims.

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0024] Example 1: Rapid Response of Emergency Supplies During Typhoon Disasters Scenario: A strong typhoon hits a coastal area, causing damage to the power grid in multiple areas, especially a major urban area (Area A) and a remote town (Township B) where large-scale power outages occurred.

[0025] Usage process: Initialization: The operator logs into the system and activates the typhoon emergency mode. The system automatically accesses meteorological data and marks potentially affected power grid equipment and areas. The digital twin environment shows that the central warehouse (Warehouse X) has sufficient inventory, but the regular road leading to Township B is blocked due to flooding.

[0026] Decision-making and prioritization: The system automatically assesses urgency based on preset rules: Area A, involving hospitals and government departments and with a high density of users, scores higher in urgency than Township B. However, Township B has a longer estimated power outage time and is an isolated power grid, resulting in a higher score for Township B as well. System-generated solution: Prioritize the allocation of generator trucks and transformers from Warehouse X to Area A; simultaneously, when allocating resources to Township B, the system automatically plans alternative routes and allocates some resources from a nearby smaller warehouse (Warehouse Y), prioritizing this task as the second priority.

[0027] Execution and Feedback: After confirming the plan, operators issued instructions. The system displayed the real-time location of the convoy heading to Area A. Simultaneously, feedback indicated insufficient inventory at Warehouse Y. The system immediately and dynamically adjusted the plan, primarily departing from Warehouse X, but re-optimized the mixed route to Township B to save time. This response data was recorded by the system and used to optimize route planning rules for future typhoon scenarios.

[0028] Results: Critical power was restored to Area A within 4 hours, and supplies were delivered and power was restored to Township B via a detour within 8 hours. The system achieved efficient utilization of emergency resources through dynamic path planning and prioritization.

[0029] Example 2: Coordinated allocation of multiple warehouses triggered by snow and ice disasters Scenario: A severe snow and ice disaster has occurred in the northern mountainous area, causing multiple power transmission lines to become icy and malfunction. Multiple disaster-stricken areas are simultaneously issuing requests for supplies, and the inventory of a single warehouse cannot meet all the needs.

[0030] Usage process: Initialization: The operator inputs disaster site information. The system displays that the three main disaster sites (C, D, and E) have similar urgency levels, but different types and quantities of required materials. The digital twin environment shows that warehouses A and B within the area each have their own strengths in terms of inventory.

[0031] Decision Making and Prioritization: The system initiates a collaborative decision-making mechanism. The initial plan is for warehouse A to supply goods to points C and D, and warehouse B to supply goods to point E. However, simulation by the collaborative module reveals that the transportation route from warehouse A to point D is extremely risky. The module fine-tunes the plan: warehouse B supplies goods to both points D and E simultaneously, while warehouse A focuses its supply on point C, and prompts the system to assign vehicles capable of navigating icy and snowy conditions to the task at point D. The system ultimately outputs a priority task list that considers both warehouse inventory balance and transportation safety.

[0032] Execution and Feedback: The operator approved the slightly adjusted plan. During execution, the system monitored the slow progress of the convoy heading to point D, automatically triggering an alarm and suggesting that a small amount of urgently needed supplies be temporarily allocated from another backup warehouse and airlifted to point D. This collaborative data was recorded, enriching the multi-warehouse supply rule base.

[0033] Results: Through cross-warehouse collaborative supply and dynamic adjustments, the material needs of the three disaster-stricken areas were met within the stipulated time, avoiding delays caused by inventory issues in a single warehouse or single route.

[0034] Example 3: Dynamic Adjustment of Priority After Sudden Geological Disasters Scenario: A landslide triggered by torrential rain caused power grid towers to collapse, blocking the only road leading to the core fault location.

[0035] Usage process: Initialization: The operator enters fault information, and the system determines that the fault point (point F) has the highest urgency. However, the digital twin environment shows that the main road is interrupted, and the estimated repair time is 12 hours.

[0036] Decision-making and prioritization: The system did not rigidly adhere to sending a single large piece of equipment to point F. Instead, it dynamically reassessed all tasks based on real-time road conditions: First, it issued instructions to dispatch an engineering team to repair the road; simultaneously, the large equipment originally planned for point F was temporarily reassigned to another less urgent but accessible fault point (point G), and a batch of smaller equipment available for temporary power supply was allocated to attempt to approach the area surrounding point F via side roads. The task priority sequence dynamically changed as the road repair progressed.

[0037] Execution and Feedback: Operators execute deployments according to a dynamically generated priority list. When half of the road repairs are completed, the system updates the estimated time and notifies the heavy equipment to prepare to reroute to point F. Throughout the process, the system ensures that there are always executable tasks running.

[0038] Results: In the event of a main road disruption, dynamic priority adjustment ensured the continued progress of emergency work in other areas and made full preparations for the final repair of point F, minimizing overall power outage losses.

[0039] Example 4: Preventive Simulation for Power Supply Security During Major Events Scenario: A city is about to host a large-scale international conference, and it is necessary to ensure the absolute safety of the power grid. A detailed emergency material support plan needs to be developed.

[0040] Usage process: Initialization: Instead of starting the real-time decision-making mode, the operator enters the "simulation and deduction" module. Input parameters such as the location of the event venue and the distribution of important loads, and assume various fault scenarios (such as main power cable failure, venue substation failure, etc.).

[0041] Simulation and Evaluation: The system runs these scenarios in a digital twin environment. For example, when simulating a main power cable failure, the system automatically generates multiple plans for allocating generator trucks and cables from various backup warehouses, and calculates the estimated response time, resource consumption, and coverage for each plan. Operators can compare the advantages and disadvantages of different plans.

[0042] Contingency Plan Generation: Based on the simulation results, the system assists operators in developing several optimal contingency plans. For example, it determines to pre-deploy specific types of generator vehicles at several key nodes closest to the venue during the event, and clarifies the activation procedures and priorities under different fault levels.

[0043] Results: Through prior simulation and deduction, a scientific and quantifiable emergency response plan was developed, which improved the predictability and reliability of power supply protection work for major events.

[0044] Example 5: Rapid Suppression of Chain Faults in Local Power Grid Equipment Scenario: A critical substation in the city center malfunctions, triggering a chain reaction that causes multiple surrounding substations to overload, requiring emergency allocation of resources for load transfer and equipment replacement.

[0045] Usage process: Initialization: Operators report the initial fault point. Based on the power grid topology model, the system quickly predicts potentially affected secondary sites and automatically generates virtual material requirements for these sites. The digital twin environment clearly demonstrates the scope and path of the fault's impact.

[0046] Decision-making and prioritization: The system not only handles current faults but also proactively pre-allocates resources for potential secondary faults. In its generated plans, the highest priority task is isolating the initial fault point and restoring power to the main grid; simultaneously, tasks such as pre-allocating maintenance resources and switching equipment for medium- and high-risk sites are also given high priority to prevent the fault from escalating. The system automatically calculates and prompts for the minimum warning inventory level of required resources.

[0047] Execution and Feedback: Operators issued instructions in an orderly manner according to the priority sequence given by the system. When a secondary site actually showed signs of overload as predicted in the warning, the pre-set resources and solutions were immediately activated, quickly suppressing the spread of the fault. This process verified the effectiveness of the system's predictive allocation rules.

[0048] Results: By deeply integrating resource allocation with fault prediction, a shift from passive response to proactive defense was achieved, rapidly suppressing the expansion of local faults and minimizing the impact of power outages.

[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for dynamic allocation and priority ranking of smart supply chains for power grid emergency materials based on reinforcement learning, characterized by: Includes the following steps: S1: Construct a digital twin environment model of the power grid emergency material supply chain. This model integrates geographic information system, real-time traffic network data, warehouse inventory status and power grid fault prediction information to simulate the entire process from material out of the warehouse and transportation to the damaged power grid point, and provides an interactive environment for reinforcement learning agents. S2: Define the core elements of the reinforcement learning framework, including a state space, an action space, and a reward function. The state space includes warehouse inventory, the amount of materials in transit, the urgency level and predicted demand of materials at each disaster-stricken point, and the real-time accessibility status of transportation routes. The action space is a set of decisions that the agent can execute, including allocating specific types and quantities of materials from a designated warehouse to a designated disaster-stricken point, and setting execution priorities for multiple concurrent allocation tasks. The reward function is calculated based on factors such as the overall degree of relief of material shortages in the system after the action is executed, the timeliness of delivery of key materials, and transportation costs. S3: Train the reinforcement learning agent so that it can learn through trial and error in the digital twin environment model and finally obtain a policy network that can maximize long-term cumulative rewards. This policy network can directly output the optimal dynamic allocation instructions for resources and the sequence of task priorities based on real-time state input.

2. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 1, characterized in that: In step S1, the assessment of the urgency level of the material demand is a dynamic process, which is based on, but is not limited to: the voltage level of the faulty power grid equipment, the criticality of the equipment in the power grid topology, the importance of the affected users, and the estimated duration of the power outage caused by the fault. The urgency score of each disaster point is calculated and updated in real time through a multi-factor weighted assessment model.

3. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 1, characterized in that: The "priority sorting" action in the action space in step S2 is implemented through an independent sorting network. This network receives the feature vectors of all currently pending allocation tasks and outputs an optimal task execution sequence. The optimization goal of this sequence is to ensure that tasks with high urgency can obtain material and transportation resources first, while avoiding low-priority tasks from being indefinitely postponed, thereby maximizing the overall emergency efficiency of the system.

4. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 1, characterized in that: The reward function in step S2 is designed in a multi-objective weighted form, which specifically includes: giving a high positive reward for successfully delivering supplies to disaster-stricken areas with high urgency, imposing a high negative reward for failure of the task due to delayed delivery of supplies or insufficient inventory, and at the same time, appropriately penalizing the total cost of the transportation route, so as to guide the agent to seek a balance between meeting urgent needs and controlling operating costs.

5. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 1, characterized in that: The method also includes a collaborative decision-making mechanism. When the allocation plan output by the agent involves cross-regional and multi-warehouse collaborative supply, the system will activate a virtual "collaborative decision-making module". This module simulates the information interaction between various warehouse management units, fine-tunes the plan based on the agent's main plan, ensures the feasibility and efficiency of the plan at the distributed execution level, and feeds back the fine-tuned results to the agent as new learning experience.

6. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 1, characterized in that: The reinforcement learning agent is trained using proximal policy optimization or deep deterministic policy gradient algorithm. Its network structure includes a convolutional neural network branch for perceiving the environmental state and a recurrent neural network branch for processing serialized task information, so as to achieve effective fusion and utilization of spatial and temporal dimension information.

7. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claims 1 to 6, characterized in that: The system includes: The environmental perception and modeling layer is responsible for accessing multi-source heterogeneous data, constructing and updating the digital twin environment model in real time; The reinforcement learning intelligent decision-making layer has a built-in trained policy network, receives real-time status information from the environmental perception layer, and outputs resource allocation plans and task priority ranking lists. The solution execution and feedback layer issues instructions from the decision-making layer to the actual warehouse management system and logistics execution system, and collects the actual effect data after the instructions are executed, which is then sent back to the intelligent decision-making layer for online fine-tuning and adaptive optimization of the strategy model.

8. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 7, characterized in that: The system also includes a simulation and contingency plan evaluation module. Before an actual disaster occurs, operators can input different disaster scenario parameters to drive the system to conduct rapid simulations in a digital twin environment, generate multiple alternative dynamic contingency plans, and evaluate the expected effects of each plan, providing forward-looking decision support for actual combat.

9. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning as described in claim 7, characterized in that: The system adopts a layered distributed architecture. The intelligent decision-making layer serves as the cloud brain, while the environmental perception and modeling layer and the scheme execution and feedback layer can be deployed on the edge. Data synchronization and instruction transmission are carried out through a secure and encrypted network, which not only ensures the centralized optimization of the core algorithm but also meets the requirements of low latency in emergency response.

10. The method for dynamic allocation and priority ranking of smart supply chain for power grid emergency materials based on reinforcement learning according to claim 1, characterized in that: When the program is executed by the processor, it implements the steps of the method for dynamic allocation and priority ranking of smart supply chain of power grid emergency materials based on reinforcement learning interaction as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Industrial gas optimization scheduling system based on digital twinning and control method thereof

    CN122243150A