A power dispatch optimization method and system based on deep reinforcement learning

By adopting a power dispatch optimization method based on deep reinforcement learning, the problems of insufficient multi-source data governance and single indicator evaluation in power dispatch are solved. This method enables rapid identification and accurate matching of power grid faults, and improves the efficiency of multi-objective collaborative optimization and fault handling in power dispatch.

CN121618635BActive Publication Date: 2026-04-21STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
Filing Date
2026-02-02
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing power dispatch optimization methods lack systematic governance of multi-source historical fault data of the power grid system, resulting in a lack of historical experience support when facing new faults or complex scenarios, leading to large decision-making biases. Furthermore, the evaluations often focus on a single indicator, ignoring multi-objective requirements and lacking dynamic monitoring and hazard prediction capabilities, resulting in low efficiency in solution generation and failing to meet the timeliness requirements for early detection and early handling of power grid faults.

Method used

A deep reinforcement learning-based approach is used to collect historical multi-source fault data of the power grid, classify and process the data, and label the changes in electrical parameters and the effectiveness indicators of the solutions before the fault. The state and effect sets are defined, the power grid status is evaluated in real time, and suitable scheduling solutions are selected. Combined with the technician matching database, the solution is optimized to achieve multi-objective collaborative optimization and rapid response.

Benefits of technology

It achieves multi-objective collaborative optimization of power dispatching schemes, improves fault recovery speed and power supply security, reduces the risk of fault escalation, improves the efficiency of handling new faults, and enriches the dispatching optimization level of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121618635B_ABST
    Figure CN121618635B_ABST
Patent Text Reader

Abstract

This invention discloses a power dispatch optimization method and system based on deep reinforcement learning, specifically relating to the field of power dispatch technology. This invention achieves multi-objective collaborative optimization of dispatch schemes by constructing a multi-dimensional quantitative evaluation system. Unlike traditional methods that rely on single-index evaluation, this invention calculates fault recovery coefficients, safety and stability coefficients, and economic evaluation coefficients in conjunction with rule matching coefficients. Based on the calculation results, schemes are selected, achieving a balance between multiple objectives such as fault recovery speed, power supply security, operational economy, and fault similarity, thereby improving the comprehensive adaptability of dispatch schemes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power dispatching technology, and more specifically, to a power dispatching optimization method and system based on deep reinforcement learning. Background Technology

[0002] As the power grid continues to expand in scale and its topology becomes increasingly complex, fault scenarios such as line short circuits, generator failures, and transformer overloads are becoming more diversified and complex, placing higher demands on the rapid response capability and multi-objective optimization capability of power dispatching schemes.

[0003] However, existing power dispatch optimization methods still have the following shortcomings in practical applications:

[0004] First, traditional dispatching relies heavily on manual experience or single real-time data to formulate solutions. It does not systematically manage multi-source historical fault data such as power grid system, fault recorder, and dispatch log, and fails to effectively extract the correlation between the changes in electrical parameters before the fault and the corresponding dispatching solutions. As a result, when facing new faults or complex scenarios, the solutions lack historical experience support and are prone to decision-making bias.

[0005] In addition, the evaluation of dispatching schemes often focuses on a single indicator, ignoring the multi-objective requirements of power grid operation, which need to ensure power supply reliability while also taking into account economic efficiency and equipment safety.

[0006] Finally, existing technologies mostly formulate response plans passively after a fault occurs, lacking the ability to dynamically monitor the real-time operating status of the power grid and predict potential hazards. Even if a potential hazard is identified, it is difficult to quickly match and adapt historical dispatch plans based on fault characteristics. They often rely on manual screening one by one, resulting in low efficiency in plan generation and failing to meet the timeliness requirements for early detection and early handling of power grid faults.

[0007] To address this, a power dispatch optimization method and system based on deep reinforcement learning is proposed. Summary of the Invention

[0008] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a power dispatch optimization method and system based on deep reinforcement learning.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A power dispatch optimization method based on deep reinforcement learning includes:

[0011] S1: Collect historical multi-source fault data of the power grid, classify and process each fault, extract historical dispatch solutions, and mark the changes in electrical parameters before the fault and the effectiveness indicators of the solution for each fault.

[0012] S2: Define a state set based on the changes in electrical parameters before the fault is marked, where the state set includes voltage, current, power, and frequency; define an effect set including the fault type based on the marked scheme effect indicators; where the effect set includes fault type number, fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient.

[0013] S3: Real-time acquisition of changes in real-time electrical parameters of the power grid, and comprehensive evaluation with state sets of different fault types. Based on the evaluation results, determine whether there are potential fault hazards in the power grid. If there are potential fault hazards, further select candidate scheduling schemes from each state set based on the evaluation results.

[0014] S4: For the set of effects of each group of candidate scheduling schemes, combine the rule matching coefficient to determine the current power grid's hidden fault types and preliminary selection schemes;

[0015] S5: Output the current power grid's potential fault types and preliminary solutions to the pre-built technician matching database, select technicians with higher optimization evaluation coefficients, and send the real-time changes in the power grid's electrical parameters, potential fault types, and preliminary solutions.

[0016] Specifically, the calculation process for the fault recovery coefficient within the effect set of step S2 is as follows:

[0017] The fault recovery time is calculated by taking the time point of formal implementation of the plan as the starting point and the time point when all parameters in the power grid state set remain within the rated range for a set duration as the ending point.

[0018] Obtain the total grid loss during the fault recovery process, and calculate the difference between it and the average grid loss in the set time zone before the fault to obtain the grid loss increment;

[0019] The fault recovery time and network loss increment are comprehensively processed to determine the fault recovery coefficient of the solution.

[0020] Specifically, the calculation process for the safety stability coefficient and economic evaluation coefficient within the effect set of step S2 is as follows:

[0021] Within the set evaluation period from the formal implementation of the plan to the recovery from the fault, the cumulative time during which the voltage of each node is within the voltage qualification range is counted, and the proportion of the cumulative time within the set evaluation period is calculated as the voltage qualification rate.

[0022] The cumulative time that the system frequency is within the acceptable range is statistically analyzed, and the proportion of time within the set evaluation period is calculated as the frequency pass rate.

[0023] The percentage of equipment that experienced no power grid overload during the assessment period is used as the equipment qualification rate.

[0024] The voltage qualification rate, frequency qualification rate, and equipment qualification rate are comprehensively processed to determine the safety and stability coefficient of the scheme.

[0025] The additional costs incurred by the increase in network loss during the period from the formal implementation of the plan to the fault recovery are statistically analyzed. The ratio is calculated with the additional cost as the numerator and the normal operation cost as the denominator. The economic evaluation coefficient is obtained by subtracting the ratio from 1.

[0026] Specifically, the process of comprehensively evaluating the real-time electrical parameter changes of the power grid and the state sets of different fault types in step S3 is as follows:

[0027] For the voltage, current, power, and frequency in the real-time electrical parameters of the power grid and the voltage, current, power, and frequency in the state sets corresponding to different fault types, a comprehensive processing is performed to obtain the rule matching coefficient between the current real-time electrical parameters of the power grid and each set of state sets.

[0028] Specifically, the process of determining whether there are potential faults and screening candidate scheduling schemes in step S3 is as follows:

[0029] If one or more sets of rule matching coefficients are less than the preset coefficient threshold, it is determined that there is a potential fault. The set of states with rule matching coefficients less than the coefficient threshold is selected as the candidate set, and the historical scheduling solutions corresponding to each set of candidate sets are identified as the candidate scheduling schemes for the current power grid.

[0030] Specifically, the process of determining the current power grid potential fault type and preliminary selection scheme in step S4 is as follows:

[0031] The fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient are extracted from the effect set of each group of candidate scheduling schemes, and combined with the rule matching coefficient for comprehensive processing to determine the comprehensive scheme coefficient of each group of candidate scheduling schemes.

[0032] The candidate scheduling schemes with higher comprehensive coefficients are extracted as preliminary schemes, and the fault type numbers in the effect set corresponding to the preliminary schemes are identified as the potential fault types of the current power grid.

[0033] Specifically, the calculation process for the technician optimization evaluation coefficient in step S5 is as follows:

[0034] In the skilled worker matching database, online skilled workers are identified as candidates. The special matching rate, scheduling performance rate and optimization efficiency of each candidate are analyzed. After comprehensive processing, the optimization evaluation coefficient of each candidate is obtained.

[0035] Specifically, the calculation process for the special project matching rate, scheduling performance rate, and optimization efficiency is as follows:

[0036] Extract the types of potential faults in the current power grid, and calculate the proportion of cases that match the types of potential faults in the current power grid among the dispatching schemes that each candidate personnel has previously undertaken and optimized, as the special matching rate;

[0037] From the optimized scheduling plans completed by each candidate, extract the performance coefficient of each group of scheduling plans after implementation and take the average value to obtain the scheduling performance rate;

[0038] Identify the time points at which candidate personnel accept the scheduling plan and at which they complete the optimization of the scheduling plan. Calculate the difference between the two sets of time points to obtain the time taken by the candidate personnel. Calculate the average of the time taken by each set of candidate personnel to obtain the optimized time efficiency.

[0039] Specifically, the calculation process for the scheme performance coefficient is as follows:

[0040] Extract the fault recovery coefficient of each scheduling scheme after its implementation. and safety stability coefficient The fault recovery coefficient of each group of scheduling schemes and safety stability coefficient Each candidate is multiplied by its corresponding preset weight coefficient, and then the results are summed to determine the performance coefficient of each candidate for each group of scheduling schemes.

[0041] A power dispatch optimization system based on deep reinforcement learning includes:

[0042] Data construction module: Stores historical power grid fault data, classifies and processes each fault, extracts historical dispatch solutions, marks the changes in electrical parameters before the fault and the effectiveness indicators of the solution for each fault, and builds a knowledge base that associates faults with solutions.

[0043] Learning Definition Module: Based on the changes in electrical parameters before the fault and the effectiveness indicators of the solution, define a state set and an effectiveness set containing the fault type, respectively; the state set includes voltage, current, power, and frequency; based on the marked effectiveness indicators of the solution; the effectiveness set includes fault type number, fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient;

[0044] Real-time evaluation module: Collects real-time changes in electrical parameters of the power grid and performs a comprehensive evaluation with the state sets of different fault types. Based on the evaluation results, it determines whether there are potential faults in the power grid. If there are potential faults, it selects candidate scheduling schemes from each state set based on the evaluation results.

[0045] The scheduling determination module: For the effect set of each group of candidate scheduling schemes, it determines the current power grid's hidden fault type and preliminary scheme by combining the rule matching coefficient. It outputs the current power grid's hidden fault type and preliminary scheme to the pre-built technician matching database, selects technicians with higher optimization evaluation coefficients, and sends the real-time electrical parameter changes of the power grid, hidden fault type, and preliminary scheme.

[0046] The technical effects and advantages of this invention are as follows:

[0047] (1) By constructing a multi-dimensional quantitative evaluation system, the scheduling scheme is optimized in a multi-objective collaborative manner. Unlike the limitations of the single indicator evaluation in traditional methods, the scheme is calculated by defining the fault recovery coefficient, safety and stability coefficient, economic evaluation coefficient and rule matching coefficient. Based on the calculation results, the scheme is selected to achieve a balance of multiple objectives such as fault recovery speed, power supply safety, operation economy and fault similarity, and improve the comprehensive adaptability of the scheduling scheme.

[0048] (2) By collecting power grid electrical parameters in real time and calculating the matching coefficient with the historical fault status set, the early identification of potential faults and accurate matching of solutions can be achieved, which can significantly shorten the response time for handling potential faults and reduce the risk of fault escalation.

[0049] (3) By establishing a scientific skilled worker matching mechanism, namely, based on the special matching rate, scheduling performance rate, and optimization efficiency quantitative screening, the implementation effect of the plan is guaranteed, and the skilled workers' adaptive adjustment experience to the plan will be stored back to the historical database to continuously enrich the knowledge base and continuously improve the scheduling optimization level.

[0050] (4) Through systematic data governance and knowledge base construction, multi-source historical fault data is transformed into structured related data to form a reusable and iterative fault and solution related knowledge base. The knowledge base can directly match the historical best solution, thereby improving the efficiency of new fault handling. Attached Figure Description

[0051] Figure 1 This is a flowchart of a power dispatch optimization method based on deep reinforcement learning according to the present invention.

[0052] Figure 2 This is a schematic diagram of a power dispatch optimization system based on deep reinforcement learning according to the present invention. Detailed Implementation

[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] Example 1

[0055] like Figure 1 As shown, a power dispatch optimization method based on deep reinforcement learning is as follows:

[0056] Data governance construction: Collect and clean historical multi-source fault data of the power grid, build a multi-level fault classification system, extract historical dispatch solutions after classifying each fault, and mark the changes in electrical parameters before the fault and the effectiveness indicators of the solution for each fault, and build a knowledge base for the association between fault and solution.

[0057] Collect data from the power grid SCADA system, fault recorder data, dispatch operation logs, and environmental parameters, covering basic fault information, changes in electrical quantities, and the handling process;

[0058] The collection and cleaning of historical multi-source fault data of the power grid specifically involves: collecting historical multi-source fault data including SCADA system monitoring data, fault recorder data, dispatch operation logs, and equipment operation ledgers. Among them, the SCADA system monitoring data covers node voltage, line current, generator power, and load demand time series data within one hour before and after the fault occurred; data cleaning includes removing outliers (using the 3σ principle to identify and delete values ​​that deviate from the data mean by more than 3 times the standard deviation), filling missing values ​​(using linear interpolation to fill missing data in continuous time series with a duration of no more than 5 minutes), and data standardization (normalizing parameters such as voltage, current, and power to the [0,1] interval).

[0059] The multi-level fault classification system is divided into three levels. The first level includes line faults, generator faults, load faults, and transformer faults. The second level further refines the first level classification. For example, the second level of line faults includes line short circuits, line trips, and line overloads, while the second level of generator faults includes generator shutdowns, generator output fluctuations, and generator excitation faults. The third level further clarifies the fault characteristics. For example, the third level of line short circuits includes single-phase ground faults, two-phase short circuits, and three-phase short circuits, and each type of third-level fault is assigned a unique fault type number. The fault-solution association knowledge base is built by storing mapping relationships using triples (fault type, electrical parameter changes, and effect indicators).

[0060] Define the learning environment: Define a set of states based on the changes in electrical parameters before the fault is marked, where the set of states includes voltage, current, power, and frequency; Define an effect set containing fault types based on the marked scheme effect indicators; where the effect set includes fault type number, fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient.

[0061] Specifically:

[0062] The starting point is the time when the formal implementation plan is taken as the starting point, and the ending point is the time when all parameters in the power grid status set remain within the rated range for a set duration; the set duration is 2 minutes.

[0063] The fault recovery time is calculated by taking the time difference between the end point and the start point; the key parameter stability standards are: node voltage deviation ≤ ±5%, line current ≤ 1.05 times the rated value;

[0064] Obtain the total grid loss during the fault recovery process, and calculate the difference between it and the average grid loss in the set time zone before the fault to obtain the grid loss increment;

[0065] For each fault type number within the set of solutions, the average fault recovery time of each group of solutions belonging to the same number is calculated and used as a reference recovery time.

[0066] The average value of the network loss increments among the schemes belonging to the same category number is calculated and used as a reference network loss increment;

[0067] Using formula Perform weighted calculations to determine the fault recovery coefficient of the solution. ;in Indicates the reference recovery time and reference network loss increment for schemes belonging to the same category number; Indicates the fault recovery time and network loss increment of the solution; These are preset weighting coefficients;

[0068] Taking a three-phase short-circuit fault (fault type number: LS-3-04) on 110kV line L5 as an example, the calculation process is fully presented through a comparison of three sets of historical dispatch schemes and one set of new schemes to be evaluated:

[0069] Extract historical scheduling schemes (numbers: LS-3-01, LS-3-02, LS-3-03) belonging to the same fault type from the fault and scheme association knowledge base.

[0070] The average values ​​of the fault recovery time and network loss increment for the three schemes LS-3-01, LS-3-02 and LS-3-03 are calculated to obtain the reference recovery time and reference network loss increment.

[0071] And calculate the fault recovery coefficient to reflect the advantages and disadvantages of different solutions under the same fault type. Assumptions:

[0072] The fault recovery coefficient of the historical scheme LS-3-02 is 0.776;

[0073] The fault recovery coefficient of the historical scheme LS-3-01 is 0.908;

[0074] The fault recovery coefficient of the historical scheme LS-3-03 is 0.899;

[0075] The fault recovery coefficient of the historical scheme LS-3-04 is 0.863;

[0076] The ranking results are: LS-3-01 (0.908) > LS-3-03 (0.899) > LS-3-04 (0.863) > LS-3-02 (0.776), indicating that the historical solution LS-3-01 has the best solution effect under this fault type.

[0077] Within the set evaluation period from the formal implementation of the plan to the recovery from the fault, the cumulative time during which the voltage of each node is within the qualified voltage range is counted, and the proportion of the cumulative time within the set evaluation period is calculated as the voltage qualification rate; this reflects the stability of the node voltage (voltage is the basis for power grid balance and safe operation of equipment; unqualified voltage may lead to equipment burnout and abnormal power consumption for users).

[0078] The system calculates the cumulative time that the frequency is within the acceptable range and uses the proportion of time within the set evaluation period as the frequency pass rate; this reflects the active power balance of the power grid (frequency fluctuations directly affect the lifespan and operating efficiency of rotating equipment such as generators and motors).

[0079] The percentage of equipment that experienced no overload within the power grid during the assessment period is used as the equipment qualification rate, reflecting the operational safety of the equipment (overload can lead to excessive temperature rise, insulation aging, and even secondary failures).

[0080] The criteria for determining no overload conditions are: after the fault is resolved, within the set evaluation period, all key equipment in the power grid (generators, lines, transformers, switches, etc.) continuously operate within the rated parameter range;

[0081] For the fault type number within the solution effect set, the average voltage qualification rate of each group of solutions belonging to the same number is calculated as the reference voltage qualification rate.

[0082] The average of the pass rates of frequencies belonging to the same category number is calculated and used as the reference frequency pass rate; the average of the pass rates of equipment belonging to the same category number is calculated and used as the reference equipment pass rate.

[0083] Using formula Perform weighted calculations to determine the safety and stability coefficient of the scheme. ;in This indicates the voltage qualification rate, frequency qualification rate, and equipment qualification rate of the scheme. This indicates the pass rate of reference voltage, pass rate of reference frequency, and pass rate of reference equipment that belong to the same category number as the scheme. These are preset weighting coefficients;

[0084] Fault type: The main transformer T4 (rated capacity 630MVA) of the 220kV substation M2 was overloaded (load rate reached 118%) due to a sudden increase in summer load. The fault type number is TO-220-4, which affects the power supply of three downstream 110kV substations.

[0085] A new dispatching scheme (No.: TO-220-04) for the overload fault of the T4 main transformer in the M2 substation was implemented. After execution, 24 hours of operational data were collected. The core indicators are as follows:

[0086] Voltage pass rate: 99.3% (only during the evening peak period of 19:00-19:30, the voltage of two 110kV nodes briefly dropped to 101.5kV, with a total non-compliance time of 36 minutes).

[0087] Frequency pass rate: 99.5% (during the morning peak period of 7:30-7:45, the frequency dropped to 49.75Hz due to a sudden increase in load, with a cumulative non-compliance time of 18 minutes).

[0088] Equipment pass rate: 98.6% (only a few pieces of equipment were overloaded during the midday period of 12:00-12:40, lasting for 40 minutes; all other equipment was operating normally).

[0089] For historical schemes with the same fault type number, calculate the "reference voltage qualification rate", "reference frequency qualification rate", and "reference equipment qualification rate" (average of similar schemes).

[0090] For example, historical schemes with the same fault type numbering include TO-220-1, TO-220-2, TO-220-3, and TO-220-4. The safety and stability coefficient is calculated. This reflects the advantages and disadvantages of different solutions under the same fault type.

[0091] The fault recovery coefficient of the historical scheme TO-220-1 is 0.891;

[0092] The fault recovery factor for the historical scheme TO-220-2 is 0.908;

[0093] The fault recovery coefficient of the historical scheme TO-220-3 is 0.912;

[0094] The fault recovery factor for the historical solution TO-220-4 is 0.987;

[0095] The ranking results are: Solution TO-220-4 > TO-220-3 > TO-220-2 > TO-220-1, indicating that the historical solution TO-220-4 has the best solution effect under this fault type.

[0096] The additional costs incurred by the increase in network loss during the period from the formal implementation of the plan to fault recovery are calculated. The ratio is calculated with the additional cost as the numerator and the normal operation cost as the denominator. The economic evaluation coefficient is obtained by subtracting the ratio from 1. ;

[0097] The difference between the total grid loss and the "normal operating grid loss during the same period before the fault" from the time the dispatch plan is officially implemented (such as the time when the faulty line L9 switch is disconnected) to the time when the grid is restored to stability (all node voltages and line currents meet the qualified standards and last for 3 minutes). (Only the grid loss increment during the fault handling period is calculated).

[0098] The economic expenditure generated by the increase in network loss is calculated using the formula "increase in network loss × on-grid electricity price" (the on-grid electricity price is the cost of electricity purchased by the power grid, reflecting the economic value of a unit of electricity).

[0099] Normal operating cost: The total cost of electricity purchased when the power grid is operating normally during the same period before the fault (i.e., "normal operating time × average load power × grid connection price", which should be consistent with the fault handling time to ensure fair comparison); the closer it is to 1, the better the network loss control effect and the stronger the economic efficiency of the scheme.

[0100] Dispatch basis determination: Real-time changes in power grid electrical parameters are collected and comprehensively evaluated with state sets of different fault types. Based on the evaluation results, it is determined whether there are potential faults in the power grid. If potential faults exist, candidate dispatch schemes are further selected from each state set based on the evaluation results.

[0101] Specifically:

[0102] For changes in real-time electrical parameters of the power grid, voltage, current, power, and frequency are denoted as... ;

[0103] The voltage, current, power, and frequency within the state set corresponding to different fault types are marked as follows: ;

[0104] Using formula Calculate the rule matching coefficients between the current real-time electrical parameters of the power grid and each set of states. ;in These are the weighting coefficients for each electrical parameter;

[0105] If one or more sets of rule matching coefficients If the coefficient is less than the preset threshold, a potential fault is identified; the rule matching coefficient is then selected. The set of states with values ​​less than the coefficient threshold is taken as the candidate set. The historical scheduling solutions corresponding to each candidate set are identified as the candidate scheduling solutions for the current power grid.

[0106] The power distribution network real-time monitoring system collects operating parameters for a certain period of time. After data cleaning (removing instantaneous jump values ​​from sensors), the average value is taken as the real-time parameter.

[0107] Taking three sets of states F01, F02 and F03 as examples, the rule matching coefficients between real-time parameters and the three types of fault state sets F01, F02 and F03 are calculated respectively.

[0108] Among them, F01 is a line overload fault, F02 is a single-phase ground fault, and F03 is a transformer light gas alarm fault.

[0109] The rule matching coefficient for F01 is 0.0084;

[0110] The rule matching coefficient for F02 is 0.5132;

[0111] The rule matching coefficient for F03 is 0.2594;

[0112] The matching coefficients of each group of rules are ranked as follows: F01 < F03 < F02;

[0113] Threshold comparison: Assuming the coefficient threshold is 0.15, then F01 < 0.15, which meets the threshold condition. The fault type corresponding to F01 is identified, and the line overload fault is the hidden fault type of the current power grid.

[0114] Preliminary plan determination: Based on the effect set of each group of candidate scheduling schemes, and combined with the rule matching coefficient, determine the current power grid's potential fault types and preliminary selection schemes;

[0115] Specifically:

[0116] Extract the fault recovery coefficient from the effect set of each group of candidate scheduling schemes. Safety and stability coefficient Economic evaluation coefficient And combined with rule matching coefficients Substitute into the formula Perform weighted calculations to determine the overall coefficient of each group of candidate scheduling schemes. ;in These are the fault recovery coefficients. Safety and stability coefficient Economic evaluation coefficient and rule matching coefficient Weighting coefficients;

[0117] Extraction scheme comprehensive coefficient The higher-ranking candidate scheduling schemes are used as preliminary schemes, and the fault type numbers in the effect set corresponding to the preliminary schemes are identified as the potential fault types of the current power grid.

[0118] The candidate with the highest comprehensive coefficient was selected as the initial candidate. This candidate takes into account "high degree of matching with potential hazards, fast recovery speed, safety and stability, and economic feasibility" in the current scenario.

[0119] Extract the fault type numbers from the effect set corresponding to the preliminary solution to clarify the types of potential faults in the current power grid, providing a precise direction for subsequent solution optimization.

[0120] Personnel Collaborative Optimization: Output the current power grid's potential fault types and preliminary solutions to a pre-built technician matching database, select technicians with higher optimization evaluation coefficients, and send the real-time changes in the power grid's electrical parameters, potential fault types, and preliminary solutions; after receiving and reviewing the information, the technicians determine whether adaptive adjustments to the preliminary solutions are needed, and the adjusted solutions become the final power grid dispatching solutions;

[0121] If the technician makes adaptive adjustments to the initial selected solution, the adjusted solution will be stored in the historical fault data for updating and continuous learning and optimization.

[0122] Specifically:

[0123] In the skilled worker matching database, online skilled workers are identified as candidates. The types of potential faults in the current power grid are extracted, and the number of dispatch schemes that match the current potential fault type number in the past dispatch schemes undertaken and optimized by each candidate is counted. The proportion of the number of schemes in the total number undertaken is calculated as the special matching rate. This directly reflects the expert's familiarity with the current fault. The higher the matching rate, the more accurate the control over the fault mechanism, the scope of impact, and key limiting parameters (such as node voltage and equipment load rate).

[0124] From the optimized scheduling plans completed by each candidate, extract the fault recovery coefficient of each group of scheduling plans after implementation. and safety stability coefficient The fault recovery coefficient of each group of scheduling schemes and safety stability coefficient Each candidate is multiplied by a corresponding preset weight coefficient, and then summed to determine the scheme performance coefficient for each group of scheduling schemes. The mean of the scheme performance coefficients for each candidate is calculated to obtain the scheduling performance rate for each candidate.

[0125] The optimization "hard power" of quantitative experts, for example, after an expert optimized the solution for the "main transformer overload" fault, the safety and stability coefficient increased from 0.86 to 0.98, an improvement rate of 14%, which shows that it can effectively solve the shortcomings in voltage qualification rate, equipment overload and other aspects.

[0126] Identify the time points at which candidate personnel accept the dispatch plan and at which they complete the optimization of the dispatch plan. Calculate the difference between the two sets of time points to obtain the time taken by the candidate personnel. Calculate the average of the time taken by each candidate personnel in each set to obtain the optimization time efficiency of the candidate personnel. Power grid fault hazards are time-sensitive. If experts respond too slowly, the hazard may escalate into an actual fault. Therefore, response timeliness is the key to ensuring smooth coordination of the dispatch process.

[0127] Extract the specific matching rate, scheduling performance rate, and optimization efficiency for each candidate, and label them after normalization. Using the formula Calculate the optimal evaluation coefficient for each candidate. ; These are the preset weighting coefficients;

[0128] By using online status screening and optimized efficiency sorting, we ensure that the selected technicians can respond in real time and the average time is reduced to less than 1 hour.

[0129] Due to the high matching rate of specific projects and the excellent scheduling performance, technicians can understand and adjust the solutions more efficiently, reducing the time cost of repeated communication and trial and error operations, and significantly reducing the risk of hidden dangers escalating into actual failures.

[0130] Example 2

[0131] Please see Figure 2 As shown, based on Embodiment 1 of this application, a power dispatch optimization method based on deep reinforcement learning is provided. Embodiment 2 of this application proposes a power dispatch optimization system based on deep reinforcement learning. Embodiment 2 is merely a preferred embodiment of Embodiment 1, and its implementation will not affect the individual implementation of Embodiment 1.

[0132] Specifically, the difference in the power dispatch optimization system based on deep reinforcement learning provided in Embodiment 2 of this application lies in that it includes:

[0133] Data construction module: Stores historical power grid fault data, classifies and processes each fault, extracts historical dispatch solutions, marks the changes in electrical parameters before the fault and the effectiveness indicators of the solution for each fault, and builds a knowledge base that associates faults with solutions.

[0134] Learning Definition Module: Based on the changes in electrical parameters before the fault and the effectiveness indicators of the solution, define a state set and an effectiveness set containing the fault type, respectively; the state set includes voltage, current, power, and frequency; based on the marked effectiveness indicators of the solution; the effectiveness set includes fault type number, fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient;

[0135] Real-time evaluation module: Collects real-time changes in electrical parameters of the power grid and performs a comprehensive evaluation with the state sets of different fault types. Based on the evaluation results, it determines whether there are potential faults in the power grid. If there are potential faults, it selects candidate scheduling schemes from each state set based on the evaluation results.

[0136] Scheduling Determination Module: For the effect set of each group of candidate scheduling schemes, the module determines the current power grid's hidden fault type and preliminary scheme by combining the rule matching coefficient. It outputs the current power grid's hidden fault type and preliminary scheme to the pre-built technician matching database, selects technicians with higher optimization evaluation coefficients, and sends the real-time electrical parameter changes of the power grid, hidden fault type, and preliminary scheme.

[0137] The above formulas are all dimensionless calculations. Dimensionless calculations can be performed using various methods such as standardization, which will not be elaborated here. The formulas are derived from software simulations based on a large amount of collected data, and the preset parameters in the formulas can be set by those skilled in the art according to the actual situation.

[0138] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, ATA hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state ATA hard disk.

[0139] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0140] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0143] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0144] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable ATA hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0145] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A power dispatch optimization method based on deep reinforcement learning, characterized in that, include: S1: Collect historical multi-source fault data of the power grid, classify and process each fault, extract historical dispatch solutions, and mark the changes in electrical parameters before the fault and the effectiveness indicators of the solution for each fault. S2: Define a state set based on the changes in electrical parameters before the fault is marked, where the state set includes voltage, current, power, and frequency; define an effect set including the fault type based on the marked scheme effect indicators; where the effect set includes fault type number, fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient. The calculation process for the fault recovery coefficient is as follows: The fault recovery time is calculated by taking the time point of formal implementation of the plan as the starting point and the time point when all parameters in the power grid state set remain within the rated range for a set duration as the ending point. Obtain the total grid loss during the fault recovery process, and calculate the difference between it and the average grid loss in the set time zone before the fault to obtain the grid loss increment; The fault recovery time and network loss increment are comprehensively processed to determine the fault recovery coefficient of the solution; The calculation process for the safety stability coefficient and the economic evaluation coefficient is as follows: Within the set evaluation period from the formal implementation of the plan to the recovery from the fault, the cumulative time during which the voltage of each node is within the voltage qualification range is counted, and the proportion of the cumulative time within the set evaluation period is calculated as the voltage qualification rate. The cumulative time that the system frequency is within the acceptable range is statistically analyzed, and the proportion of time within the set evaluation period is calculated as the frequency pass rate. The percentage of equipment that experienced no power grid overload during the assessment period is used as the equipment qualification rate. The voltage qualification rate, frequency qualification rate, and equipment qualification rate are comprehensively processed to determine the safety and stability coefficient of the scheme. The additional costs incurred by the increase in network loss during the period from the formal implementation of the plan to the fault recovery are statistically analyzed. The ratio is calculated with the additional cost as the numerator and the normal operation cost as the denominator. The economic evaluation coefficient is obtained by subtracting the ratio from 1. S3: Real-time acquisition of changes in real-time electrical parameters of the power grid, and comprehensive evaluation with state sets of different fault types. Based on the evaluation results, determine whether there are potential fault hazards in the power grid. If there are potential fault hazards, further select candidate scheduling schemes from each state set based on the evaluation results. S4: For the set of effects of each group of candidate scheduling schemes, combine the rule matching coefficient to determine the current power grid's hidden fault types and preliminary selection schemes; S5: Output the current power grid's potential fault types and preliminary solutions to the pre-built technician matching database, select technicians with higher optimization evaluation coefficients, and send the real-time changes in the power grid's electrical parameters, potential fault types, and preliminary solutions.

2. The power dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, The specific process of comprehensively evaluating the real-time electrical parameter changes of the power grid and the state sets of different fault types in step S3 is as follows: For the voltage, current, power, and frequency in the real-time electrical parameters of the power grid and the voltage, current, power, and frequency in the state sets corresponding to different fault types, a comprehensive processing is performed to obtain the rule matching coefficient between the current real-time electrical parameters of the power grid and each set of state sets.

3. The power dispatch optimization method based on deep reinforcement learning according to claim 2, characterized in that, The specific process for determining whether there are potential faults and screening candidate scheduling schemes in step S3 is as follows: If one or more sets of rule matching coefficients are less than the preset coefficient threshold, it is determined that there is a potential fault. The set of states with rule matching coefficients less than the coefficient threshold is selected as the candidate set, and the historical scheduling solutions corresponding to each set of candidate sets are identified as the candidate scheduling schemes for the current power grid.

4. The power dispatch optimization method based on deep reinforcement learning according to claim 3, characterized in that, The specific process for determining the current power grid potential fault type and preliminary selection scheme in step S4 is as follows: The fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient are extracted from the effect set of each group of candidate scheduling schemes, and combined with the rule matching coefficient for comprehensive processing to determine the comprehensive scheme coefficient of each group of candidate scheduling schemes. The candidate scheduling schemes with higher comprehensive coefficients are extracted as preliminary schemes, and the fault type numbers in the effect set corresponding to the preliminary schemes are identified as the potential fault types of the current power grid.

5. The power dispatch optimization method based on deep reinforcement learning according to claim 1, characterized in that, The specific calculation process for the technician optimization evaluation coefficient in step S5 is as follows: In the skilled worker matching database, online skilled workers are identified as candidates. The special matching rate, scheduling performance rate and optimization efficiency of each candidate are analyzed, and the optimization evaluation coefficient of each candidate is obtained after comprehensive processing.

6. The power dispatch optimization method based on deep reinforcement learning according to claim 5, characterized in that, The specific calculation process for the special matching rate, scheduling performance rate, and optimization efficiency is as follows: Extract the types of potential faults in the current power grid, and calculate the proportion of cases that match the types of potential faults in the current power grid among the dispatching schemes that each candidate personnel has previously undertaken and optimized, as the special matching rate; From the optimized scheduling plans completed by each candidate, extract the performance coefficient of each group of scheduling plans after implementation and take the average value to obtain the scheduling performance rate; Identify the time points at which candidate personnel accept the scheduling plan and at which they complete the optimization of the scheduling plan. Calculate the difference between the two sets of time points to obtain the time taken by the candidate personnel. Calculate the average of the time taken by each set of candidate personnel to obtain the optimized time efficiency.

7. The power dispatch optimization method based on deep reinforcement learning according to claim 6, characterized in that, The specific calculation process for the performance coefficient of the scheme is as follows: Extract the fault recovery coefficient of each scheduling scheme after its implementation. and safety stability coefficient The fault recovery coefficient of each group of scheduling schemes and safety stability coefficient Each candidate is multiplied by its corresponding preset weight coefficient, and then the results are summed to determine the performance coefficient of each candidate for each group of scheduling schemes.

8. A power dispatch optimization system based on deep reinforcement learning, applied to the power dispatch optimization method based on deep reinforcement learning proposed in any one of claims 1-7, characterized in that, include: Data construction module: Stores historical power grid fault data, classifies and processes each fault, extracts historical dispatch solutions, marks the changes in electrical parameters before the fault and the effectiveness indicators of the solution for each fault, and builds a knowledge base that associates faults with solutions. Learning Definition Module: Based on the changes in electrical parameters before the fault and the effectiveness indicators of the solution, define the state set and the effect set containing the fault type, respectively; The state set includes voltage, current, power, and frequency; based on the labeled scheme effect indicators; the effect set includes fault type number, fault recovery coefficient, safety and stability coefficient, and economic evaluation coefficient; Real-time evaluation module: Collects real-time changes in electrical parameters of the power grid and performs a comprehensive evaluation with the state sets of different fault types. Based on the evaluation results, it determines whether there are potential faults in the power grid. If there are potential faults, it selects candidate scheduling schemes from each state set based on the evaluation results. The scheduling determination module determines the current power grid's potential fault types and preliminary solutions based on the effect set of each group of candidate scheduling schemes, combined with the rule matching coefficient. It outputs the current power grid's potential fault types and preliminary solutions to the pre-built technician matching database, selects technicians with higher optimization evaluation coefficients, and sends the real-time changes in the power grid's electrical parameters, potential fault types, and preliminary solutions.

Citation Information

Patent Citations

  • Power dispatching optimization method and system based on deep reinforcement learning

    CN119378733A

  • Distribution network fault scheduling decision generation method fusing knowledge graph

    CN120911584A