Maintenance operation ticket automatic ticketing method and system based on markov process
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明的目的是克服自动生成电网检修操作票时,依赖静态的设备拓扑与固定操作规则,难以适配复杂运行场景下检修操作的准确时序化风险管控需求,所生成的操作票准确性较低的缺点,提供一种基于马尔科夫过程的检修操作票自动成票方法及系统,通过引入马尔科夫过程优化检修操作规划中的时序决策和风险量化,提高所生成操作票的准确性,并结合对模型静态参数的动态化改造与长时序后效性修正,增强其与电网检修场景的动态适配能力,最终使生成的检修操作票兼顾准确性与现场适配性
通过引入马尔科夫过程,利用其状态、动作、下一状态的时序转移机制与概率化决策特征,以检修操作空间构建限定检修操作的状态边界及时序依赖关系,避免操作步骤的时序逻辑疏漏。同时,以静态状态转移规则为约束,生成对应动态状态转移决策权重,实现对电网运行不确定性带来的检修操作风险的量化预判,优化了检修操作规划中的时序决策与风险量化能力,可有效提升检修操作票的准确性。且进一步结合电网实时运行数据和检修任务信息对马尔科夫过程的静态参数进行动态化改造,使得状态、约束等模型参数均可随电网工况变化动态调整,并通过长时序后效性修正弥补传统马尔科夫过程忽略历史操作累积影响的缺陷,避免因短期决策导致的潜在安全隐患。同时,在模型中融入极端场景路径转移和电网边界约束,提升模型对复杂、动态电网检修场景的适配能力,从而使得生成的检修操作序列可准确贴合电网实际运行工况与实际检修任务需求,在保障检修操作票准确性的基础上,提高其现场适配性。
Smart Images

Figure CN122066413B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid maintenance technology, and in particular to a method and system for automatically generating maintenance operation tickets based on Markov processes. Background Technology
[0002] Power grid maintenance operation tickets are the basis for executing power system maintenance operations. The accuracy, security, and dynamic adaptability of their preparation are directly related to the normal conduct of power grid maintenance operations and the stable operation of the power system.
[0003] Currently, the generation of power grid maintenance operation tickets mostly adopts a method of manual compilation combined with traditional rule-based algorithm-assisted planning. The operation framework is determined by manually sorting out the maintenance tasks and basic power grid operation data, and then simple algorithms are used to optimize the operation steps, combined with fixed power grid constraints to complete the compilation of the operation ticket.
[0004] However, most of these methods are based on static equipment topology and fixed operating rules for ticketing, failing to consider the uncertainties such as operating condition fluctuations and measurement noise during power grid operation. They are only suitable for traditional power grid scenarios with simple topologies and single maintenance tasks, and are difficult to adapt to the accurate and time-sequential risk management requirements of maintenance operations in complex operating scenarios. When power grid operating conditions change in real time or face extreme operating scenarios, the prepared maintenance operation tickets are prone to omissions in timing logic, resulting in low accuracy and making it difficult to ensure the safety of maintenance operations. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of automatically generating power grid maintenance operation tickets, which relies on static equipment topology and fixed operating rules, making it difficult to adapt to the accurate time-sequential risk management requirements of maintenance operations in complex operating scenarios, resulting in low accuracy of the generated operation tickets. This invention provides an automatic operation ticket generation method and system based on Markov processes. By introducing Markov processes to optimize the time-sequential decision-making and risk quantification in maintenance operation planning, the accuracy of the generated operation tickets is improved. Furthermore, by combining dynamic modification of the static parameters of the model and correction of long-term aftereffects, its dynamic adaptability to power grid maintenance scenarios is enhanced, ultimately enabling the generated maintenance operation tickets to balance accuracy and on-site adaptability.
[0006] The objective of this invention is achieved through the following technical solution: The method for automatically generating maintenance operation tickets based on Markov processes includes: Based on the basic state and operation timing-dependent state of the target maintenance equipment, a maintenance operation state space is constructed. The initial maintenance status and global target maintenance status are obtained based on real-time power grid operation data and maintenance task information, and the corresponding dynamic constraints are obtained based on the current power grid operating conditions. Using static state transition rules as constraints, combined with constraint priority quantification and conflict handling, dynamic state transition decision weights are generated based on real-time power grid operation data, dynamic constraints, and maintenance task information. Based on the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, a temporary Markov decision model is constructed by combining long-term aftereffect correction, extreme scenario path transition, and power grid boundary constraints. Based on the temporary Markov decision model, the optimal state transition sequence is solved by the hierarchical pruning global search algorithm, and the optimal state transition sequence is mapped to the maintenance operation sequence. Based on a preset operation ticket template, maintenance operation tickets are generated through a maintenance operation sequence.
[0007] By introducing Markov processes and utilizing their temporal transition mechanism of states, actions, and next states, along with their probabilistic decision-making characteristics, a state boundary and temporal dependencies for maintenance operations are constructed within the maintenance operation space, avoiding oversights in the temporal logic of operation steps. Simultaneously, using static state transition rules as constraints, corresponding dynamic state transition decision weights are generated, enabling quantitative prediction of maintenance operation risks arising from grid operation uncertainties. This optimizes the temporal decision-making and risk quantification capabilities in maintenance operation planning, effectively improving the accuracy of maintenance operation tickets. Furthermore, by combining real-time grid operation data and maintenance task information, the static parameters of the Markov process are dynamically modified, allowing model parameters such as states and constraints to be dynamically adjusted according to changes in grid operating conditions. Long-term aftereffect correction compensates for the shortcomings of traditional Markov processes that ignore the cumulative impact of historical operations, ensuring that maintenance operation planning can take into account the long-term benefits of historical operations and avoid potential safety hazards caused by short-term decisions. Meanwhile, extreme scenario path transfers and power grid boundary constraints are incorporated into the model to enhance its adaptability to complex and dynamic power grid maintenance scenarios. This allows the generated maintenance operation sequences to accurately match the actual operating conditions of the power grid and the actual maintenance task requirements, thereby improving its on-site adaptability while ensuring the accuracy of the maintenance operation tickets.
[0008] Furthermore, the construction of the maintenance operation state space based on the basic state and operational timing-dependent state of the target maintenance equipment includes: Extract the open / closed status, energized status, grounded status, and operating mode status of the target maintenance equipment to form a basic status; Based on the maintenance task information and the current power grid operating conditions, the maintenance operation of the target maintenance equipment is obtained. Combined with the time-series dependency relationship of the maintenance operation, the corresponding preceding operation completion identifier and subsequent operation permission condition are encoded into the operation time-series dependency state. Based on the basic state and the operation timing-dependent state, a maintenance operation state space is formed.
[0009] Furthermore, the step of obtaining the initial maintenance status and the global target maintenance status based on real-time power grid operation data and maintenance task information, and obtaining corresponding dynamic constraints based on the current power grid operating conditions, includes: The actual operating status of the target maintenance equipment is determined based on real-time power grid operation data, and the actual operating status is taken as the initial maintenance status. Based on maintenance task information, predict the target stable operating condition of the target maintenance equipment, and take the target stable operating condition as the global target maintenance state. Identify the current power grid operating condition type, and extract the corresponding power flow constraints, voltage constraints, and equipment operation constraints based on the operating condition type to form dynamic constraint conditions.
[0010] Furthermore, the dynamic state transition decision weights are generated based on static state transition rules, combined with constraint priority quantification and conflict handling, and according to real-time power grid operation data, dynamic constraints, and maintenance task information. This includes: Extract constraint parameters associated with state transitions from dynamic constraints, determine the priority level of each constraint parameter by combining maintenance task information, and perform initial quantization and assignment. Remove constraint parameters that do not satisfy the static state transition rules, and identify priority conflicts for the remaining constraint parameters; Priority overrides are applied to constraint parameters that have priority conflicts, or parameter adaptation and adjustment are performed in conjunction with maintenance task information to form an effective constraint set; The current operating scenario is determined based on real-time power grid operation data, and the effective constraint set is adjusted accordingly. Based on the adjusted set of effective constraints, dynamic state transition decision weights are generated by combining real-time operational data deviation values and maintenance task suitability parameters.
[0011] Furthermore, the generation of dynamic state transition decision weights based on the adjusted effective constraint set, combined with real-time operational data deviation values and maintenance task adaptability parameters, includes: Based on static state transition rules, initial maintenance state, and global target maintenance state, several alternative state transition paths are derived. Based on real-time power grid operation data, dynamic constraints, and global target maintenance status, calculate the real-time operation data deviation value and maintenance task adaptability parameter for each alternative state transition path; Based on the adjusted set of effective constraints, the constraint parameters corresponding to each candidate state transition path are matched. The adjusted quantized values of each matched constraint parameter, the real-time running data deviation value, and the maintenance task adaptability parameter are weighted and summed to obtain the dynamic state transition decision weight of each candidate state path.
[0012] Furthermore, the provisional Markov decision model constructed based on the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, combined with long-term aftereffect correction, extreme scenario path transitions, and power grid boundary constraints, includes: The maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights are the basic parameters. Based on the basic parameters, the aftereffect correction factor is calculated through long-term state trajectory, and the path transfer strategy is constructed according to the preset extreme working condition state transfer constraint rules. Based on the basic parameters, a temporary Markov decision model is obtained by combining the aftereffect correction factor, path transition strategy and power grid boundary constraints to form the state transition probability and reward function.
[0013] Furthermore, the state transition probability and reward function formed based on fundamental parameters, combined with aftereffect correction factors, path transition strategies, and grid boundary constraints, includes: Using the dynamic state transition decision weights as the initial baseline values, the time-corrected transition values are obtained by weighting with aftereffect correction factors. Based on the power grid boundary constraints and path transition strategy, each time series correction transition value is adjusted in compliance with regulations to form the state transition probability; The reward value is based on the difference between the initial maintenance state and the global target maintenance state, and the time-series correction reward value is obtained by weighting it with a follow-up correction factor. Based on the power grid boundary constraints and path transfer strategy, the reward value of each time series correction is adjusted by reward and penalty to obtain the reward value after reward and penalty. The reward value after reward and punishment is adjusted by scaling the dynamic state transition decision weights to form a reward function.
[0014] Furthermore, the method based on the temporary Markov decision model, using a hierarchical pruning global search algorithm to solve for the optimal state transition sequence, and mapping the optimal state transition sequence to a maintenance operation sequence, includes: The hierarchical structure of the hierarchical pruning global search algorithm is divided according to the maintenance operation sequence. The initial maintenance state is the search starting point, the global target maintenance state is the search ending point, and the algorithm search path is initialized. Based on the state transition probability and reward function of the temporary Markov decision model, the candidate state transition paths at each level are traversed, and the paths with a transition probability of zero and a cumulative reward lower than the preset reward threshold are pruned in a hierarchical manner. Based on the remaining paths after pruning, the state transition path with the largest cumulative reward value is selected to obtain the optimal state transition sequence. The action set in the optimal state transition sequence is matched and mapped with the preset maintenance operation instruction library. Based on the mapping and matching results, the corresponding equipment operation instruction combination is generated to obtain the optimal maintenance operation sequence.
[0015] Furthermore, the step of generating a maintenance operation ticket based on a preset operation ticket template and a maintenance operation sequence includes: Identify the current maintenance scenario type based on maintenance task information and match the corresponding preset operation ticket template; Map the optimal maintenance operation sequence to the corresponding fields in the template to generate a maintenance operation ticket.
[0016] An automated maintenance operation ticket generation system based on Markov processes, for any of the above-mentioned ticket generation methods, includes: The parameter preprocessing module is used to construct the maintenance operation state space of the target maintenance equipment, extract the corresponding initial maintenance state, global target maintenance state and dynamic constraints, and generate dynamic state transition decision weights. The model building module is used to construct a temporary Markov decision model based on the maintenance operation state space, the initial maintenance state, the global target maintenance state, and the dynamic state transition decision weights. The maintenance operation ticket generation module is used to obtain the maintenance operation sequence based on the temporary Markov decision model and the hierarchical pruning global search algorithm, and generate the corresponding maintenance operation ticket.
[0017] The beneficial effects of this invention are: By introducing Markov processes and leveraging their temporal transition mechanism of states, actions, and next states, along with their probabilistic decision-making characteristics, a state boundary and temporal dependencies for maintenance operations are constructed within the maintenance operation space, avoiding oversights in the temporal logic of operation steps. Simultaneously, using static state transition rules as constraints, corresponding dynamic state transition decision weights are generated, enabling quantitative prediction of maintenance operation risks arising from grid operation uncertainties. This optimizes the temporal decision-making and risk quantification capabilities in maintenance operation planning, effectively improving the accuracy of maintenance operation tickets. Furthermore, by combining real-time grid operation data and maintenance task information, the static parameters of the Markov process are dynamically modified, allowing model parameters such as states and constraints to be dynamically adjusted according to changes in grid operating conditions. Long-term aftereffect correction compensates for the shortcomings of traditional Markov processes that ignore the cumulative impact of historical operations, avoiding potential safety hazards caused by short-term decisions. Meanwhile, extreme scenario path transfers and power grid boundary constraints are incorporated into the model to enhance its adaptability to complex and dynamic power grid maintenance scenarios. This allows the generated maintenance operation sequences to accurately match the actual operating conditions of the power grid and the actual maintenance task requirements, thereby improving its on-site adaptability while ensuring the accuracy of the maintenance operation tickets. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of a process of the present invention; Figure 2 This is a schematic diagram of a dynamic state transition decision weight generation process according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the construction process of a temporary Markov decision model according to an embodiment of the present invention. Detailed Implementation
[0019] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] Example: The maintenance operation ticket is a standardized instruction document used in the maintenance of power equipment to standardize the operation process and ensure the safety and compliance of electrical operations. Its generation needs to be based on the initial operating status of the equipment and the maintenance objectives, and follow electrical safety rules and power grid topology constraints to form a coherent and orderly operation sequence. The selection of each operation depends only on the current state of the equipment and has no direct relationship with the previous operation process, but is only related to the result of the previous operation.
[0021] The Markov process is characterized by its lack of aftereffects, meaning that the future state of the system is determined solely by the current state, independent of historical states and evolutionary paths. This essential characteristic aligns perfectly with the state-dependent nature of maintenance operations. Furthermore, this process can mathematically model the discrete states of equipment and the transition relationships between states. It can describe both deterministic state changes brought about by operations and the stochastic evolution of equipment operating states. Its modeling approach, which relies on state transitions to construct sequential paths, abstracts the maintenance operation ticket generation problem into a state transition path derivation problem, thus meeting the requirements of rule constraints, step sequence, and state correlation in operation ticket generation. Compared to traditional rule-driven algorithms, it enables time-series decision-making and risk quantification for maintenance operation planning, ensuring the accuracy of automatic maintenance operation ticket generation.
[0022] Based on this, this embodiment proposes an automatic maintenance operation ticket generation method based on Markov processes, such as... Figure 1 As shown, it includes: Based on the basic state and operation timing-dependent state of the target maintenance equipment, a maintenance operation state space is constructed. The initial maintenance status and global target maintenance status are obtained based on real-time power grid operation data and maintenance task information, and the corresponding dynamic constraints are obtained based on the current power grid operating conditions. Using static state transition rules as constraints, combined with constraint priority quantification and conflict handling, dynamic state transition decision weights are generated based on real-time power grid operation data, dynamic constraints, and maintenance task information. Based on the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, a temporary Markov decision model is constructed by combining long-term aftereffect correction, extreme scenario path transition, and power grid boundary constraints. Based on the temporary Markov decision model, the optimal state transition sequence is solved by the hierarchical pruning global search algorithm, and the optimal state transition sequence is mapped to the maintenance operation sequence. Based on a preset operation ticket template, maintenance operation tickets are generated through a maintenance operation sequence.
[0023] By introducing Markov processes to optimize the timing decision-making and risk quantification in maintenance operation planning, the accuracy of the generated operation tickets is improved. Furthermore, by combining dynamic modification of the model's static parameters with long-term aftereffect correction, its dynamic adaptability to power grid maintenance scenarios is enhanced, ultimately enabling the generated maintenance operation tickets to balance accuracy and on-site adaptability.
[0024] Markov processes rely on discrete and standardized sets of states for sequential decision-making and state transition deduction. General state definitions cannot match the execution state characteristics of equipment condition switching and operational sequence interlocking in power maintenance operations. Furthermore, traditional maintenance ticketing methods often suffer from fragmented and disjointed state descriptions, making it difficult to support the unified computation and dynamic decision-making of Markov models. Therefore, this paper integrates the basic state of equipment with the temporal dependencies of operations to establish a maintenance operation space, unifying the state definition and expression of the entire maintenance process. This transforms the actual business logic of power maintenance into state elements recognizable by the Markov model, ensuring that the model can deduce operation sequences based on the current state and achieve dynamic automatic generation of maintenance operation tickets.
[0025] The construction of the maintenance operation state space based on the basic state and operation timing dependency state of the target maintenance equipment includes: Extract the open / closed status, energized status, grounded status, and operating mode status of the target maintenance equipment to form a basic status; Based on the maintenance task information and the current power grid operating conditions, the maintenance operation of the target maintenance equipment is obtained. Combined with the time-series dependency relationship of the maintenance operation, the corresponding preceding operation completion identifier and subsequent operation permission condition are encoded into the operation time-series dependency state. Based on the basic state and the operation timing-dependent state, a maintenance operation state space is formed.
[0026] The on / off state, energized state, grounded state, and operating mode state of the target maintenance equipment are extracted. These four states are then discretized and assigned values to form the basic states characterizing the equipment's operating conditions. The on / off state characterizes the on / off attribute of the target maintenance equipment; the energized state characterizes the electrical energized attribute of the target maintenance equipment; the grounded state characterizes the connection status of the grounding switch of the target maintenance equipment; and the operating mode state characterizes the operating, standby, or maintenance working mode of the target maintenance equipment. These four states together constitute the basic state dimensions of the maintenance operation state space.
[0027] Based on maintenance task information, the set of maintenance operation contents is determined. Combined with the current power grid operating conditions, the execution prerequisites for each operation are determined, and the sequential execution constraints between maintenance operations are extracted to obtain time-series dependencies. Furthermore, the preconditions required for each subsequent maintenance operation are defined as the precondition completion identifier, and the conditions for allowing the execution of the corresponding subsequent operation after the preconditions are met are defined as the subsequent operation permission conditions. A unified coding rule is used to numerically encode the precondition completion identifiers and subsequent operation permission conditions for each maintenance operation. The state variables formed by integrating the unified coding of the precondition completion identifiers and subsequent operation permission conditions for each maintenance operation constitute the operation time-series dependent state.
[0028] The acquired basic state is used as the basic state dimension of the state space. The operation timing-dependent state is used as the constraint state dimension and then spliced and merged to form a maintenance operation state space that only contains the equipment's own operating condition characteristics and the constraints of the operation timing.
[0029] Considering that constructing a state space based solely on basic states and operation timing-dependent states can only reflect the operating conditions of a single device and the sequential execution relationship between operations, while power grid maintenance operations are determined by the power grid physical topology and equipment collaborative maintenance rules, the resulting basic maintenance operation state space cannot reflect the constraints of the topology and maintenance rules in addition to the operation content. Therefore, topology mutually exclusive states and timing redundancy constraint states are further introduced on the basis of the formed maintenance operation state space.
[0030] Among them, topological mutual exclusion states originate from the inherent electrical connection relationships of the primary wiring of the power grid. There are natural electrical safety conflicts and state mutual exclusion rules between different devices. These constraints reflect the electrical topological association restrictions between devices, which cannot be covered by the operating conditions or operation sequence of a single device. They must be independently encoded to form topological mutual exclusion states in order to incorporate such physical constraints into the state space. Temporal redundancy constraint states originate from the inherent rules of collaborative maintenance of power equipment. There are redundant operation restrictions in the maintenance process with multi-level temporal interlocking and repeated verification. These constraints limit the rationality of operation execution and the standardization of the process. They are fundamentally different from the simple operation sequence logic and are difficult to express through temporal dependent states. They also need to be separately encoded to form corresponding constraint states.
[0031] Specifically, by reading the primary wiring topology data of the power grid, the electrical connection relationships between the target maintenance equipment and adjacent busbars, branch switches, grounding devices, and other related equipment are traversed to identify operation combinations and state combinations that have electrical safety conflicts, thus obtaining topological exclusion relationships. Equipment state combinations that cannot coexist and operation combinations that cannot be executed synchronously are encoded and expressed to form mutually exclusive topological states, thereby limiting the compliant scope of state transitions during maintenance.
[0032] Based on the equipment type, voltage level, and maintenance method of the target maintenance equipment, the system matches the rules with a pre-defined collaborative maintenance rule library to extract multi-level timing interlocking requirements and operational redundancy verification rules applicable to the equipment. Timing interlocking conditions that need to be met at each level are encoded, as are redundant operation judgment rules that are repeatedly executed or have no actual constraint effect. These rules are then integrated through methods such as splicing to form a timing redundancy constraint state, thereby limiting the execution flow and verification frequency of maintenance operations.
[0033] By incorporating the above two types of states, the maintenance operation state space can cover the equipment's operating conditions, operation timing relationships, electrical topology safety constraints, and maintenance process specification constraints, so as to accurately provide state boundaries that conform to the actual rules of maintenance operations.
[0034] The generation process of the basic state, operation timing-dependent state, topological mutually exclusive state, and timing redundancy constraint state all adopts the same state identification and encoding rules. The specific state identification and encoding rules can be set according to actual needs.
[0035] Considering that the description dimensions of the four states are independent and the constraint information is fragmented, directly combining them to form the maintenance operation state space would lead to inconsistent state definitions, ambiguous state boundaries, and the inability of various maintenance constraints to work synergistically in state determination during subsequent Markov decision-making. This would result in a disconnect between the state transition rules and the actual maintenance constraints. Therefore, we first perform vector concatenation on the basic state, operation timing-dependent state, topology mutually exclusive state, and timing redundant constraint state according to a preset dimension order. This generates fused state data that includes equipment operating conditions, timing constraints, topology constraints, and redundant constraints. A single state vector carries all the constraint information corresponding to the maintenance operation, allowing various constraints to work together to define the state.
[0036] The electrical safety boundaries and maintenance operation specifications of equipment at different voltage levels have inherent differences. The complexity of maintenance operations and the patterns of state changes corresponding to different maintenance intervals also vary significantly. Unclassified fused state data can result in a mixture of different operating condition characteristics, leading to the same state identifier corresponding to multiple maintenance scenarios, causing confusion in state identification and insufficient scenario adaptability. Furthermore, unfiltered fused state data contains a large number of redundant states that represent equivalent operating conditions, have repetitive constraints, or have no practical decision-making significance. These redundant states directly expand the overall size of the state space, increasing the computational burden of constructing state transition relationships and searching for optimal paths in the Markov decision process. At the same time, these redundant states can also interfere with the definition of state intervals, reducing the accuracy of state transition inference and decision-making efficiency.
[0037] Therefore, the acquired fusion status data is further clustered hierarchically according to the maintenance interval and voltage level, and redundant states in the fusion status data are removed based on the hierarchical clustering results, thereby forming the final maintenance operation status space.
[0038] First, voltage level and maintenance interval are pre-classified. Voltage level is set as a three-level clustering standard of high voltage, medium voltage and low voltage, and maintenance interval is set as a two-level clustering standard of regular maintenance, temporary maintenance and condition maintenance. Based on the hierarchical clustering algorithm, the fused status data is classified step by step. First, the voltage level is used as the first-level clustering index to divide the fused status data into status data clusters of the corresponding voltage level. Then, the maintenance interval is used as the second-level clustering index to further split the data clusters of each voltage level to form status data subsets of subdivided scenarios.
[0039] Vector-wise comparisons are performed on the state data subsets completed by each cluster. State data with the same encoded components and consistent constraint meanings are marked as equivalent redundant states. Encoding differences between vectors are identified. State data with only differences in equivalent representation, duplicate marking, or encoding order among the four types of constraints, and whose differences do not affect the execution permission determination of maintenance operations, the compliance determination of state transitions, and the verification conclusion of safety constraints, are marked as invalid redundant states. All marked equivalent redundant states and invalid redundant states are traversed and eliminated, and valid state data with independent encoding characteristics and differences in constraint expression are retained.
[0040] The valid state data after removing redundancy from each subset are uniformly collected, and a unique state number is assigned to each group of valid state data. A structured state set containing the state number, encoding vector, corresponding clustering scenario and constraint features is established. This structured state set is the final maintenance operation state space.
[0041] The real-time operation data of the power grid is constantly changing. The initial operating conditions of the target maintenance equipment, such as switching on / off, energizing, and grounding, will change with the real-time operation status. However, the maintenance operation state space is only a static constraint set that integrates equipment operating conditions, timing, topology, and redundancy verification rules. It is difficult to adapt to the dynamic actual scenario of power grid operation and maintenance work. If the static maintenance operation state space is relied upon alone, the starting state of the maintenance process cannot be accurately determined. This can easily lead to a mismatch between the initial state and the actual on-site operating conditions, which in turn causes the starting point of the subsequent state transition deviates from the actual operating conditions.
[0042] Furthermore, different maintenance tasks correspond to different operational objectives, and the final maintenance status, outage scope, and recovery requirements of the equipment vary. The constructed static maintenance operation state space does not actually include these operational objectives, resulting in a lack of termination criteria for the subsequent Markov decision-making process, making it impossible to form a complete and task-compliant maintenance operation sequence. Simultaneously, the current operating conditions of the power grid include dynamic information such as real-time load levels, main and distribution network interaction methods, and tie switch status. This information generates real-time safety constraints and operational limitations, which are also not included in the static maintenance operation state space. Without extracting these dynamic constraints in conjunction with real-time operating conditions, the generated maintenance operation sequence is prone to conflicting with the real-time operational safety requirements of the power grid, failing to meet the dynamic compliance requirements of on-site operations.
[0043] Therefore, by further obtaining the initial maintenance state and the global target maintenance state based on the real-time operation data of the power grid and the maintenance task information, the static state space is bound to the real-time operation of the power grid and the maintenance task objectives, providing clear start and end states for the Markov process. Based on the static constraints of the static state space, corresponding dynamic constraints are added based on the real-time power grid operating conditions, providing accurate boundary conditions for subsequent maintenance state transitions and operation sequence generation based on the state space.
[0044] The step of obtaining the initial maintenance status and the global target maintenance status based on real-time power grid operation data and maintenance task information, and obtaining corresponding dynamic constraints based on the current power grid operating conditions, includes: The actual operating status of the target maintenance equipment is determined based on real-time power grid operation data, and the actual operating status is taken as the initial maintenance status. Based on maintenance task information, predict the target stable operating condition of the target maintenance equipment, and take the target stable operating condition as the global target maintenance state. Identify the current power grid operating condition type, and extract the corresponding power flow constraints, voltage constraints, and equipment operation constraints based on the operating condition type to form dynamic constraint conditions.
[0045] The system collects real-time measurement information and equipment status monitoring data uploaded from the power grid. This data includes the switching position, energized attributes, grounding status, operating mode, and real-time operating information such as associated node voltage and branch power flow of the target maintenance equipment. The real-time data is matched and mapped with the coding rules of the basic states in the maintenance operation state space. The values of each basic state dimension of the equipment are calibrated item by item through the real-time data to obtain the actual operating state that can truly reflect the equipment's operating condition before the maintenance starts. This state is the initial maintenance state for subsequent maintenance state transition simulation and the starting state point of the Markov decision process.
[0046] The maintenance task information is used to extract details such as the maintenance target, maintenance type, maintenance scope, and post-maintenance recovery requirements. This extraction can be achieved through keyword targeting and other methods. Then, combined with power grid operation rules and equipment maintenance operation rules, the stable operating conditions that the target maintenance equipment should meet after all maintenance operations are completed are deduced. This operating condition is the final fixed state without temporary transitional operations or temporary safety measures, including the final configuration results of equipment connection / disconnection, energization, grounding, and operating mode. This target stable operating condition is set as the global target maintenance state, serving as the basis for determining the termination of subsequent state transitions, thus ensuring that the maintenance operation sequence has a termination indication.
[0047] Based on real-time grid load levels, main and distribution network interaction methods, tie switch status, and power supply transfer methods, the current operating conditions are categorized into types such as normal operation, peak load, tie-switching, and maintenance-restricted operation. For different operating condition types, corresponding power flow constraints, voltage constraints, and equipment operation constraints are retrieved from a preset constraint library. Power flow constraints limit the allowable range of active and reactive power flow in each branch; voltage constraints limit the acceptable range of voltage amplitude at each topology node; and equipment operation constraints limit the types of equipment operations prohibited or restricted under the current operating conditions.
[0048] By integrating the above three types of constraints, dynamic constraints are formed. In the subsequent state transition process, dynamic constraints can effectively eliminate states and operation paths that conflict with the current power grid operating conditions, thus ensuring the reliability of the formed operation sequence.
[0049] Power grid maintenance operation scenarios mainly include routine maintenance scenarios and extreme maintenance scenarios. In routine maintenance scenarios, although the power grid operating conditions are relatively stable, there are still real-time changes such as load fluctuations, minor adjustments to the status of tie switches, and slight adjustments to the maintenance scope. Static state transition rules and fixed decision weights are only based on preset rules and fixed topology relationships, and cannot adaptively adjust to these subtle changes in operating conditions. This leads to deviations between state transition decisions and the current actual operating conditions, affecting the execution efficiency and accuracy of the maintenance operation sequence, and even causing redundancy in the operation process and slight disconnection from real-time constraints. In extreme maintenance scenarios such as peak load operation, limited power transfer between main and distribution networks, critical voltage limits at nodes, and parallel cross-maintenance of multiple devices, the power grid safety margin is extremely low. The conflicts between dynamic power flow constraints, voltage safety constraints, field equipment operation restrictions, and static maintenance sequence rules are significantly amplified. If state transition decisions are made directly based on static state transition rules, not only will the decision results be disconnected from the actual operating conditions of the power grid, but the decision results may also prioritize secondary constraints, violating the safety constraints under the current operating conditions, thereby causing safety accidents such as power grid overruns and live-line misoperation.
[0050] Therefore, by combining real-time power grid operation data, dynamic constraints, and maintenance task information, the priority of constraints is quantified and conflicts are handled. Based on the optimized constraints, corresponding dynamic state transition decision weights are generated. This approach can adjust the decision logic according to subtle changes in operating conditions in routine maintenance scenarios, ensuring the accuracy and efficiency of decisions and adapting to the stable operation requirements of routine maintenance. In extreme maintenance scenarios, it can coordinate multiple constraint conflicts, limit the priority of safety constraints, adapt to the special characteristics of extreme operating conditions, ensure the safety and feasibility of state transition decisions, and achieve comprehensive adaptation to both routine and extreme maintenance scenarios. This avoids the problem of unreasonable decisions and unexecuted operation sequences in different scenarios due to fixed static rules and fixed decision weights.
[0051] Among them, such as Figure 2 As shown, the dynamic state transition decision weights are generated based on static state transition rules, combined with constraint priority quantification and conflict handling, and according to real-time power grid operation data, dynamic constraints, and maintenance task information. This includes: Extract constraint parameters associated with state transitions from dynamic constraints, determine the priority level of each constraint parameter by combining maintenance task information, and perform initial quantization and assignment. Remove constraint parameters that do not satisfy the static state transition rules, and identify priority conflicts for the remaining constraint parameters; Priority overrides are applied to constraint parameters that have priority conflicts, or parameter adaptation and adjustment are performed in conjunction with maintenance task information to form an effective constraint set; The current operating scenario is determined based on real-time power grid operation data, and the effective constraint set is adjusted accordingly. Based on the adjusted set of effective constraints, dynamic state transition decision weights are generated by combining real-time operational data deviation values and maintenance task suitability parameters.
[0052] The dynamic constraints are the power flow constraints, voltage constraints, and equipment operation constraints identified above. Only some constraint parameters in the dynamic constraints are directly related to equipment state transitions and operation execution decisions during maintenance. To improve subsequent decision-making efficiency, constraint parameters related to state transitions, such as branch power flow limits, node voltage allowable ranges, and equipment operation prohibition conditions, are first selected. Based on the maintenance object, scope, safety control level, and operational objectives of the maintenance task, and combined with weighted summation scoring of indicators, the relevant constraint parameters are assigned priority levels. Specifically, power flow and voltage limit exceedance constraints that ensure safe grid operation are set as high priority; timing interlocking constraints that follow maintenance rules are set as medium priority; and redundancy verification constraints that optimize operation procedures are set as low priority. Initial quantization is performed according to the preset fixed numerical ranges corresponding to each priority level, forming a set of constraint parameters with priority weights.
[0053] Static state transition rules are basic state transition criteria pre-set based on maintenance operation rules and the primary topology of the power grid. They constitute the fundamental constraints for the execution of maintenance operations. Each priority-weighted constraint parameter is compared and verified against the static state transition rules. Constraint parameters that violate the basic criteria or fail to meet the static constraint requirements are directly eliminated, retaining only those that comply with the rules. The retained constraint parameters are then paired for verification to identify priority conflicts arising from contradictory execution conditions or overlapping constraint ranges. If different constraint parameters impose mutually exclusive restrictions on the same state transition path, or if different priority constraints yield conflicting compliance judgments for the same maintenance operation, a priority conflict is identified.
[0054] For conflicts between identified high-priority safety constraints and low-priority process constraints, the high-priority overriding method is used to handle the conflict. For conflicts between identified constraints of the same priority, the corresponding parameter limits are adjusted in conjunction with the maintenance task to handle the conflicting constraint relationships.
[0055] After conflict resolution, the constraint parameters that are conflict-free, conform to the static state transition rules, and are suitable for the current maintenance task are integrated into a valid constraint set to limit the constraint execution logic, avoid constraint conflicts in Markov process decisions, and ensure the accuracy and reliability of subsequent state transition decisions.
[0056] Based on the formation of an effective constraint set, real-time power grid load, power flow distribution, and node voltage data are collected. This real-time data is compared with preset operating condition thresholds to determine the current operating condition scenario. According to the safety margins and constraint requirements of different operating conditions, the constraint parameters in the effective constraint set are dynamically adjusted. In extreme operating conditions such as limited power transfer or voltage criticality, the limits of safety constraint parameters such as power flow and voltage are reduced to increase constraint strength. Under normal and stable operating conditions, the normal limits of constraint parameters are maintained, balancing maintenance efficiency and grid operation safety, thus ensuring that the effective constraint set matches the current actual grid operating conditions.
[0057] Relying solely on the priority and limits of constraints can only determine whether a state transition is compliant. If fixed decision weights are still used, even when combined with an optimized set of effective constraints, it is still difficult to reflect the real-time safety margin of the power grid and maintenance objectives, and it cannot adapt to fluctuations in operating conditions. Therefore, based on the effective constraint set, real-time operational data deviations and maintenance task suitability are incorporated to form dynamic state transition decision weights that can balance the real-time operational safety of the power grid with the execution objectives of maintenance tasks.
[0058] The step of generating dynamic state transition decision weights based on the adjusted effective constraint set, combined with real-time operational data deviation values and maintenance task adaptability parameters, includes: Based on static state transition rules, initial maintenance state, and global target maintenance state, several alternative state transition paths are derived. Based on real-time power grid operation data, dynamic constraints, and global target maintenance status, calculate the real-time operation data deviation value and maintenance task adaptability parameter for each alternative state transition path; Based on the adjusted set of effective constraints, the constraint parameters corresponding to each candidate state transition path are matched. The adjusted quantized values of each matched constraint parameter, the real-time running data deviation value, and the maintenance task adaptability parameter are weighted and summed to obtain the dynamic state transition decision weight of each candidate state path.
[0059] Dynamic state transition decision weights are essentially comprehensive quantitative evaluation values oriented towards state transition paths. They rely on specific state transition paths as the calculation carrier and cannot be generated independently of the paths. Therefore, starting with the initial maintenance state and ending with the global target maintenance state, all state transition methods conforming to static state transition rules are traversed within the maintenance operation state space. Several candidate state transition paths that can gradually transition from the starting point to the ending point are selected, obtaining a set of compliant paths that meet the requirements of basic maintenance procedures, topological constraints, and start and end conditions. This set serves as the basis for calculating dynamic state transition decision weights, effectively ensuring the effectiveness and reliability of subsequent dynamic state transition decision weights. Simultaneously, by forming a limited set of candidate paths through pre-deduction, the scope of weight calculation is limited to valid paths, avoiding indiscriminate traversal calculations across the entire state space, reducing computational redundancy in constraint matching, deviation value, and fitness calculations, and improving weight generation efficiency.
[0060] Based on the power flow and voltage constraints in the dynamic constraints, real-time power grid measurement data corresponding to each state node in each alternative state transition path is extracted. The absolute or relative deviation between the real-time measurement values and the constraint limits is calculated, and the overall deviation value of each alternative state transition path is obtained, i.e., the real-time operating data deviation value. This quantifies the degree to which each node on the path deviates from the safety constraint boundary, reflecting the real-time safety level of the path. Furthermore, based on the global target maintenance status and maintenance task requirements, the matching degree between the execution steps and state transition process of each alternative state transition path and the task requirements is compared. Specifically, this can be calculated using preset quantitative formulas, such as the proportion of adaptation items and deviation deductions, to obtain the adaptation quantification value of each alternative state transition path, thereby quantifying the degree of fit between the path and the maintenance task target.
[0061] Next, the complete execution process of each candidate state transition path is matched one by one with the adjusted set of valid constraints. All constraint parameters that the path must satisfy at each state node are selected, and the quantified values of these constraint parameters after working condition adjustment are extracted, i.e., priority quantification values, to ensure that constraint priority is integrated into the path evaluation. Subsequently, according to a preset weight allocation ratio, the matched constraint parameter quantification values, real-time operating data deviation values, and maintenance task adaptability parameters are weighted and summed to finally obtain the dynamic state transition decision weight for each candidate path.
[0062] By weighted summing of constraint parameter quantization values, real-time operational data deviation values, and maintenance task suitability parameters, reverse quantization of real-time operational data deviation values can be achieved. That is, the smaller the deviation value, the higher the reverse quantized value, thus transforming it into a safety evaluation index that can participate in weighted calculations. This index maps the real-time safety margin of the power grid to a weight correction factor, optimizing the shortcomings of traditional Markov processes with fixed weights, which cannot respond to real-time operating condition changes and are difficult to avoid safety risks. Simultaneously, the maintenance task suitability parameter can be used as an indicator reflecting the degree of fit between the path and the maintenance task objective, guiding the decision weights of the Markov process to converge towards the global target maintenance state. Optimizing fixed weights lacks task orientation and is prone to detours in state transitions. The constraint parameter quantization values, however, ensure that weight optimization does not deviate from the safety boundary, guaranteeing the reliability of dynamic state transition decision weights, and ultimately obtaining dynamic state transition decision weights that balance constraint compliance, real-time power grid safety levels, and maintenance task objective orientation.
[0063] Although the acquired dynamic state transition decision weights are pre-calculated based on the effective constraint set, real-time operating data deviation values, and maintenance task suitability parameters for alternative state transition paths, their essence is to provide a unified value measurement standard for the Markov optimization process that fits real-time operating conditions and balances constraint priorities and task objectives. During subsequent Markov model optimization, the model does not directly copy alternative paths but uses these dynamic weights as the evaluation basis to perform value assessment, iterative comparison, and global ranking of each state transition branch in the state space. In routine maintenance scenarios, referring to these weights can quickly eliminate redundant branches, guiding the optimization process towards an efficient and concise path. In extreme maintenance scenarios, these weights also embed safety constraint priorities and conflict handling logic, allowing the optimization process to avoid high-risk transition branches and ensuring the reliability and safety of the optimization path.
[0064] After optimizing the state space, initial state, decision weights, and constraints of the maintenance operation, the necessary state carrier, decision start and end benchmarks, constraint basis, and dynamic value evaluation criteria for maintenance state transition decisions have been formed. Further integration of these elements will be used to construct a temporary Markov decision model for the current maintenance task, providing a foundation for iterative optimization of the optimal maintenance path in the future.
[0065] However, this integration still results in a traditional, general-purpose Markov decision model framework. While its lack of aftereffects allows for simplified modeling of simple maintenance operation tickets, in complex maintenance scenarios with long-term temporal coupling and multiple operating conditions, the model's lack of aftereffects can lead to neglecting the long-term cumulative impact of maintenance operations on grid conditions, resulting in locally optimal but subsequently invalid paths that exceed limits. Furthermore, when facing extreme operating conditions in complex maintenance scenarios, although extreme operating condition optimization has been incorporated into the decision weight optimization, this type of general-purpose model only supports conventional state transition logic. Due to the extremely low grid safety margin in extreme operating conditions, all conventional state transition paths will fail. Even if the dynamic state transition weights indicate a safe direction, the model lacks corresponding emergency transition rules to implement the jump, making it unable to find a feasible path or resulting in a path with significant safety risks.
[0066] In addition, the dynamic constraints used in building the model are actually flexible optimization constraints that calculate decision weights and coordinate constraint priorities. They are used for path value scoring and conflict coordination, and do not have mandatory restriction attributes. When the model seeks optimization, it may still flexibly break through these constraints in order to pursue high-scoring paths, thus solving for invalid paths with significant safety risks.
[0067] Therefore, based on the basic model architecture built from the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, the temporary Markov decision model is further combined with long-term aftereffect correction, extreme scenario path transition, and power grid boundary constraints. The resulting model retains the modeling simplification advantage brought by no aftereffect while making up for its inherent defects in long-term decision-making, extreme scenario adaptation, and safety boundary control. Ultimately, the constructed temporary Markov decision model can adapt to the maintenance needs of complex maintenance scenarios.
[0068] Among them, such as Figure 3 As shown, the temporary Markov decision model constructed based on the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, combined with long-term aftereffect correction, extreme scenario path transition, and power grid boundary constraints, includes: The maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights are the basic parameters. Based on the basic parameters, the aftereffect correction factor is calculated through long-term state trajectory, and the path transfer strategy is constructed according to the preset extreme working condition state transfer constraint rules. Based on the basic parameters, a temporary Markov decision model is obtained by combining the aftereffect correction factor, path transition strategy and power grid boundary constraints to form the state transition probability and reward function.
[0069] The maintenance operation state space is a collection of all equipment operating states, maintenance operation nodes, and power grid topology states during power grid maintenance, and it serves as the state carrier of the model.
[0070] Using the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights as the basic parameters of the temporary Markov decision model, the maintenance operation state space limits the range of state values in the model, determines all possible maintenance operation states and potential transition relationships between states, and provides the model with a complete state traversal and optimization boundary, avoiding state overflow or optimization range deviation. The initial maintenance state is used as the starting point for model optimization, ensuring that the model starts from the current actual maintenance state and fits the real-time maintenance field of the power grid. The global target maintenance state is used as the endpoint for model optimization, providing decision guidance and ensuring that all optimization processes revolve around achieving the global maintenance goal, avoiding aimless state transitions. The dynamic state transition decision weights serve as the model's value evaluation standard, providing a quantitative basis for subsequent state transition probability calculations and reward functions, ensuring that the model can balance the real-time safety level of the power grid, constraint compliance, and maintenance task objectives.
[0071] By combining historical transition data and real-time operation data of each state in the maintenance operation state space, the long-term operation trajectory of each alternative state transition path obtained from previous simulations is simulated. The long-term operation trajectory includes the operation parameters, maintenance operation steps and constraint changes of each state node at each time point of the corresponding path, as well as the time series and resource consumption sequence of the transition between each node.
[0072] Based on various data from long-term operational trajectories, this study analyzes the long-term impact of each path on three dimensions throughout its entire lifecycle: power grid safe operation, maintenance task progress, and constraint compliance. The quantitative indicators for the long-term impact on power grid safe operation include cumulative operational deviation, peak operational deviation, duration of operational deviation, and equipment failure risk trend. The quantitative indicators for the long-term impact on maintenance task progress include cumulative maintenance progress deviation and quantitative maintenance resource consumption. The quantitative indicators for the long-term impact on constraint compliance include the number of violations, duration of violations, severity of violations, and constraint compliance rate. The long-term cumulative effect of the three dimensions is calculated by weighting the quantitative values using long-term impact accumulation coefficients and deviation transmission coefficients, thereby obtaining a post-effect correction factor to quantify the long-term value differences of different state transition paths.
[0073] Simultaneously, considering the operational characteristics and maintenance rules of the power grid under extreme conditions, pre-defined state transition constraint rules for extreme conditions are established, such as prohibiting transitions in high-risk states, escalating priority constraints, and prioritizing emergency maintenance paths. Based on these rules, corresponding path transition strategies for extreme scenarios are constructed to limit executable state transition branches, evasive state transition branches, and preferred state transition branches under extreme scenarios, ensuring that the model can respond quickly and avoid safety risks in extreme scenarios. After obtaining the path transition strategies for extreme scenarios, they are integrated with the model's own conventional path transition rules to form the final path transition strategy.
[0074] Based on the aforementioned basic parameters, combined with the calculated aftereffect correction factor, the constructed path transition strategy, and the power grid boundary constraints, the state transition probability and reward function of the temporary Markov decision model are further formed, thus completing the model construction.
[0075] Among them, based on the basic parameters, and combined with the aftereffect correction factor, path transition strategy, and power grid boundary constraints, the state transition probability and reward function are formed, including: Using the dynamic state transition decision weights as the initial baseline values, the time-corrected transition values are obtained by weighting with aftereffect correction factors. Based on the power grid boundary constraints and path transition strategy, each time series correction transition value is adjusted in compliance with regulations to form the state transition probability; The reward value is based on the difference between the initial maintenance state and the global target maintenance state, and the time-series correction reward value is obtained by weighting it with a follow-up correction factor. Based on the power grid boundary constraints and path transfer strategy, the reward value of each time series correction is adjusted by reward and penalty to obtain the reward value after reward and penalty. The reward value after reward and punishment is adjusted by scaling the dynamic state transition decision weights to form a reward function.
[0076] First, taking into account the deviation of real-time power grid operation data, the adaptability of maintenance tasks, and the quantitative values of constraint parameters, the initial quantitative dynamic state transition decision weights that reflect the value of the path are used as the initial benchmark value. Then, the subsequent effect correction factor is used as the corresponding correction weight. The initial benchmark value is weighted and calculated to obtain the time-series corrected transition value. In this way, the long-term impact of the entire life cycle of the path is integrated into the calculation of the transition value. The correction of the initial benchmark value only focuses on the limitations of short-term operating conditions, ensuring that the transition value can take into account both short-term compliance and long-term value.
[0077] By combining power grid boundary constraints and preset path transition strategies, the obtained time-series corrected transition values are verified and adjusted for compliance. For time-series corrected transition values that comply with boundary constraints and are compatible with the path transition strategy, their quantization characteristics are retained. For time-series corrected transition values that do not comply with constraints or conflict with the transition strategy, the quantization amplitude is adjusted according to the corresponding verification anomaly type, or the corresponding values of the non-compliant branches are removed. Finally, a quantization value that meets all constraints and matches the path transition strategy is obtained, i.e., the state transition probability. That is, the higher the dynamic state transition decision weight, the better the aftereffect, and the higher the state transition probability of the transition branch that complies with the path transition strategy and boundary constraints.
[0078] The difference between the initial maintenance state and the global target maintenance state is used as the basic reward value. The greater the difference, the lower the basic reward value, reflecting the value orientation of the state transition toward the target. The smaller the difference, the higher the basic reward value, so as to give positive value quantification to the transition to the target state.
[0079] Based on the basic reward value, and with the subsequent effect correction factor as the corresponding correction weight, the basic reward value is weighted and calculated to incorporate the long-term impact of the path into the reward calculation. This avoids the reward focusing only on short-term approach to the goal while ignoring long-term grid safety and maintenance efficiency, and ensures that the reward value can reflect the long-term value of the path.
[0080] By combining power grid boundary constraints and path transfer strategies, the reward values for each time series correction are adjusted through rewards and penalties. For time series correction reward values that meet the boundary constraints and are compatible with the path transfer strategy, the corresponding reward value is increased, and the increase can be set according to actual needs. For time series correction reward values that do not meet the constraints or conflict with the transfer strategy, a reverse penalty is applied, such as reducing the reward value.
[0081] Furthermore, the calculation objects for the state transition probability and the specific reward value are both the state transition branches on each candidate state transition path.
[0082] For the reward value after the penalty, it is scaled to a preset range according to the quantization range of the dynamic state transition decision weights to avoid the model optimization deviation caused by the reward value being too high or too low. All reward values after amplitude scaling are integrated, and with the specific state transition branch as input and the scaled reward value as output, a one-to-one mapping relationship is established. The final reward function is formed by the reward value mapping relationship of each state transition branch.
[0083] Finally, using the maintenance operation state space as the corresponding state set and action set, and combining the constructed state transition probabilities and reward functions, a final temporary Markov decision model is obtained. Here, the transitions between different states in the maintenance operation state space are the state transition branches of the Markov process, i.e., actions. The resulting temporary Markov decision model is not a fixed, static model, but rather a temporary decision-making vehicle that can dynamically update the state transition probabilities and reward functions based on real-time data, constraint adjustments, and changes in operating conditions, flexibly adapting to the needs of different maintenance scenarios. In the subsequent model optimization process, the model will use the state transition probability as the basis and the reward function as the optimization direction to iteratively optimize within the maintenance operation state space. This can avoid the shortcomings of traditional models whose fixed decision logic cannot respond to real-time changes in operating conditions. It can also balance short-term maintenance efficiency and long-term power grid safety through aftereffect correction and extreme scenario strategies. At the same time, it relies on power grid boundary constraints to ensure the compliance of the optimization path. Ultimately, it achieves the optimization of the optimal state transition path from the initial maintenance state to the global target maintenance state, ensuring the accuracy and reliability of the final output maintenance operation sequence, and thus ensuring the accuracy and reliability of the generated maintenance operation tickets.
[0084] Considering that the state set of the constructed temporary Markov decision model is defined by the maintenance operation state space, which includes all possible maintenance operation states and transition branches between states, if a conventional global search is used, all state branches need to be calculated and traversed one by one, resulting in a huge amount of computation and a large number of invalid branches. Therefore, a hierarchical pruning global search algorithm is further adopted to realize the optimization iteration of the temporary Markov decision model, reduce the amount of computation, avoid redundant calculations, improve the solution efficiency, and fit the dynamic update characteristics of the model.
[0085] The method of using a temporary Markov decision model to solve for the optimal state transition sequence through a hierarchical pruning global search algorithm, and mapping the optimal state transition sequence to a maintenance operation sequence, includes: The hierarchical structure of the hierarchical pruning global search algorithm is divided according to the maintenance operation sequence. The initial maintenance state is the search starting point, the global target maintenance state is the search ending point, and the algorithm search path is initialized. Based on the state transition probability and reward function of the temporary Markov decision model, the candidate state transition paths at each level are traversed, and the paths with a transition probability of zero and a cumulative reward lower than the preset reward threshold are pruned in a hierarchical manner. Based on the remaining paths after pruning, the state transition path with the largest cumulative reward value is selected to obtain the optimal state transition sequence. The action set in the optimal state transition sequence is matched and mapped with the preset maintenance operation instruction library. Based on the mapping and matching results, the corresponding equipment operation instruction combination is generated to obtain the optimal maintenance operation sequence.
[0086] Starting with the initial maintenance state as layer 1, i.e., the search starting point, the maintenance stages corresponding to each adjacent state transition branch are divided into layers according to the order of maintenance operations, until the global target maintenance state is reached, i.e., the search endpoint. Each layer corresponds to a maintenance stage, containing all possible states and transition branches (actions) of that stage. The layers progress sequentially, gradually transitioning from the initial maintenance state to the global target maintenance state, ensuring the adaptability of the hierarchical structure to the model's state transition logic and actual maintenance needs.
[0087] Starting from the initial maintenance state, all possible initial search branches are preset, i.e., the transition branches in the model action set, thus forming an initial search path set. Moreover, the paths in the initial search path set are not the complete candidate state transition paths previously pre-deduced, but rather initial branches and subsequent potential extension directions extracted from the candidate state transition paths, starting from the initial maintenance state, which are candidate paths to be screened.
[0088] Based on the state transition probabilities and reward functions of a temporary Markov decision model, candidate paths at each level are traversed one by one. During the traversal, paths with zero state transition probabilities and paths with accumulated rewards below a preset reward threshold are pruned. These paths are considered invalid paths with no possibility of transition or too low value, and therefore do not require further search. Furthermore, the pruning operation is performed hierarchically, rather than pruning all paths at once. The traversal and pruning of the first level are completed first, and then, based on the remaining branches after pruning at the first level, the corresponding transition branches of the second level are traversed, and so on, until all levels are pruned. This layered pruning not only aligns with the actual maintenance operation process but also accurately removes invalid branches, avoids computational redundancy, and improves solution efficiency.
[0089] After pruning at all levels, the cumulative reward value of the reward function is used as the screening criterion to select the state transition path with the largest cumulative reward value from the remaining paths after pruning. This path is the optimal state transition sequence from the initial maintenance state to the global target maintenance state.
[0090] The optimal state transition sequence is essentially composed of a series of state transition branches. Each state transition branch, i.e., the action, in the sequence is matched and mapped with a preset maintenance operation instruction library. Based on the maintenance operation requirements corresponding to each action, the corresponding equipment operation instruction is matched from the instruction library. Finally, all the matched instructions are integrated to form a complete optimal maintenance operation sequence.
[0091] After optimizing the maintenance operation sequence, it is further combined with the preset operation ticket template to convert it into a maintenance operation ticket, thereby improving maintenance efficiency and ensuring the safe operation of the power grid.
[0092] The step of generating a maintenance operation ticket based on a preset operation ticket template and a maintenance operation sequence includes: Identify the current maintenance scenario type based on maintenance task information and match the corresponding preset operation ticket template; Map the optimal maintenance operation sequence to the corresponding fields in the template to generate a maintenance operation ticket.
[0093] Keyword matching is used to extract relevant keyword information from maintenance task information, thereby determining the current maintenance scenario type, such as routine maintenance, emergency repair, switching operation, etc., and then extracting a matching preset operation ticket template. The obtained optimal maintenance operation sequence is then filled into the associated columns of the operation steps in the preset operation ticket template to form the final maintenance operation ticket.
[0094] Another aspect of this embodiment provides an automatic maintenance operation ticket generation system based on Markov processes, including: The parameter preprocessing module is used to construct the maintenance operation state space of the target maintenance equipment, extract the corresponding initial maintenance state, global target maintenance state and dynamic constraints, and generate dynamic state transition decision weights. The model building module is used to construct a temporary Markov decision model based on the maintenance operation state space, the initial maintenance state, the global target maintenance state, and the dynamic state transition decision weights. The maintenance operation ticket generation module is used to obtain the maintenance operation sequence based on the temporary Markov decision model and the hierarchical pruning global search algorithm, and generate the corresponding maintenance operation ticket.
[0095] The parameter preprocessing module is connected to the model building module, and the model building module is connected to the maintenance operation ticket generation module.
[0096] The parameter preprocessing module, model building module, and maintenance operation ticket generation module are all data processing devices such as computers with corresponding data processing capabilities, and are equipped with external communication interfaces to extract the required data from external platforms.
[0097] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Other variations and modifications are possible without departing from the technical solutions described in the claims.
Claims
1. A method for automatically generating maintenance operation tickets based on Markov processes, characterized in that, include: Based on the basic state and operation timing-dependent state of the target maintenance equipment, a maintenance operation state space is constructed. The initial maintenance status and global target maintenance status are obtained based on real-time power grid operation data and maintenance task information, and the corresponding dynamic constraints are obtained based on the current power grid operating conditions. Using static state transition rules as constraints, combined with constraint priority quantification and conflict handling, dynamic state transition decision weights are generated based on real-time power grid operation data, dynamic constraints, and maintenance task information. Extract constraint parameters associated with state transitions from dynamic constraints, determine the priority level of each constraint parameter by combining maintenance task information, and perform initial quantization and assignment. Remove constraint parameters that do not satisfy the static state transition rules, and identify priority conflicts for the remaining constraint parameters; Priority overrides are applied to constraint parameters that have priority conflicts, or parameter adaptation and adjustment are performed in conjunction with maintenance task information to form an effective constraint set; The current operating scenario is determined based on real-time power grid operation data, and the effective constraint set is adjusted accordingly. Based on static state transition rules, initial maintenance state, and global target maintenance state, several alternative state transition paths are derived. Based on real-time power grid operation data, dynamic constraints, and global target maintenance status, calculate the real-time operation data deviation value and maintenance task adaptability parameter for each alternative state transition path; Based on the adjusted set of effective constraints, the constraint parameters corresponding to each candidate state transition path are matched. The adjusted quantized values of each matched constraint parameter, the deviation values of real-time operating data, and the maintenance task adaptability parameters are weighted and summed to obtain the dynamic state transition decision weight of each candidate state path. Based on the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, a temporary Markov decision model is constructed by combining long-term aftereffect correction, extreme scenario path transition, and power grid boundary constraints. Based on the temporary Markov decision model, the optimal state transition sequence is solved by the hierarchical pruning global search algorithm, and the optimal state transition sequence is mapped to the maintenance operation sequence. Based on a preset operation ticket template, maintenance operation tickets are generated through a maintenance operation sequence.
2. The method for automatically generating maintenance operation tickets based on Markov processes according to claim 1, characterized in that, The maintenance operation state space is constructed based on the basic state and operational timing dependency state of the target maintenance equipment, including: Extract the open / closed status, energized status, grounded status, and operating mode status of the target maintenance equipment to form a basic status; Based on the maintenance task information and the current power grid operating conditions, the maintenance operation of the target maintenance equipment is obtained. Combined with the time-series dependency relationship of the maintenance operation, the corresponding preceding operation completion identifier and subsequent operation permission condition are encoded into the operation time-series dependency state. Based on the basic state and the operation timing-dependent state, a maintenance operation state space is formed.
3. The method for automatically generating maintenance operation tickets based on Markov processes according to claim 1, characterized in that, The process of obtaining the initial maintenance status and global target maintenance status based on real-time power grid operation data and maintenance task information, and obtaining corresponding dynamic constraints based on the current power grid operating conditions, includes: The actual operating status of the target maintenance equipment is determined based on real-time power grid operation data, and the actual operating status is taken as the initial maintenance status. Based on maintenance task information, predict the target stable operating condition of the target maintenance equipment, and take the target stable operating condition as the global target maintenance state. Identify the current power grid operating condition type, and extract the corresponding power flow constraints, voltage constraints, and equipment operation constraints based on the operating condition type to form dynamic constraint conditions.
4. The method for automatically generating maintenance operation tickets based on Markov processes according to claim 1, characterized in that, The temporary Markov decision model, constructed based on the maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights, combined with long-term aftereffect correction, extreme scenario path transitions, and power grid boundary constraints, includes: The maintenance operation state space, initial maintenance state, global target maintenance state, and dynamic state transition decision weights are the basic parameters. Based on the basic parameters, the aftereffect correction factor is calculated through long-term state trajectory, and the path transfer strategy is constructed according to the preset extreme working condition state transfer constraint rules. Based on the basic parameters, a temporary Markov decision model is obtained by combining the aftereffect correction factor, path transition strategy and power grid boundary constraints to form the state transition probability and reward function.
5. The method for automatically generating maintenance operation tickets based on Markov processes according to claim 4, characterized in that, The process of forming the state transition probability and reward function based on fundamental parameters, combined with aftereffect correction factors, path transition strategies, and power grid boundary constraints, includes: Using the dynamic state transition decision weights as the initial baseline values, the time-corrected transition values are obtained by weighting with aftereffect correction factors. Based on the power grid boundary constraints and path transition strategy, each time series correction transition value is adjusted in compliance with regulations to form the state transition probability; The reward value is based on the difference between the initial maintenance state and the global target maintenance state, and the time-series correction reward value is obtained by weighting it with a follow-up correction factor. Based on the power grid boundary constraints and path transfer strategy, the reward value of each time series correction is adjusted by reward and penalty to obtain the reward value after reward and penalty. The reward value after reward and punishment is adjusted by scaling the dynamic state transition decision weights to form a reward function.
6. The method for automatically generating maintenance operation tickets based on Markov processes according to claim 1, characterized in that, The method based on the temporary Markov decision model, which uses a hierarchical pruning global search algorithm to solve for the optimal state transition sequence and maps the optimal state transition sequence to a maintenance operation sequence, includes: The hierarchical structure of the hierarchical pruning global search algorithm is divided according to the maintenance operation sequence. The initial maintenance state is the search starting point, the global target maintenance state is the search ending point, and the algorithm search path is initialized. Based on the state transition probability and reward function of the temporary Markov decision model, the candidate state transition paths at each level are traversed, and the paths with a transition probability of zero and a cumulative reward lower than the preset reward threshold are pruned in a hierarchical manner. Based on the remaining paths after pruning, the state transition path with the largest cumulative reward value is selected to obtain the optimal state transition sequence. The action set in the optimal state transition sequence is matched and mapped with the preset maintenance operation instruction library. Based on the mapping and matching results, the corresponding equipment operation instruction combination is generated to obtain the optimal maintenance operation sequence.
7. The method for automatically generating maintenance operation tickets based on Markov processes according to claim 1, characterized in that, The process of generating maintenance operation tickets based on a preset operation ticket template and a maintenance operation sequence includes: Identify the current maintenance scenario type based on maintenance task information and match the corresponding preset operation ticket template; The optimal maintenance operation sequence is mapped and filled into the corresponding fields of the template to generate a maintenance operation ticket.
8. An automatic maintenance operation ticket generation system based on Markov processes, used to execute the ticket generation method according to any one of claims 1 to 7, characterized in that, include: The parameter preprocessing module is used to construct the maintenance operation state space of the target maintenance equipment, extract the corresponding initial maintenance state, global target maintenance state and dynamic constraints, and generate dynamic state transition decision weights. The model building module is used to construct a temporary Markov decision model based on the maintenance operation state space, the initial maintenance state, the global target maintenance state, and the dynamic state transition decision weights. The maintenance operation ticket generation module is used to obtain the maintenance operation sequence based on the temporary Markov decision model and the hierarchical pruning global search algorithm, and generate the corresponding maintenance operation ticket.
Citation Information
Patent Citations
Power transformation maintenance decision optimization method and device, storable medium and computing equipment
CN114493041A
Power grid equipment time-phased power failure strategy self-adaptive generation system and method based on GRPO algorithm
CN121860809A