OCS system operation mode adjustment method and device, terminal equipment and storage medium
By constructing a simulation model of the OCS system and determining the model based on preset operating modes, the optimal operating adjustment sequence is automatically determined, which solves the problem of low adjustment efficiency caused by reliance on human experience in existing technologies and achieves efficient elimination of the impact of overlapping conflicts in power outage sections.
Patent Information
- Application Number
- CN202511738444.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies rely on human experience to determine scheduling schemes that eliminate the conflict effects caused by overlapping power outage sections, resulting in low adjustment efficiency.
By acquiring the topology data, equipment parameters, and operating condition data of the OCS system, an initial simulation model is constructed. The model is then used to determine candidate operation adjustment sequences and their Q values using preset operating modes. Finally, the optimal operation adjustment sequence is determined to achieve automated adjustment.
This greatly improves adjustment efficiency, reduces reliance on human experience, and enhances the accuracy and efficiency of adjustments.
Smart Images

Figure CN121566631A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system dispatch automation technology, and in particular to a method, apparatus, terminal equipment and storage medium for adjusting the operation mode of an OCS system. Background Technology
[0002] The OCS system, or integrated dispatch automation system, is an important tool in the power system for unified dispatch and management of power generation, transmission, and distribution. Its purpose is to improve the operating efficiency of the power grid, reduce operating costs, and enhance the reliability and security of power supply. In the power system, the phenomenon of overlapping power outage sections refers to the situation where multiple power outage areas overlap due to maintenance, faults, or natural disasters under certain specific circumstances. This situation has a significant impact on power dispatch and power supply reliability.
[0003] Overlapping power outage sections mean that multiple constraints are coupled together. Traditionally, the operation adjustment scheme to eliminate the conflict caused by overlapping power outage sections is determined by human experience. However, this method is time-consuming and has a high error rate, so it has the problem of low adjustment efficiency. Summary of the Invention
[0004] This invention provides a method, apparatus, terminal device, and storage medium for adjusting the operation mode of an OCS system, which can solve the problem of low adjustment efficiency in the prior art, which relies on manual determination of the scheduling scheme to eliminate the conflict caused by overlapping power outage sections.
[0005] An embodiment of the present invention provides a method for adjusting the operating mode of an OCS system, comprising: Obtain the current topology data, equipment parameters, operating status data, and planned power outage information of the OCS system to be adjusted; Based on the topology data, equipment parameters, operating condition data, and planned power outage information, an initial OCS system simulation model is constructed. Based on the initial OCS system simulation model and the preset operation mode, a model is determined, and several candidate operation adjustment operation sequences and the Q value corresponding to each candidate operation adjustment operation sequence are determined. Based on the Q value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode, the model is determined, and the optimal operation adjustment sequence is determined. The operating mode of the OCS system to be adjusted is adjusted according to the optimal operating adjustment sequence.
[0006] Furthermore, the step of determining several candidate operation adjustment sequences and the Q-value corresponding to each candidate operation adjustment sequence based on the initial OCS system simulation model and the preset operation mode determination model includes: Obtain the first initial simulation topology data and the first initial simulation operating condition data of the initial OCS system simulation model; Based on the first initial simulation topology data, the first initial simulation operating condition data, the initial OCS system simulation model, and the preset operating mode determination model, the first operating mode adjustment operation is repeatedly executed to obtain the candidate operating mode adjustment operation sequence, and the final Q value corresponding to the last operating mode adjustment operation in each candidate operating mode adjustment operation sequence. The final Q value is used as the Q value for each candidate run adjustment operation sequence; The first operation mode adjustment operation includes: Obtain the current first simulation topology data and the current first simulation running condition data of the current first OCS system simulation model; wherein, the initial first OCS system simulation model is the initial OCS system simulation model, the initial first simulation topology data is the first initial simulation topology data, and the initial first simulation running condition data is the first initial simulation running condition data; Based on the current first simulation running condition data, the current first simulation topology data, and the preset running mode, the model is determined, and several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount are obtained. The simulation model of the first OCS system is updated according to each current candidate operation to obtain the current first simulation topology data and the current first simulation running condition data corresponding to each current candidate operation. Based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, the current target candidate operation is determined from several current candidate operations, and the current power outage section coupling constraint model is constructed. If the first objective function of the current power outage section coupling constraint model does not converge, the current first constraint violation exceeds the preset range, or any first constraint violation exceeds the preset range in the previous N-1 consecutive iterations, the updated first OCS system simulation model corresponding to the current target candidate operation is used as the first OCS system simulation model at the next moment, the current first simulation topology data corresponding to the current target candidate operation is used as the first simulation topology data at the next moment, and the current first simulation operating condition data corresponding to the current target candidate operation is used as the first simulation operating condition data at the next moment. Otherwise, the operation sequence corresponding to each current candidate operation is taken as the candidate operation adjustment sequence.
[0007] Furthermore, the step of determining the model based on the current first simulation running condition data, the current first simulation topology data, and the preset running mode, and obtaining several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount, includes: Based on the current first simulation running condition data, the current first constraint violation amount is calculated; The current first simulation topology data, the current first simulation running condition data, and the current first constraint violation are input into the preset running mode determination model to obtain several current candidate operations and the initial Q value of each current candidate operation.
[0008] Furthermore, the preset operating mode determines the model construction, including: Acquire several historical operation scheduling data of the OCS system to be adjusted under different historical operation scenarios; wherein, the historical operation scenarios include: normal operation scenario, overload operation scenario and dynamic instability operation scenario; the historical operation scheduling data includes: historical operation adjustment operation, historical operation condition data corresponding to the historical operation adjustment operation, and topology data corresponding to the historical operation adjustment operation. Based on the historical operating data, the historical reward value and the corresponding historical constraint violation amount for each historical operating adjustment operation are calculated. The historical operation scheduling data, historical reward value, and historical constraint violation amount corresponding to each historical operation adjustment operation are used as a set of candidate training sample data, and then several sets of candidate training sample data are obtained. From several sets of candidate training sample data, several sets of selected training sample data are obtained by sampling according to the preset historical operation scenario sampling ratio; Several sets of selected training sample data are input into the model to be trained for iterative training until the loss function converges, thereby generating the preset operating mode determination model; In each iteration of training, the model is determined based on the current operating mode and the currently selected training sample data to obtain the current predicted Q value and the current target Q value; the current loss function is calculated based on the current predicted Q value and the current target Q value, and it is determined whether the current loss function has converged; if it has converged, the model determined by the current operating mode is used as the preset operating mode model; otherwise, the model parameters in the current operating mode model are adjusted and training continues.
[0009] Furthermore, determining the current target candidate operation from a plurality of current candidate operations based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the device parameters includes: Based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, a coupling constraint model of the current power outage section is constructed with the goal of minimizing the power outage range, load loss, and equipment overload risk. Based on the first current constraint condition of the current power outage section coupling constraint model, generate the current dynamic mask for each current candidate operation; Based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q value, determine the current final Q value of each candidate operation; The candidate operation corresponding to the largest current final Q value is taken as the current target candidate operation.
[0010] Furthermore, determining the current final Q-value of each candidate operation based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q-value includes: For each current candidate operation, the current feasibility score of each current candidate operation is calculated based on the current first simulation running condition data corresponding to the current candidate operation and the corresponding first current constraint condition. Update the initial Q value of the current candidate operation whose current feasibility score is less than the preset score threshold to negative infinity; Based on all updated initial Q values and all unupdated initial Q values, construct the current Q value matrix; Based on the current Q-value matrix and the current dynamic mask, the current final Q-value of each candidate operation is calculated.
[0011] Furthermore, determining the optimal operation adjustment sequence based on the Q-value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model includes: The candidate run adjustment operation sequence corresponding to the maximum Q value is selected as the run adjustment operation sequence; Based on the initial OCS system simulation model, obtain the second OCS system simulation model after executing the last operation adjustment in the selected operation adjustment sequence, the second initial simulation topology data of the second OCS system simulation model, and the second initial simulation operation condition data; Based on the second OCS system simulation model and the preset operating mode, the model is determined by repeatedly executing the second operating mode adjustment operation until the second objective function of the current power outage section coupling constraint model converges and the second constraint violation amount of N consecutive iterations does not exceed the preset range, thus obtaining the optimal operating adjustment operation sequence. The second operating mode adjustment operation includes: Obtain the current second simulation topology data and the current second simulation running condition data of the current second OCS system simulation model; wherein, the initial OCS system simulation model is the second OCS system simulation model after the last running adjustment operation in the selected running adjustment operation sequence is executed, the initial second simulation topology data is the second initial simulation topology data, and the initial second simulation running condition data is the second initial simulation running condition data; Based on the current second simulation running condition data, the current second simulation topology data, and the preset running mode, the model is determined, and several current candidate operations, the initial Q value of each current candidate operation, and the current second constraint violation amount are obtained. The simulation model of the second OCS system is updated according to each current candidate operation to obtain the current second simulation topology data and the current second simulation operating condition data corresponding to each current candidate operation. Based on the current second simulation topology data corresponding to each current candidate operation, the current second simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, the current target candidate operation is determined from several current candidate operations, and the current power outage section coupling constraint model is constructed. If the second objective function of the current power outage section coupling constraint model does not converge, the current second constraint violation exceeds the preset range, or any second constraint violation exceeds the preset range in the previous N-1 consecutive iterations, the updated second OCS system simulation model corresponding to the current target candidate operation is used as the second OCS system simulation model at the next moment, the current second simulation topology data corresponding to the current target candidate operation is used as the second simulation topology data at the next moment, and the current second simulation operating condition data corresponding to the current target candidate operation is used as the second simulation operating condition data at the next moment. Otherwise, the operation sequence corresponding to the current target candidate operation is taken as the optimal operation adjustment sequence.
[0012] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments; This invention provides an OCS system operation mode adjustment device, comprising: The module includes a data acquisition module, a simulation model construction module, a Q-value calculation module, an optimal operation adjustment sequence filtering module, and an operation mode adjustment module. The data acquisition module is used to acquire the current topology data, equipment parameters, operating condition data, and planned power outage information of the OCS system to be adjusted; The simulation model construction module is used to construct an initial OCS system simulation model based on the topology data, equipment parameters, operating condition data, and planned power outage information. The Q-value calculation module is used to determine a number of candidate operation adjustment sequences and the Q-value corresponding to each candidate operation adjustment sequence based on the initial OCS system simulation model and the preset operation mode. The optimal operation adjustment sequence filtering module is used to determine the optimal operation adjustment sequence based on the Q value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model. The operation mode adjustment module is used to adjust the operation mode of the OCS system to be adjusted according to the optimal operation adjustment sequence.
[0013] Based on the above method embodiments, the present invention provides a corresponding terminal device embodiment; The present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the operation mode adjustment method of an OCS system described in any embodiment of the present invention.
[0014] Based on the above method embodiments, the present invention provides a corresponding storage medium embodiment; The present invention provides a storage medium including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the operation mode adjustment method of an OCS system described in any embodiment of the present invention.
[0015] The embodiments of the present invention have the following beneficial effects: This invention provides a method, apparatus, terminal device, and storage medium for adjusting the operation mode of an OCS system. The method includes: acquiring the current topology data, equipment parameters, operating condition data, and planned power outage information of the OCS system to be adjusted; subsequently, constructing an initial OCS system simulation model based on the topology data, equipment parameters, operating condition data, and planned power outage information; then, determining several candidate operation adjustment operation sequences and a Q-value corresponding to each candidate operation adjustment operation sequence based on the initial OCS system simulation model and a preset operation mode determination model; then, determining an optimal operation adjustment operation sequence based on the Q-value corresponding to each candidate operation adjustment operation sequence, the initial OCS system simulation model, each candidate operation adjustment operation sequence, and the preset operation mode determination model; and finally, adjusting the operation mode of the OCS system to be adjusted according to the optimal operation adjustment operation sequence. Therefore, the entire process of this invention does not require the use of human experience to obtain operation adjustment operations, greatly reducing the possibility of low adjustment efficiency caused by reliance on manual methods. Attached Figure Description
[0016] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a method for adjusting the operation mode of an OCS system according to an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of the structure of an OCS system operation mode adjustment device provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0021] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0023] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0024] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0025] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0026] See Figure 1To address the problem of low efficiency in existing technologies that rely on manual determination of scheduling schemes to eliminate conflicts caused by overlapping power outage sections, this invention provides a method for adjusting the operation mode of an OCS system, comprising: Step S101: Obtain the current topology data, equipment parameters, operating condition data, and planned power outage information of the OCS system to be adjusted; Specifically, the aforementioned topology data includes the number of devices and their connections in the entire OCS system to be adjusted; the aforementioned device parameters are the power parameters of each power device itself, such as the maximum allowable power angle deviation of the generator and the maximum allowable voltage of the device; the aforementioned operating condition data includes the real-time load level of the OCS system to be adjusted during operation (e.g., the active and reactive power injection of each node), line power flow distribution (active and reactive power flow values of each line), node voltage amplitude (actual value of voltage at each node), system frequency (overall operating frequency of the power grid), and the current operating mode of the power grid equipment (output status of generators, tap position of transformers, open / closed status of switches, etc.), used to reflect the current operating status and stability of the power grid; the aforementioned planned power outage information mainly includes the outage area (the specific substations, lines, and equipment involved) used to clarify the scope of the outage, the outage time (the start and end times of the outage), the reason for the outage (equipment maintenance, upgrades, etc.), and the equipment numbers involved in the outage (line numbers, transformer numbers, etc.).
[0027] Step S102: Based on the topology data, equipment parameters, operating condition data, and planned power outage information, construct an initial OCS system simulation model; Specifically, simulations are performed using topology data, equipment parameters, operating condition data, and planned power outage information to obtain an OCS system with the same state as the current OCS system to be adjusted. The simulation process is based on existing technology and will not be described in detail here.
[0028] Step S103: Determine the model based on the initial OCS system simulation model and the preset operation mode, and determine several candidate operation adjustment operation sequences and the Q value corresponding to each candidate operation adjustment operation sequence; Specifically, the preset operating mode determines that the model is a deep Q-network. When the corresponding state is input, a set of action spaces and the corresponding Q value (i.e., the initial Q value in this invention) will be obtained.
[0029] In a preferred embodiment, the step of determining a plurality of candidate operation adjustment sequence and the Q value corresponding to each candidate operation adjustment sequence based on the initial OCS system simulation model and the preset operation mode determination model includes: Obtain the first initial simulation topology data and the first initial simulation operating condition data of the initial OCS system simulation model; Based on the first initial simulation topology data, the first initial simulation operating condition data, the initial OCS system simulation model, and the preset operating mode determination model, the first operating mode adjustment operation is repeatedly executed to obtain the candidate operating mode adjustment operation sequence, and the final Q value corresponding to the last operating mode adjustment operation in each candidate operating mode adjustment operation sequence. The final Q value is used as the Q value for each candidate run adjustment operation sequence; The first operation mode adjustment operation includes: Obtain the current first simulation topology data and the current first simulation running condition data of the current first OCS system simulation model; wherein, the initial first OCS system simulation model is the initial OCS system simulation model, the initial first simulation topology data is the first initial simulation topology data, and the initial first simulation running condition data is the first initial simulation running condition data; Based on the current first simulation running condition data, the current first simulation topology data, and the preset running mode, the model is determined, and several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount are obtained. The simulation model of the first OCS system is updated according to each current candidate operation to obtain the current first simulation topology data and the current first simulation running condition data corresponding to each current candidate operation. Based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, the current target candidate operation is determined from several current candidate operations, and the current power outage section coupling constraint model is constructed. Specifically, the power outage section coupling constraint model forms a multi-objective optimization problem by defining the constraint conditions under overlapping power outage sections. Based on topology data, equipment parameters, operating conditions, and planned power outage information, a set of components affected by the power outage operation is abstracted and defined as power outage sections. The physical boundaries and interaction relationships of the main and distribution network power outage sections are clarified. Combined with protection configuration and operating rules, key constraint-related nodes in the overlapping section scenario are defined, forming the basic constraint conditions. From three dimensions—static safety, dynamic stability, and protection coordination—the constraint coupling effect caused by overlapping power outage sections is quantified, and the nonlinear constraint relationships between multiple sections are described. The constraint conditions corresponding to the power outage section coupling constraint model are constructed, and the corresponding objective function is constructed with the optimization objectives of minimizing the power outage range, load loss, and equipment overload risk.
[0030] If the first objective function of the current power outage section coupling constraint model does not converge, the current first constraint violation exceeds the preset range, or any first constraint violation exceeds the preset range in the previous N-1 consecutive iterations, the updated first OCS system simulation model corresponding to the current target candidate operation is used as the first OCS system simulation model at the next moment, the current first simulation topology data corresponding to the current target candidate operation is used as the first simulation topology data at the next moment, and the current first simulation operating condition data corresponding to the current target candidate operation is used as the first simulation operating condition data at the next moment. Otherwise, the operation sequence corresponding to each current candidate operation is taken as the candidate operation adjustment sequence.
[0031] Specifically, in each iteration, the preset operating mode determines the model, which generates initial Q-values for all current actions based on the current state of the first OCS system simulation model. Then, based on all current actions, it determines the actual operation the system needs to execute, marking it as the current target candidate operation. The updated system state is then obtained based on this target candidate operation. This closed-loop process dynamically adjusts the search direction through state feedback, ensuring that each exploration step satisfies relevant conditions, gradually narrowing the optimal solution search range. As the iteration progresses, the convergence of constraint violations and the objective function is continuously monitored. When N consecutive state updates do not trigger new constraint violations (i.e., the system has stabilized in a safe state), and the objective function value tends to stabilize (fluctuating within a very small threshold), the solution space is determined to have converged to a locally optimal region satisfying all safety constraints, thus terminating the iteration. Furthermore, the search path only covers the effective region within the constraint subspace, avoiding ineffective exploration in high-dimensional continuous control problems and significantly improving solution efficiency and reliability.
[0032] In this preferred embodiment, based on the initial OCS system simulation model and the preset operation mode, several candidate operation adjustment operation sequences and the Q value corresponding to each candidate operation adjustment operation sequence are determined.
[0033] In another preferred embodiment, the step of determining the model based on the current first simulation running condition data, the current first simulation topology data, and the preset running mode, and obtaining several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount, includes: Based on the current first simulation running condition data, the current first constraint violation amount is calculated; Specifically, the constraint violation amount is determined based on the number of violations of static safety constraints and dynamic stability constraints in the power outage section coupled constraint model. That is, if both static safety constraints and dynamic stability constraints are violated simultaneously, the constraint violation amount is 2; if either static safety constraint or dynamic stability constraint is violated, the constraint violation amount is 1; and if neither is violated, the constraint violation amount is 0. The static safety constraints mainly focus on power flow distribution and equipment capacity limitations, and their expression is: In the formula, Indicates static safety constraints. Indicates the power flow of the line. Indicates the maximum capacity of the line. Indicates node voltage. Indicates the upper limit of the node voltage. This indicates the lower limit of the node voltage.
[0034] Specifically, dynamic stability constraints mainly focus on frequency fluctuations and power angle instability risks during transient processes, which can be expressed as: In the formula, This represents a dynamic stability constraint. Indicates frequency deviation. Indicates the maximum permissible frequency deviation. Indicates the difference in work angle. This indicates the maximum permissible power angle difference.
[0035] The current first simulation topology data, the current first simulation running condition data, and the current first constraint violation are input into the preset running mode determination model to obtain several current candidate operations and the initial Q value of each current candidate operation.
[0036] Specifically, the first simulation topology data, the first simulation operating condition data, and the first constraint violation quantity are transformed into a three-dimensional structure, which serves as the state vector of the model to determine the preset operating mode. This structure is then fused through a fully connected layer and an LSTM timing module within the model to capture the spatiotemporal coupling characteristics of the power grid operation. The output is an initial Q-value matching the action space dimension, along with the corresponding action (i.e., the operation in this invention). The action space covers common adjustment strategies in power grid operation, forming a discrete set including operations such as topology reconfiguration, load transfer, generator output regulation, and reactive power compensation equipment switching. A hierarchical encoding method is used: the first layer is the operation type (e.g., switch opening / closing, transformer tap adjustment); the second layer is the operation object (e.g., specific line or bus number); and the third layer is the operation magnitude (e.g., load transfer ratio or generator output increment). The output action space is dynamically constrained through a dynamic masking mechanism, excluding inoperable switches based on the current topology or filtering out over-limit actions based on equipment capacity limitations.
[0037] In this preferred embodiment, based on the current first simulation running condition data, the current first simulation topology data, and the model determined by the preset running mode, several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount are obtained.
[0038] In another preferred embodiment, the preset operating mode determines the construction of the model, including: Acquire several historical operation scheduling data of the OCS system to be adjusted under different historical operation scenarios; wherein, the historical operation scenarios include: normal operation scenario, overload operation scenario and dynamic instability operation scenario; the historical operation scheduling data includes: historical operation adjustment operation, historical operation condition data corresponding to the historical operation adjustment operation, and topology data corresponding to the historical operation adjustment operation. Based on the historical operating data, the historical reward value and the corresponding historical constraint violation amount for each historical operating adjustment operation are calculated. Specifically, the calculation method for historical constraint violations is the same as that for the first constraint violation, and will not be repeated here. Historical reward values are calculated based on a reward function, which comprehensively evaluates the improvement effects of actions on grid security, economy, and efficiency. A weighted summation method is used to quantify multi-objective conflicts. The reward function includes constraint satisfaction rewards, adjustment cost penalties, and operational efficiency rewards. Constraint satisfaction rewards are divided into two parts: static security rewards are calculated by reducing line overload rates and voltage exceedances, while dynamic stability rewards are designed based on frequency deviation and power angle stability indicators. Adjustment cost penalties cover the number of switching operations, generator regulation amplitude, and reactive power equipment switching frequency, reflecting equipment losses and operational costs in actual operation. Operational efficiency rewards are achieved by improving line utilization, reducing network losses, and optimizing power transmission paths. A dynamic adjustment mechanism is used for weight allocation, adaptively updating according to the real-time operating status of the grid to avoid local optima caused by fixed weights. The reward function accumulates future returns through discount factors, guiding the model to learn long-term optimal strategies to achieve coordinated optimization of grid operation constraints and scheduling objectives.
[0039] Specifically, the expression for the reward based on the degree of constraint satisfaction is: In the formula, This represents the reward value for the degree of constraint satisfaction at time t. This represents the value of the static security reward at time t. This represents the value of the dynamic stable reward at time t. , , and Indicates the weighting coefficient. This represents the state vector at time t. Represents the action space at time t. Represents a set of routes. This represents the power of the i-th line. This represents the maximum allowable power of the i-th line. Represents a set of nodes. This represents the voltage at node j. This represents the maximum allowable voltage at node j. Indicates frequency deviation. Indicates the maximum permissible frequency deviation. Represents a set of generators. This indicates the power angle deviation of generator k. This represents the maximum permissible power angle deviation of generator k.
[0040] Specifically, the expression for adjusting the cost penalty is: In the formula, This indicates the value of the adjustment cost penalty. This represents the penalty value for the number of switch operations. This represents the value of the generator regulation range penalty. This represents the value of the frequency penalty for reactive power equipment switching. This indicates the penalty weight for the switching operation. Indicates the number of switch operations. This indicates the penalty weight for generator regulation. Indicates generator Output adjustment amount, This represents the maximum allowable adjustment of generator k. This indicates the penalty weight for switching reactive power equipment. This indicates the number of times reactive power equipment is switched on and off.
[0041] Specifically, the expression for the performance bonus is: In the formula, This represents the value of the performance bonus. The reward weighting represents the line utilization rate. This represents the value of the reward for improving line utilization. This represents the value of the power transmission path optimization reward. This indicates the reward weight for network loss. This represents the network loss of line i. Indicates total power. This represents the reward weight for power transmission path optimization.
[0042] In summary, the expression for the entire reward function is: The historical operation scheduling data, historical reward value, and historical constraint violation amount corresponding to each historical operation adjustment operation are used as a set of candidate training sample data, and then several sets of candidate training sample data are obtained. Specifically, a set of candidate training sample data includes four data points: power grid state, executed action, immediate reward, and next state. Therefore, this set of data can be stored as a standardized experience quadruple in the experience replay memory pool. Preferably, the proportion of historical running scenarios in the experience replay memory pool can be analyzed periodically, a hierarchical sampling strategy can be used to balance the category distribution, a duplicate storage mechanism can be set for low-frequency, high-value samples, and old data can be periodically eliminated and new samples can be introduced to maintain the diversity and timeliness of data distribution.
[0043] From several sets of candidate training sample data, several sets of selected training sample data are obtained by sampling according to the preset historical operation scenario sampling ratio; Specifically, the preset historical operating scenario sampling ratio can be set as follows: 60% normal operating scenarios, 30% overload operating scenarios, and 10% dynamic instability operating scenarios. After extracting several sets of selected training sample data, their dimensional differences are eliminated and normalized to a unified numerical range, forming standardized batch data that can be directly input into the model. During normalization, continuous variables are normalized to [-1, 1], discrete actions use One-Hot encoding, and historical reward values are compressed to a unified range through quantile standardization. Based on this, operating scenario labels can be added to each quadruple as an additional dimension to embed the data, supporting differentiated loss calculations for different operating conditions in constraint-aware training. Dynamic weights are assigned to samples under different operating conditions based on the scenario labels, strengthening learning in extreme scenarios.
[0044] Several sets of selected training sample data are input into the model to be trained for iterative training until the loss function converges, thereby generating the preset operating mode determination model; In each iteration of training, the model is determined based on the current operating mode and the currently selected training sample data to obtain the current predicted Q value and the current target Q value; the current loss function is calculated based on the current predicted Q value and the current target Q value, and it is determined whether the current loss function has converged; if it has converged, the model determined by the current operating mode is used as the preset operating mode model; otherwise, the model parameters in the current operating mode model are adjusted and training continues.
[0045] Specifically, after the selected training sample data is input into the model to be trained, the model will calculate the predicted Q value in the corresponding state and generate the target Q value of the next state through its internal target network. The loss function is constructed using the temporal difference error between the two, and the model parameters are updated through the backpropagation algorithm. The model parameters are iteratively optimized until the loss function converges stably.
[0046] Specifically, after the model outputs the predicted Q-value for the current state, it inputs the next state data into the target network, calculates its maximum Q-value, and combines it with the immediate reward to generate the corresponding target Q-value. In the formula, Indicates the target Q value. This represents the instantaneous reward at time t (i.e., the historical reward value mentioned above). This represents a discount factor used to weigh the importance of current rewards against future rewards. Indicates the predicted Q value, This represents the state vector at time t+1. This represents the action space corresponding to time t+1. This represents the parameters of the target network at time t+1. Indicates the target network's next state All possible actions The maximum Q-value estimate.
[0047] Specifically, taking the mean squared error as the loss function, the expression for the loss function is: In the formula, This represents the value of the loss function. This represents the expectation, which is the average of the sample loss. This represents the parameters of the target network at time t.
[0048] Specifically, during training, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters, and the model parameters are updated by gradient descent to gradually narrow the gap between the predicted value and the target value. The network parameters are iteratively optimized, and the changes in the loss function are monitored until the loss function converges (for example, the loss function value is less than a preset threshold after 50 consecutive iterations).
[0049] Preferably, to stabilize the training process, a fixed interval of steps is set every 1000 gradient updates to completely copy the model parameters to the target network, avoiding training divergence caused by drastic fluctuations in the target Q value. After synchronization, the target network keeps the parameters frozen and is only used to generate a stable target Q value before the next synchronization.
[0050] Preferably, during the model training process, the reward function accumulates future returns through a discount factor, guiding the model to learn the long-term optimal strategy, so as to achieve the coordinated optimization of power grid operation constraints and scheduling objectives.
[0051] In this preferred embodiment, a preset operation mode determination model is trained using several historical operation scheduling data of the OCS system to be adjusted under different historical operation scenarios.
[0052] In another preferred embodiment, determining the current target candidate operation from a plurality of current candidate operations based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the device parameters includes: Based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, a coupling constraint model of the current power outage section is constructed with the goal of minimizing the power outage range, load loss, and equipment overload risk. Specifically, by analyzing topology data, the set of components affected by power outage operations is identified and abstracted into power outage sections. The physical boundaries between the main grid and distribution network power outage sections are clarified, and the interaction relationship between the two at locations such as tie switches and distributed power supply access points is defined. In conjunction with the configuration of protection devices (differential protection, overcurrent protection, etc.) and operating rules (islanding operation restrictions), key constraint-related nodes in the case of overlapping sections are identified, including components that may cause protection maloperation or over-limit due to multi-section coupling.
[0053] Subsequently, the constraint coupling effect caused by overlapping power outage sections was quantified from three dimensions: static safety, dynamic stability, and protection coordination. Among them, the static safety dimension focuses on changes in power flow distribution, analyzes the superimposed impact of simultaneous operation of multiple sections on line load rate and node voltage deviation, and describes nonlinear constraint relationships. The dynamic stability dimension considers frequency fluctuations and power angle instability risks caused by section switching during transient processes. The protection coordination dimension combines protection settings and section topology to identify risk points of false tripping / failure to trip caused by overlapping protection ranges.
[0054] In summary, the constraint conditions for the power outage section coupling constraint model are as follows: In the formula, The constraints represent the coupled constraint model of the power outage section. express Its main focus is on the risk of false activation or failure to activate caused by overlapping protection ranges, so as to ensure that the protection device will not falsely activate or fail to activate.
[0055] Specifically, minimizing the power outage area, load loss, and equipment overload risk are taken as optimization objectives. To address the conflicting characteristics between different objectives, a weighted sum method is used for normalization, integrating multiple objectives into a single comprehensive optimization objective. The resulting objective function of the power outage section coupling constraint model is: In the formula, This represents the value of the objective function. This represents the weighting coefficient corresponding to the power outage area, where f is the quantified value of the power outage area. It can be quantified by the number of affected nodes within the power outage area or the area of the power outage area. The smaller the value, the better. This represents the weighting coefficient corresponding to the load loss, where 's' represents the quantified value of the load loss. This value can be quantified by the total unmet load within the power outage area; a smaller value is better. Let z represent the weighting coefficient corresponding to the equipment overload risk, z represent the quantified value of the equipment overload risk, and b represent the equipment b. This indicates the power of device b. This indicates the maximum allowable power of device b. Indicates equipment The degree of overload.
[0056] Based on the first current constraint condition of the current power outage section coupling constraint model, generate the current dynamic mask for each current candidate operation; Specifically, by encoding the constraints, the corresponding dynamic mask is obtained.
[0057] Based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q value, determine the current final Q value of each candidate operation; The candidate operation corresponding to the largest current final Q value is taken as the current target candidate operation.
[0058] Preferably, the candidate operation with the highest final Q value is selected for execution, and the power grid state is updated. This closed-loop process dynamically adjusts the search direction through state feedback to ensure that each step of exploration meets real-time safety conditions, gradually narrowing the search range of the optimal solution. As the iteration progresses, the convergence of constraint violations and objective function is continuously monitored. Its search path only covers the effective region within the constraint subspace, avoiding invalid exploration in high-dimensional continuous control problems and significantly improving solution efficiency and reliability.
[0059] In this preferred embodiment, the current target candidate operation is determined from a number of current candidate operations based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the device parameters.
[0060] In another preferred embodiment, determining the current final Q value of each candidate operation based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q value includes: For each current candidate operation, the current feasibility score of each current candidate operation is calculated based on the current first simulation running condition data corresponding to the current candidate operation and the corresponding first current constraint condition. Specifically, the formula for calculating the feasibility score is as follows: In the formula, Indicates the state at time t. Next, candidate actions The corresponding feasibility score, , ,and All represent weighting coefficients. This represents the static safety constraint score at time t, used to evaluate candidate actions. Whether static safety constraints are met, i.e., evaluating candidate actions. Will this lead to line overload or node voltage exceeding limits? The dynamic stability constraint score at time t is used to evaluate candidate actions. Whether the dynamic stability constraint is satisfied, i.e., evaluating candidate actions. Will this lead to frequency deviation or power angle instability? The protection coordination constraint score at time t is used to evaluate candidate actions. Whether the protection coordination constraints are met, i.e., evaluating candidate actions. Will this cause the protection device to malfunction or fail to operate? Representing state Downline The trend Representing state Next node voltage, Representing state The frequency deviation below, Representing state The status of the lower protection device.
[0061] Update the initial Q value of the current candidate operation whose current feasibility score is less than the preset score threshold to negative infinity; Specifically, as can be seen from the above formula for calculating the feasibility score, the calculation process is based on the constraints of the power outage section coupling constraint model. Through this calculation process, the feasibility score may be 0, or it may be a non-zero number close to 1 or close to 0. If the feasibility score is 0 or close to 0, it indicates that the action... Violating at least one constraint renders the action infeasible; a feasibility score close to 1 indicates that the action is feasible. Therefore, in summary, a preset score threshold is set, which can be adjusted according to the actual power grid conditions. Feasibility is then judged based on this threshold, and the initial Q value of unreliable candidate actions is updated to negative infinity.
[0062] Based on all updated initial Q values and all unupdated initial Q values, construct the current Q value matrix; Based on the current Q-value matrix and the current dynamic mask, the current final Q-value of each candidate operation is calculated.
[0063] Specifically, all current dynamic masks can also be constructed as a matrix, and the current final Q value of each candidate operation can be obtained by matrix multiplication.
[0064] Preferably, by forcing the initial Q-value of infeasible candidate actions to negative infinity, and through the element-wise multiplication operation of the dynamic mask and the Q-value matrix, the filtered Q-value matrix retains only the candidate actions that satisfy all constraints and their predicted rewards, thereby guiding the exploration direction of the Q-network to the safe solution space and reducing the interruption of the search path caused by the violation of constraints.
[0065] In this preferred embodiment, the current final Q value of each candidate operation is determined based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q value.
[0066] Step S104: Determine the optimal operation adjustment sequence based on the Q value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode. Specifically, since the state of the OCS system simulation model is constantly changing after the aforementioned steps, in order to determine the optimal operation sequence from several candidate operation adjustment operation sequences, it is still necessary to combine the constantly updated state and perform iterative operations similar to the aforementioned steps to determine the unique optimal operation adjustment operation sequence.
[0067] In a preferred embodiment, determining the optimal operation adjustment sequence based on the Q-value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model includes: The candidate run adjustment operation sequence corresponding to the maximum Q value is selected as the run adjustment operation sequence; Based on the initial OCS system simulation model, obtain the second OCS system simulation model after executing the last operation adjustment in the selected operation adjustment sequence, the second initial simulation topology data of the second OCS system simulation model, and the second initial simulation operation condition data; Specifically, by executing the selected operation adjustment sequence in the initial OCS system simulation model, the second OCS system simulation model after the last operation adjustment, the second initial simulation topology data of the second OCS system simulation model, and the second initial simulation operation condition data are obtained.
[0068] Based on the second OCS system simulation model and the preset operating mode, the model is determined by repeatedly executing the second operating mode adjustment operation until the second objective function of the current power outage section coupling constraint model converges and the second constraint violation amount of N consecutive iterations does not exceed the preset range, thus obtaining the optimal operating adjustment operation sequence. The second operating mode adjustment operation includes: Obtain the current second simulation topology data and the current second simulation running condition data of the current second OCS system simulation model; wherein, the initial OCS system simulation model is the second OCS system simulation model after the last running adjustment operation in the selected running adjustment operation sequence is executed, the initial second simulation topology data is the second initial simulation topology data, and the initial second simulation running condition data is the second initial simulation running condition data; Based on the current second simulation running condition data, the current second simulation topology data, and the preset running mode, the model is determined, and several current candidate operations, the initial Q value of each current candidate operation, and the current second constraint violation amount are obtained. The simulation model of the second OCS system is updated according to each current candidate operation to obtain the current second simulation topology data and the current second simulation operating condition data corresponding to each current candidate operation. Based on the current second simulation topology data corresponding to each current candidate operation, the current second simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, the current target candidate operation is determined from several current candidate operations, and the current power outage section coupling constraint model is constructed. If the second objective function of the current power outage section coupling constraint model does not converge, the current second constraint violation exceeds the preset range, or any second constraint violation exceeds the preset range in the previous N-1 consecutive iterations, the updated second OCS system simulation model corresponding to the current target candidate operation is used as the second OCS system simulation model at the next moment, the current second simulation topology data corresponding to the current target candidate operation is used as the second simulation topology data at the next moment, and the current second simulation operating condition data corresponding to the current target candidate operation is used as the second simulation operating condition data at the next moment. Otherwise, the operation sequence corresponding to the current target candidate operation is taken as the optimal operation adjustment sequence.
[0069] Specifically, in determining the optimal operational adjustment sequence, a selected operational adjustment sequence is first obtained from several candidate operational adjustment operation sequences based on the maximum Q-value. Based on the grid state under the last operation of this sequence, and combined with a greedy strategy to balance local optima and global search capabilities, feasible candidate actions with the highest initial Q-value under the current state are progressively selected for execution. Simultaneously, the environmental state is updated and fed back to the network for the next round of evaluation and decision-making, thereby driving the search path to continuously move towards higher-reward regions and rapidly converge the solution space. When the system state stabilizes within the safe region and the objective function converges stably, it is determined that the optimal solution space region has been reached, the search terminates, and the final target candidate operation is obtained. Subsequently, through reverse analysis, the operation sequence corresponding to this final target candidate operation is determined, resulting in the optimal operational adjustment sequence.
[0070] Therefore, in each iteration, the search path moves towards higher reward regions based on the Q-value gradient, while the dynamic mask continuously shrinks the solution space, eliminating new inactions caused by state changes, quickly adapting to real-time changes in the power grid, and gradually approaching the optimal solution within the safety region. The search terminates when the iteration process meets the dual convergence conditions: first, the system state stabilizes within the safety region, i.e., the second constraint violation amount does not exceed the preset range for N consecutive iterations; second, the objective function value tends to stabilize, and the change amount in multiple consecutive steps is less than a small threshold. At this point, it is determined that the optimal solution space region has been reached, the search process converges, and the action sequence from the initial state to the target state is output as the above-mentioned optimal operation adjustment sequence.
[0071] In this preferred embodiment, the optimal operation adjustment sequence is determined based on the Q value corresponding to each candidate operation adjustment sequence, the OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model.
[0072] Step S105: Adjust the operating mode of the OCS system to be adjusted according to the optimal operating adjustment sequence.
[0073] Specifically, the optimal operating adjustment sequence is converted into operating instructions that the OCS system can understand and sent to the OCS system to be adjusted. The system will then execute the switching operation, generator output regulation and load transfer actions in sequence according to the order of the operating instructions.
[0074] Preferably, while performing the operation, the power flow convergence and constraint compliance after each step can be verified in real time. Once an anomaly is detected, subsequent operations are immediately stopped and an early warning is issued to ensure the safety and controllability of the entire simulation process, avoid power grid accidents caused by improper operation, and ensure the safety and controllability of the entire process. After all operations are completed, it can be verified whether the final operation mode of the OCS system has completely eliminated all constraint conflicts caused by the overlap of power outage sections, and confirm whether key operating indicators, such as system loss and voltage quality, have reached the optimal target, so as to realize the automatic, safe and optimal adjustment of the operation mode.
[0075] Based on the above method embodiments, the present invention provides corresponding apparatus embodiments.
[0076] like Figure 2 As shown, an embodiment of the present invention provides an OCS system operation mode adjustment device, comprising: The module includes a data acquisition module, a simulation model construction module, a Q-value calculation module, an optimal operation adjustment sequence filtering module, and an operation mode adjustment module. The data acquisition module is used to acquire the current topology data, equipment parameters, operating condition data, and planned power outage information of the OCS system to be adjusted; The simulation model construction module is used to construct an initial OCS system simulation model based on the topology data, equipment parameters, operating condition data, and planned power outage information. The Q-value calculation module is used to determine a number of candidate operation adjustment sequences and the Q-value corresponding to each candidate operation adjustment sequence based on the initial OCS system simulation model and the preset operation mode. The optimal operation adjustment sequence filtering module is used to determine the optimal operation adjustment sequence based on the Q value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model. The operation mode adjustment module is used to adjust the operation mode of the OCS system to be adjusted according to the optimal operation adjustment sequence.
[0077] It should be noted that the device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without creative effort. The above schematic diagrams are merely examples of an OCS system operation mode adjustment device and do not constitute a limitation on an OCS system operation mode adjustment device. It may include more or fewer components than shown, or combine certain components, or use different components.
[0078] Based on the above method embodiments, the present invention provides corresponding terminal device embodiments.
[0079] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the operation mode adjustment method of an OCS system described in any embodiment of the present invention.
[0080] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the device. The aforementioned terminal devices may be computing devices such as desktop computers, laptops, handheld computers, and cloud servers. These devices may include, but are not limited to, processors and memory. The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. This processor is the control center of the device, connecting various parts of the device via various interfaces and lines. The aforementioned memory can be used to store the aforementioned computer programs and / or modules. The aforementioned processor implements various functions of the aforementioned device by running or executing the computer programs and / or modules stored in the aforementioned memory, and by calling data stored in the memory. The aforementioned memory may mainly include a program storage area and a data storage area, wherein the program storage area may store the operating system, at least one application program required for a function, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0081] Based on the above method embodiments, the present invention provides corresponding storage medium embodiments.
[0082] Another embodiment of the present invention provides a storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the storage medium is located to execute the operating mode adjustment method of an OCS system described in any embodiment of the present invention.
[0083] In this embodiment, the storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0084] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for adjusting the operating mode of an OCS system, characterized in that, include: Obtain the current topology data, equipment parameters, operating status data, and planned power outage information of the OCS system to be adjusted; Based on the topology data, equipment parameters, operating condition data, and planned power outage information, an initial OCS system simulation model is constructed. Based on the initial OCS system simulation model and the preset operation mode, a model is determined, and several candidate operation adjustment operation sequences and the Q value corresponding to each candidate operation adjustment operation sequence are determined. Based on the Q value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode, the model is determined, and the optimal operation adjustment sequence is determined. The operating mode of the OCS system to be adjusted is adjusted according to the optimal operating adjustment sequence.
2. The method for adjusting the operation mode of an OCS system according to claim 1, characterized in that, The step of determining a number of candidate operation adjustment sequences and the Q value corresponding to each candidate operation adjustment sequence based on the initial OCS system simulation model and the preset operation mode includes: Obtain the first initial simulation topology data and the first initial simulation operating condition data of the initial OCS system simulation model; Based on the first initial simulation topology data, the first initial simulation operating condition data, the initial OCS system simulation model, and the preset operating mode determination model, the first operating mode adjustment operation is repeatedly executed to obtain the candidate operating mode adjustment operation sequence, and the final Q value corresponding to the last operating mode adjustment operation in each candidate operating mode adjustment operation sequence. The final Q value is used as the Q value for each candidate run adjustment operation sequence; The first operation mode adjustment operation includes: Obtain the current first simulation topology data and the current first simulation running condition data of the current first OCS system simulation model; wherein, the initial first OCS system simulation model is the initial OCS system simulation model, the initial first simulation topology data is the first initial simulation topology data, and the initial first simulation running condition data is the first initial simulation running condition data; Based on the current first simulation running condition data, the current first simulation topology data, and the preset running mode, the model is determined, and several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount are obtained. The simulation model of the first OCS system is updated according to each current candidate operation to obtain the current first simulation topology data and the current first simulation running condition data corresponding to each current candidate operation. Based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, the current target candidate operation is determined from several current candidate operations, and the current power outage section coupling constraint model is constructed. If the first objective function of the current power outage section coupling constraint model does not converge, the current first constraint violation exceeds the preset range, or any first constraint violation exceeds the preset range in the previous N-1 consecutive iterations, the updated first OCS system simulation model corresponding to the current target candidate operation is used as the first OCS system simulation model at the next moment, the current first simulation topology data corresponding to the current target candidate operation is used as the first simulation topology data at the next moment, and the current first simulation operating condition data corresponding to the current target candidate operation is used as the first simulation operating condition data at the next moment. Otherwise, the operation sequence corresponding to each current candidate operation is taken as the candidate operation adjustment sequence.
3. The method for adjusting the operation mode of an OCS system according to claim 2, characterized in that, The step of determining the model based on the current first simulation running condition data, the current first simulation topology data, and the preset running mode, and obtaining several current candidate operations, the initial Q value of each current candidate operation, and the current first constraint violation amount, includes: Based on the current first simulation running condition data, the current first constraint violation amount is calculated; The current first simulation topology data, the current first simulation running condition data, and the current first constraint violation are input into the preset running mode determination model to obtain several current candidate operations and the initial Q value of each current candidate operation.
4. The method for adjusting the operation mode of an OCS system according to claim 3, characterized in that, The preset operating mode determines the construction of the model, including: Acquire several historical operation scheduling data of the OCS system to be adjusted under different historical operation scenarios; wherein, the historical operation scenarios include: normal operation scenario, overload operation scenario and dynamic instability operation scenario; the historical operation scheduling data includes: historical operation adjustment operation, historical operation condition data corresponding to the historical operation adjustment operation, and topology data corresponding to the historical operation adjustment operation. Based on the historical operating data, the historical reward value and the corresponding historical constraint violation amount for each historical operating adjustment operation are calculated. The historical operation scheduling data, historical reward value, and historical constraint violation amount corresponding to each historical operation adjustment operation are used as a set of candidate training sample data, and then several sets of candidate training sample data are obtained. From several sets of candidate training sample data, several sets of selected training sample data are obtained by sampling according to the preset historical operation scenario sampling ratio; Several sets of selected training sample data are input into the model to be trained for iterative training until the loss function converges, thereby generating the preset operating mode determination model; In each iteration of training, the model is determined based on the current operating mode and the currently selected training sample data to obtain the current predicted Q value and the current target Q value; the current loss function is calculated based on the current predicted Q value and the current target Q value, and it is determined whether the current loss function has converged; if it has converged, the model determined by the current operating mode is used as the preset operating mode model; otherwise, the model parameters in the current operating mode model are adjusted and training continues.
5. The method for adjusting the operation mode of an OCS system according to claim 4, characterized in that, The step of determining the current target candidate operation from a plurality of current candidate operations based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the device parameters includes: Based on the current first simulation topology data corresponding to each current candidate operation, the current first simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, a coupling constraint model of the current power outage section is constructed with the goal of minimizing the power outage range, load loss, and equipment overload risk. Based on the first current constraint condition of the current power outage section coupling constraint model, generate the current dynamic mask for each current candidate operation; Based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q value, determine the current final Q value of each candidate operation; The candidate operation corresponding to the largest current final Q value is taken as the current target candidate operation.
6. The method for adjusting the operation mode of an OCS system according to claim 5, characterized in that, The step of determining the current final Q value of each candidate operation based on the first current constraint, the current first simulation running condition data, the current dynamic mask, and the corresponding initial Q value includes: For each current candidate operation, the current feasibility score of each current candidate operation is calculated based on the current first simulation running condition data corresponding to the current candidate operation and the corresponding first current constraint condition. Update the initial Q value of the current candidate operation whose current feasibility score is less than the preset score threshold to negative infinity; Based on all updated initial Q values and all unupdated initial Q values, construct the current Q value matrix; Based on the current Q-value matrix and the current dynamic mask, the current final Q-value of each candidate operation is calculated.
7. The method for adjusting the operation mode of an OCS system according to claim 6, characterized in that, The step of determining the optimal operation adjustment sequence based on the Q-value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model includes: The candidate run adjustment operation sequence corresponding to the maximum Q value is selected as the run adjustment operation sequence; Based on the initial OCS system simulation model, obtain the second OCS system simulation model after executing the last operation adjustment in the selected operation adjustment sequence, the second initial simulation topology data of the second OCS system simulation model, and the second initial simulation operation condition data; Based on the second OCS system simulation model and the preset operating mode, the model is determined by repeatedly executing the second operating mode adjustment operation until the second objective function of the current power outage section coupling constraint model converges and the second constraint violation amount of N consecutive iterations does not exceed the preset range, thus obtaining the optimal operating adjustment operation sequence. The second operating mode adjustment operation includes: Obtain the current second simulation topology data and the current second simulation running condition data of the current second OCS system simulation model; wherein, the initial OCS system simulation model is the second OCS system simulation model after the last running adjustment operation in the selected running adjustment operation sequence is executed, the initial second simulation topology data is the second initial simulation topology data, and the initial second simulation running condition data is the second initial simulation running condition data; Based on the current second simulation running condition data, the current second simulation topology data, and the preset running mode, the model is determined, and several current candidate operations, the initial Q value of each current candidate operation, and the current second constraint violation amount are obtained. The simulation model of the second OCS system is updated according to each current candidate operation to obtain the current second simulation topology data and the current second simulation operating condition data corresponding to each current candidate operation. Based on the current second simulation topology data corresponding to each current candidate operation, the current second simulation operating condition data corresponding to each current candidate operation, and the equipment parameters, the current target candidate operation is determined from several current candidate operations, and the current power outage section coupling constraint model is constructed. If the second objective function of the current power outage section coupling constraint model does not converge, the current second constraint violation exceeds the preset range, or any second constraint violation exceeds the preset range in the previous N-1 consecutive iterations, the updated second OCS system simulation model corresponding to the current target candidate operation is used as the second OCS system simulation model at the next moment, the current second simulation topology data corresponding to the current target candidate operation is used as the second simulation topology data at the next moment, and the current second simulation operating condition data corresponding to the current target candidate operation is used as the second simulation operating condition data at the next moment. Otherwise, the operation sequence corresponding to the current target candidate operation is taken as the optimal operation adjustment sequence.
8. An operating mode adjustment device for an OCS system, characterized in that, include: The module includes a data acquisition module, a simulation model construction module, a Q-value calculation module, an optimal operation adjustment sequence filtering module, and an operation mode adjustment module. The data acquisition module is used to acquire the current topology data, equipment parameters, operating condition data, and planned power outage information of the OCS system to be adjusted; The simulation model construction module is used to construct an initial OCS system simulation model based on the topology data, equipment parameters, operating condition data, and planned power outage information. The Q-value calculation module is used to determine a number of candidate operation adjustment sequences and the Q-value corresponding to each candidate operation adjustment sequence based on the initial OCS system simulation model and the preset operation mode. The optimal operation adjustment sequence filtering module is used to determine the optimal operation adjustment sequence based on the Q value corresponding to each candidate operation adjustment sequence, the initial OCS system simulation model, each candidate operation adjustment sequence, and the preset operation mode determination model. The operation mode adjustment module is used to adjust the operation mode of the OCS system to be adjusted according to the optimal operation adjustment sequence.
9. A terminal device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement an OCS system operation mode adjustment method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium includes a stored computer program, wherein, when the computer program is running, it controls the device where the storage medium is located to execute an OCS system operation mode adjustment method as described in any one of claims 1 to 7.