Anti-error checking method, system and equipment based on deep reinforcement learning and medium
Through the anti-miss-checking method based on deep reinforcement learning, the shutter backward operation task is automatically obtained and the operation sequence that conforms to the anti-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss-miss
Patent Information
- Application Number
- CN202510002473.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-03
AI Technical Summary
The existing anti-error system relies on manual experience, has poor scalability and low maintenance, and is difficult to adapt to the rapid development of the power grid. It is impossible to effectively prevent safety problems caused by misoperation of secondary pressure plates, misoperation, misoperation, and other misoperation operations.
The anti-missive verification method based on deep reinforcement learning is adopted, and the reverse-missive verification is obtained by constructing a professional anti-missive corpus, and the in-missive learning algorithm is used to analyze and infer the reverse-missive operation sequence that conforms to the anti-missive logic, so as to achieve intelligent anti-missive verification.
Reduce the dependence on manual experience, improve the efficiency and accuracy of error prevention, reduce the rate of reverse gate error, and ensure the reliability and safety of reverse gate operation of the substation.
Smart Images

Figure CN120088090A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of anti-misoperation for substation switching operations and artificial intelligence, and particularly relates to an anti-misoperation verification method, system, device and medium based on deep reinforcement learning. Background Art
[0002] With the increasing growth of the power grid scale, the number of devices faced by power grid operators at all levels has increased rapidly, and the risk of misoperation in daily work such as writing operation tickets, switching operations, and operation mode adjustment has gradually increased. The scope of remote operation of primary and secondary equipment is getting wider and wider, and misoperations such as missing or incorrect switching of primary and secondary equipment or switching in an incorrect order often occur, resulting in power grid accidents from time to time, seriously affecting the safe and stable operation of the power grid. Substation electrical misoperations can not only cause accidents such as large-scale power outages, equipment damage, and personal injuries, but may even lead to the oscillation and collapse of the power grid, seriously threatening the safety of the power grid, equipment and personnel, and affecting the safe operation of the power system. The substation anti-misoperation locking function is an important guarantee to prevent electrical misoperations. At the current stage, the anti-misoperation system mainly establishes an anti-misoperation rule library based on expert experience for anti-misoperation constraints. The anti-misoperation rule library mainly relies on manual configuration, with a single judgment basis, poor scalability, low maintainability, and an update method that is difficult to adapt to the rapid development of the power grid. For the switching operations of secondary pressure plates, the microcomputer anti-misoperation system does not perform anti-misoperation verification, nor does it have control functions such as guiding the process of pressure plate operation and comparing states, and cannot effectively solve the safety problems caused by missing, incorrect, or wrong switching of secondary pressure plates.
[0003] With the rapid development of artificial intelligence technology, reinforcement learning (RL), as an important branch in the field of machine learning, is used to solve stochastic sequence decision-making problems. It does not need to understand the environmental changes in advance and has strong adaptability to many uncertain interferences. Compared with deep learning, deep learning focuses on recognition and expression, while reinforcement learning focuses on making optimal decisions by repeatedly interacting with the environment and has good generalization ability. In particular, deep reinforcement learning (DRL) combined with deep neural networks has good decision-making ability and adaptive learning ability for non-convex and non-linear problems. During training, the agent sends actions to interact with the environment. After each interaction, it determines whether the current behavior is conducive to achieving the goal according to the reward value feedback by the environment. By recording the state transitions and experiences experienced by the agent, it continuously iteratively updates its own policy until the policy converges to the optimal.
[0004] To cope with the complex and changing power grid forms and break through the bottleneck of switching operation decision-making and operation that relies on experience, the power system urgently needs to use advanced information technologies such as artificial intelligence to improve the intelligent error prevention level of substation switching operations. On the premise of ensuring the reliable operation of the system, automatically and intelligently obtain error prevention tasks, accurately infer the error prevention sequence that conforms to the switching regulations, perform error prevention verification, improve the error prevention efficiency and accuracy, and reduce the switching error rate. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an error prevention verification method, system, device and medium based on deep reinforcement learning, which uses the deep reinforcement learning algorithm to realize switching error prevention verification, reduces the dependence on manual experience, and avoids safety problems caused by misoperations such as omission, incorrect switching or non-switching in accordance with the specified sequence of primary and secondary equipment, and improves the reliability and safety of substation switching operations.
[0006] To achieve the above technical objectives, the present invention is implemented by the following technical solutions.
[0007] In the first aspect, an error prevention verification method based on deep reinforcement learning is provided, including the following steps: Step S1: Construct an error prevention professional corpus, including structured data and unstructured data; where the structured data includes equipment model data and power grid topology data, and the unstructured data includes error prevention rules; Step S2: Based on the error prevention professional corpus, obtain the switching operation task, and parse and infer a switching operation sequence that conforms to the error prevention logic based on the deep reinforcement learning algorithm; Step S3: Parse the input switching operation ticket, output the intelligent error prevention verification result, and realize the intelligent error prevention of the switching operation.
[0008] Furthermore, in the step S1, the structured data includes the equipment model attributes of substations, transformers, lines, buses, circuit breakers, disconnectors, and earthing switches, and the topological connection relationships between the various equipment models; the error prevention rules include the first constraint relationship between primary equipment and primary equipment, the second constraint relationship between primary equipment and secondary equipment, and the third constraint relationship between secondary equipment and secondary equipment, which are the basic principles that must be followed in power grid switching operations.
[0009] Furthermore, in the step S2, the deep reinforcement learning algorithm learns and trains in a way of continuously exploring and trial-and-error, uses the obtained rewards or punishments to guide the switching decision-making, and continuously adjusts and optimizes the action strategy according to the action behavior effect, so as to adapt to the environment and obtain the optimal solution; at time t, after the switching operation task is issued, the agent obtains the state of the equipment under the current switching operation task and then makes a switching decision selection and executes the corresponding action ,The intelligent agent changes the learning environment and continuously tries and makes errors until it obtains the maximum reward value, finds the optimal action strategy, and finally infers the switching operation sequence that conforms to the error-proofing logic.
[0010] Furthermore, in step S2, the state space ,in for t The status of the switching equipment at all times, including primary and secondary equipment; is the real-time status of each switching device in the action space during the reasoning process; action space , for t The moment action space n The switching action value of a switching device.
[0011] Furthermore, in step S2, the reward Indicates the agent's current state Next, perform the selected action The reward obtained is that the correct switching operation is to execute the switching operation task to the target state with the least number of operations, without any illegal operations during the execution process. After the switching decision is completed, the higher the consistency between the equipment operation state and the target state, the greater the corresponding reward value; the penalty Indicates the penalty for a wrong action. A wrong action refers to an action that violates the error prevention rules. If a wrong action occurs during the operation, an action penalty will be given and a new round will start; target reward ;
[0012] Where: is the number of actions; is the target state consistency factor, where M is the number of false operations; is the initial state value of the switching device; is the target state value of the switching device; the state value of the switching device in the operating state is 1, the state value in the hot standby state is 2, and the state value in the cold standby state is 3; is the value corresponding to the i-th action, the correct action value is 4, and the incorrect action value is 0;
[0013] Where: If the kth action is a false action when executing the selected action, then The value is -1. If the kth action is not a false action when executing the selected action, then The value is 0.
[0014] Further, in the step S2, the agent of the deep reinforcement learning algorithm consists of a deep double Q network (DDQN), and the calculation formula is:
[0015] In the formula: is the target Q value; is the immediate reward; is the discount factor; is the neural network function of the action value, is to make the action when taking the maximum value; are the main network parameters; are the target network parameters; The loss function for updating the neural network weights is expressed as:
[0016] In the formula: is the target Q value at time t; is the target Q value at time t-1.
[0017] Further, in the step S3, the input switching operation ticket is parsed and compared with the switching operation sequence that conforms to the anti-error logic parsed and inferred by the deep reinforcement learning algorithm, and the intelligent anti-error verification result is output to realize the intelligent anti-error of the switching operation.
[0018] In the second aspect, an anti-error verification system based on deep reinforcement learning is provided. Based on the anti-error verification method based on deep reinforcement learning as described above, it includes An anti-error professional corpus construction module for constructing an anti-error professional corpus, including structured data and unstructured data; where the structured data includes device model data and power grid topology data, and the unstructured data includes anti-error rules; A learning and training module for obtaining switching operation tasks based on the anti-error professional corpus and parsing and inferring a switching operation sequence that conforms to the anti-error logic based on the deep reinforcement learning algorithm; A result output module for parsing the input switching operation ticket, comparing it with the switching operation sequence that conforms to the anti-error logic parsed and inferred by the deep reinforcement learning algorithm, and outputting an intelligent anti-error verification result to realize the intelligent anti-error of the switching operation.
[0019] In the third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the anti-error verification method based on deep reinforcement learning as described above is implemented.
[0020] Fourthly, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are used to implement the anti-error verification method based on deep reinforcement learning as described above.
[0021] The present invention provides an anti-error verification method, system, device and medium based on deep reinforcement learning, which has the following beneficial effects: In view of the problems that the existing anti-error rule base mainly relies on manual configuration, has a single judgment basis, poor scalability, low maintainability, and a difficult-to-adapt update method to the rapid development of the power grid, the present invention proposes an anti-error verification solution based on deep reinforcement learning, which automatically obtains switching operation tasks and accurately infers a switching operation sequence that conforms to the anti-error logic to compare with the input switching operation ticket, realizes anti-error verification, copes with the complex and changeable power grid form, breaks through the bottleneck of switching decision-making and operation relying on experience, avoids safety problems caused by misoperations such as missed or incorrect switching of primary and secondary equipment or switching not in the specified order, improves the anti-error efficiency and accuracy, reduces the switching error rate, and ensures the reliability and safety of substation switching operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 is a flowchart of the anti-error verification method based on deep reinforcement learning provided by the embodiment of the present invention; Figure 2 is a structural diagram of the anti-error verification system based on deep reinforcement learning provided by the embodiment of the present invention; Figure 3 is a partial wiring diagram of a 220 kV substation provided by the embodiment of the present invention; Figure 4 is an example of intelligent anti-error verification for switching operations provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope protected by the present invention.
[0025] To cope with the complex and changeable power grid forms and break through the bottleneck of switching operation decision-making and operation relying on experience, the present invention utilizes advanced information technologies such as artificial intelligence to improve the intelligent error prevention level of substation switching operations. The technical solution of the present invention will be specifically described below in conjunction with specific embodiments.
[0026] Embodiment 1
[0027] This embodiment provides an error prevention verification method based on deep reinforcement learning, as Figure 1 shown, including the following steps: Step S1, construct an error prevention professional corpus, including structured data and unstructured data; among them, the structured data includes equipment model data and power grid topology data, and the unstructured data includes error prevention rules.
[0028] In this embodiment, the structured data includes equipment model attributes such as substations, transformers, lines, buses, circuit breakers, disconnectors, grounding switches, etc. and the topological connection relationships between various equipment models; the unstructured data (error prevention rules) includes the first constraint relationship between primary equipment and primary equipment, the second constraint relationship between primary equipment and secondary equipment, and the third constraint relationship between secondary equipment and secondary equipment, which are the basic principles that must be followed in power grid switching operations. The specific constraint relationships are determined according to the actual power grid structure, which belongs to the conventional technology in this field and is not the focus of the present invention, so it will not be elaborated here.
[0029] Step S2, based on the error prevention professional corpus, obtain the switching operation task, and parse and infer a switching operation sequence that conforms to the error prevention logic based on the deep reinforcement learning algorithm.
[0030] In this embodiment, the deep reinforcement learning algorithm learns and trains in a way of continuously exploring and trial-and-error, uses the obtained rewards or punishments to guide the switching decision-making, and continuously adjusts and optimizes the action strategy according to the action behavior effect, so as to adapt to the environment and obtain the optimal solution; at time t, after the switching operation task is issued, the intelligent agent obtains the state of the equipment under the current switching operation task, and then makes a switching decision selection and executes the corresponding action. The intelligent agent continuously tries and errors until it obtains the maximum reward value by changing the learning environment, finds the optimal action strategy, and finally infers a switching operation sequence that conforms to the error prevention logic.
[0031] State space , where is t the state of the switching equipment at time ; the switching equipment includes primary and secondary equipment; , is t the real-time state of each switching equipment in the action space during the inference process; action space nThe switching operation values of switching equipment. Since the switching operation sequence needs to be deduced finally, it is an iterative reasoning process. In each reasoning process, each action needs to be iteratively deduced round by round, and finally the switching operation sequence is formed. To improve the accuracy of reasoning, the state includes the initial state of the switching equipment at the current moment and the real-time state of each switching equipment in the action space during the reasoning process. Specifically: is t the state of the switching equipment at time, which is the initial state, In the first round of iterative reasoning process, it is the same as ; In the first round of iterative reasoning process, an action has been selected. Therefore, the state of a certain switching equipment in the action space has changed. Therefore, in the second round of iterative reasoning process, remains unchanged, but has changed, which is the real-time state of each switching equipment in the action space formed after the action selected in the first round of iterative reasoning process; And so on, until all rounds of reasoning are completed, and finally the switching operation sequence is obtained.
[0032] During the reasoning process, every time a switching operation sequence is deduced, its corresponding target reward is calculated. The reward represents the reward obtained by the agent when executing the selected action in the current state . The correct switching operation is to execute the switching operation task to the target state with as few operation times as possible, and there are no any illegal operations during the execution process. After the switching decision is completed, the higher the degree of consistency between the equipment operation state and the target state, the greater the corresponding reward value; The penalty represents the penalty obtained for misoperation. Misoperation refers to the action that violates the anti-misoperation rules. If a misoperation occurs during the execution of the operation, an action penalty will be given and a new round will start.
[0033] The reward is calculated as follows:
[0034] In the formula: is the number of actions; is the target state consistency factor, where M is the number of misoperations; is the initial state value of the switching equipment; is the target state value of the switching equipment; The state value of the switching equipment in the running state is 1, the state value in the hot standby state is 2, and the state value in the cold standby state is 3; is the value corresponding after the i-th action, and the correct action takes the value of 4, and the misoperation takes the value of 0.
[0035] For example: (1) Transferring from the operation of Bus 1 (602 closed, 6021 closed, 6023 closed) to the cold standby state (602 open, 6021 open, 6023 open): The correct action sequence is 602 / 6023 / 6021. At this time = (3 - 1)[(4 - 1) + (4 - 1) + (4 - 1)] / 3 / 1; (2) Transferring from the hot standby state of Bus 1 (602 open, 6021 closed, 6023 closed) to the cold standby state (602 open, 6021 open, 6023 open): The correct action sequence is 6023 / 6021. At this time = (3 - 2)[(4 - 2) + (4 - 2)] / 2 / 1; In the above (1) and (2), no misoperation occurs. (3) Transferring from the hot standby state of Bus 1 (602 open, 6021 closed, 6023 closed) to the cold standby state (602 open, 6021 open, 6023 open): If a misoperation occurs once at this time, the sequence is 6021 / 6023 / 6021. At this time = (3 - 2)[(0 - 2) + (4 - 2) + (4 - 2)] / 3 / 2. The judgment of misoperation is realized by anti-misoperation rules.
[0036]
[0037] In the formula: If the k-th action is a misoperation when performing a selection action, then takes the value of -1. If the k-th action is not a misoperation when performing a selection action, then takes the value of 0.
[0038] In this embodiment, the agent of the deep reinforcement learning algorithm consists of a deep double Q-network (DDQN). The calculation formula is:
[0039] In the formula: is the target Q value; is the immediate reward; is the discount factor; is the neural network function of the action value, is to make the action when taking the maximum value; are the main network parameters; are the target network parameters; The loss function for updating the neural network weights is expressed as:
[0040] In the formula: is the target Q value at time t; is the target Q value at time t - 1.
[0041] Step S3: Parse the input switching operation ticket, compare it with the switching operation sequence that conforms to the anti-error logic parsed and inferred by the deep reinforcement learning algorithm, and output the intelligent anti-error verification result to achieve intelligent anti-error for switching operations.
[0042] The anti-error verification method based on deep reinforcement learning provided in the above embodiment automatically obtains the switching operation task, accurately infers the switching operation sequence that conforms to the anti-error logic, compares it with the input switching operation ticket, realizes anti-error verification, copes with the complex and changeable power grid form, breaks through the bottleneck of switching decision-making and operation relying on experience, avoids safety problems caused by misoperations such as missing or incorrect switching of primary and secondary equipment or switching in an improper order, improves the anti-error efficiency and accuracy, reduces the switching error rate, and ensures the reliability and safety of substation switching operations.
[0043] Embodiment 2
[0044] Based on the anti-error verification method based on deep reinforcement learning provided in the foregoing embodiment, this embodiment provides an anti-error verification system based on deep reinforcement learning, as Figure 2 shown, including, An anti-error professional corpus construction module for constructing an anti-error professional corpus, including structured data and unstructured data; wherein the structured data includes equipment model data and power grid topology data, and the unstructured data includes anti-error rules; A learning and training module for obtaining the switching operation task based on the anti-error professional corpus and parsing and inferring the switching operation sequence that conforms to the anti-error logic based on the deep reinforcement learning algorithm; A result output module for parsing the input switching operation ticket, comparing it with the switching operation sequence that conforms to the anti-error logic parsed and inferred by the deep reinforcement learning algorithm, and outputting the intelligent anti-error verification result to achieve intelligent anti-error for switching operations.
[0045] It should be understood that the functional unit modules in the embodiments of the present invention can be concentrated in one processing unit, or each unit module can exist physically alone, or two or more unit modules can be integrated in one unit module, and can be implemented in the form of hardware or software.
[0046] Embodiment 3
[0047] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the anti-error verification method based on deep reinforcement learning described in the foregoing embodiment. The memory can be various types of memories, such as random access memory, read-only memory, flash memory, etc. The processor can be various types of processors, for example, a central processing unit, a microprocessor, a digital signal processor, etc.
[0048] Example 4
[0049] This embodiment provides a computer-readable storage medium on which computer-executable instructions are stored. The computer-executable instructions are used to implement the anti-error verification method based on deep reinforcement learning described in the above embodiments.
[0050] Example 5
[0051] Figure 3 The figure shows a local wiring diagram of a 220 kV substation in an embodiment of the present invention. The meanings of the numbers of the power equipment in the figure are: 1M (1 mother busbar), 2M (2 mother busbars), 6021 (6021 disconnector), 6022 (6022 disconnector), 6023 (6023 disconnector), 602 (602 circuit breaker), 6021-1 (6021-1 grounding disconnector), 6023-1 (6023-1 grounding disconnector), 6023-2 (6023-2 grounding disconnector).
[0052] Taking the switching operation task: 220kV XX line 602 circuit breaker switches from operation to cold standby (Ⅰ mother) as an example, a method for error prevention verification based on deep reinforcement learning is described, which specifically includes the following steps: The first step is to construct professional error prevention corpus, including structured data and unstructured data. The structured data includes the model attributes of equipment such as substations, transformers, lines, busbars, circuit breakers, disconnectors, and grounding switches and their topological connection relationships. The unstructured data includes the first constraint relationship between primary equipment and primary equipment, the second constraint relationship between primary equipment and secondary equipment, and the third constraint relationship between secondary equipment and secondary equipment.
[0053] The second step is to obtain the switching operation task 220kVXX line 602 circuit breaker from operation to cold standby (Ⅰ bus), the action space is A set bus protection 602 start failure receiving soft pressure plate, A set bus protection 602 trip outlet soft pressure plate, A set bus protection 602 SV receiving soft pressure plate, B set bus protection 602 start failure receiving soft pressure plate, B set bus protection 602 trip outlet soft pressure plate, B set bus protection 602 SV receiving soft pressure plate, A set line protection 602 start failure sending soft pressure plate, A set line protection 602 reclosing outlet soft pressure plate, B set line protection 602 start failure sending soft pressure plate, B set line protection 602 reclosing outlet soft pressure plate, 602 circuit breaker, 6023 disconnector, 6021 disconnector; based on the deep reinforcement learning algorithm, learning and training are carried out in a continuous exploration manner, and the rewards or penalties obtained are used to guide the switching operation decision. According to the effect of the action behavior, the action strategy is continuously adjusted and optimized to infer the switching operation sequence that conforms to the anti-error logic, as shown in Table 1.
[0054]
[0055] In the third step, the input switching operation ticket (as shown in Table 2) is intelligently parsed and compared with the switching operation sequence that conforms to the anti-error logic inferred by the deep reinforcement learning algorithm (as shown in Table 1) to output the intelligent anti-error verification result. In Table 2, there are errors in items 2 and 3. When the circuit breaker is transferred from the operating state to the cold standby state, the line-side disconnecting switch should be opened first. An example of intelligent anti-error verification for switching operations is attached Figure 4 as shown
[0056]
[0057] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content of other embodiments
[0058] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code
[0059] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks
[0060] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks
[0061] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 in one or more processes and / or boxes Figure 1 of the functions specified in one box or a plurality of boxes.
[0062] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art may make variations, modifications, substitutions, and alterations within the scope of the present invention to the above embodiments.
Claims
1. A method for preventing error verification based on deep reinforcement learning, characterized in that: The following steps are involved: Step S1, constructing error prevention professional corpus, including structured data and unstructured data; The structured data includes equipment model data and power grid topology data, and the unstructured data includes error prevention rules; Step S2: Based on the error prevention professional corpus, the switching operation task is obtained, and the switching operation sequence that conforms to the error prevention logic is analyzed and inferred based on the deep reinforcement learning algorithm; Step S3, parse the input switching operation ticket, output the intelligent error prevention verification result, and realize the intelligent error prevention of switching operation.
2. The error-proofing method based on deep reinforcement learning according to claim 1, characterized in that: In step S1, the structured data includes the equipment model attributes of the substation, transformer, line, busbar, circuit breaker, disconnector, and earthing switch, and the topological connection relationship between the equipment models; the error prevention rules include the first constraint relationship between primary equipment and primary equipment, the second constraint relationship between primary equipment and secondary equipment, and the third constraint relationship between secondary equipment and secondary equipment.
3. The error-proofing method based on deep reinforcement learning according to claim 1, characterized in that: In step S2, the deep reinforcement learning algorithm learns and trains in a continuous trial-and-error manner, uses the rewards or penalties obtained to guide the switching decision, and continuously adjusts and optimizes the action strategy according to the effect of the action behavior, so as to adapt to the environment and obtain the optimal solution; at time t, after the switching operation task is issued, the intelligent agent obtains the status of the equipment under the current switching operation task, and then makes a switching decision selection and executes the corresponding action. The intelligent agent changes the learning environment and continuously tries and errors until the maximum reward value is obtained, finds the optimal action strategy, and finally infers a switching operation sequence that conforms to the error-proof logic.
4. The error-proofing method based on deep reinforcement learning according to claim 3 is characterized in that: In step S2, the state space ,in for t The status of the switching equipment at all times, is the real-time status of each switching device in the action space during the reasoning process; action space , for t The moment action space n The switching action value of a switching device.
5. The error-proofing method based on deep reinforcement learning according to claim 3 is characterized in that: In step S2, the reward It indicates the reward obtained by the agent for executing the selected action in the current state. There is no illegal operation during the execution of the action. After the switching decision is completed, the higher the consistency between the equipment operation state and the target state, the greater the corresponding reward value. Indicates the penalty for a wrong action. A wrong action refers to an action that violates the error prevention rules. If a wrong action occurs when executing a selected action, an action penalty will be given and a new round will begin; Target reward .
6. The error-proofing method based on deep reinforcement learning according to claim 3 is characterized in that: In step S2, the agent of the deep reinforcement learning algorithm is composed of a deep double Q network, and the calculation formula is: ; Where: is the target Q value; For immediate rewards; is the discount factor; is the neural network function of action value, To make The action when the value is maximum; is the main network parameter; is the target network parameter; The loss function used to update the neural network weights It is expressed as: ; Where: is the target Q value at time t; is the target Q value at time t-1.
7. The error-proofing method based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the input switching operation ticket is parsed and compared with the switching operation sequence that conforms to the error prevention logic analyzed and inferred by the deep reinforcement learning algorithm, and the intelligent error prevention verification result is output to realize the intelligent error prevention of the switching operation.
8. A system for preventing error checking based on deep reinforcement learning, and a method for preventing error checking based on deep reinforcement learning according to any one of claims 1 to 7, characterized in that: include, The error prevention professional corpus construction module is used to construct error prevention professional corpus, including structured data and unstructured data; The structured data includes equipment model data and power grid topology data, and the unstructured data includes error prevention rules; The learning and training module is used to obtain the switching operation tasks based on the error prevention professional corpus, and parse and infer the switching operation sequence that conforms to the error prevention logic based on the deep reinforcement learning algorithm; The result output module is used to parse the input switching operation ticket, compare it with the switching operation sequence that conforms to the error prevention logic analyzed and inferred by the deep reinforcement learning algorithm, output the intelligent error prevention verification result, and realize the intelligent error prevention of switching operation.
9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the anti-error verification method based on deep reinforcement learning described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: Computer executable instructions are stored thereon, and the computer executable instructions are used to implement the anti-error verification method based on deep reinforcement learning as described in any one of claims 1 to 7.