Optimization Configuration Method and System for Distribution Network Distributed Protection Device

By calculating the data transmission error rate and reliability of the distribution network distributed protection system, and using reinforcement learning to optimize the configuration of distributed protection devices, the optimization layout of protection devices in the distribution network is solved, and the reliability and flexibility of the system are improved.

CN115513942BActive Publication Date: 2025-07-29STATE GRID ANHUI ELECTRIC POWER CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211277094.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2025-07-29
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

In the prior art, the distributed protection devices of the distribution network lack an optimized configuration method, which is difficult to meet the reliability needs of distribution network protection services, and the traditional protection device arrangement lacks quantitative evaluation indicators.

Method used

By combining the ultra-reliable low-latency communication theory, the data transmission error rate and reliability of the distribution network distributed protection system are calculated, and the neural network model is trained using reinforcement learning methods. The location and number of main station arrangements of the distribution network protection service are optimized based on the reliability requirements of distribution network protection services.

Benefits of technology

The optimized layout of distributed protection devices in distribution networks is realized, which improves the reliability and flexibility of the system, avoids the limitations of traditional manual design, and improves the convenience of installation and maintenance of the protection devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115513942B_ABST
    Figure CN115513942B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimized configuration method and system for a distribution network distributed protection device, including calculating the data transmission error rate corresponding to different configuration states of the distribution network distributed protection system based on the respective configuration states of the distributed protection devices in the distribution network and combining the theory of ultra-reliable low-latency communication applications; calculating the reliability corresponding to different configuration states of the distribution network distributed protection system based on the data transmission error rate corresponding to different configuration states; using the reliability corresponding to different configuration states as the reward basis for reinforcement learning, and training a neural network model with the reliability requirement of the distribution network protection service as the constraint condition to optimize the reward, and solving to obtain the optimized configuration of the distribution network distributed protection device. By proposing a quantitative calculation index for the reliability of the system in the case of different numbers of master station protection devices and distributed protection device layout schemes, the present invention solves the problem that traditional protection devices are manually designed and arranged and lack quantitative evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network protection, and in particular to a method and system for optimizing configuration of a distribution network distributed protection device. Background Art

[0002] Distribution networks are characterized by a large number of devices and widespread distribution. Due to limitations in funding and technology, their intelligent development is relatively lagging. Currently, most distribution network relay protection devices still adhere to half-century-old principles, setting, and configuration schemes. However, new features of distribution networks, such as the fluid hierarchical relationships between distributed switches in ring networks and the changing power flow of distributed power sources, make it difficult to improve protection accuracy and flexibility through simple current coordination. Longitudinal fiber optic protection can address the shortcomings of traditional relay protection in terms of accuracy and flexibility. However, the high installation and maintenance costs of laying point-to-point fiber optic cables across such a widespread distribution network make widespread adoption difficult.

[0003] The rapid development of 5G communication technology has provided low-latency, highly reliable information channels for distribution network protection services. In recent years, the application of 5G communication technology in distribution network protection has emerged. Leveraging 5G wireless communication network coverage, it is easy to improve the convenience of installing and maintaining communication channels for protection devices within the distribution area. Rapid logical associations between numerous distributed protection nodes within the 5G coverage area facilitate adaptation to changes in distribution network topology, greatly enhancing the operational flexibility of relay protection systems. 5G wireless networks require no construction, wiring, or maintenance, and newly added nodes quickly combine with other nodes within the 5G coverage area, easily meeting the needs of ready-to-use distributed protection expansion for distribution networks. Distribution network distributed protection systems based on 5G communication have shown broad application prospects in the practice of smart distribution networks and have received widespread attention.

[0004] At present, the distribution network distributed protection system based on 5G communication has realized the basic functions of relay protection. However, while meeting the reliability requirements of distribution network protection services, there is a lack of research on determining the location of distributed protection devices, the master / slave station ratio, and realizing the optimal configuration of distributed protection devices.

[0005] In the related technology, the Chinese invention patent document with publication number CN114678860A records a distribution network protection and control method and system based on deep reinforcement learning, including an intelligent agent obtaining local measurement data and using the obtained perception information as environmental state information for deep reinforcement learning; the intelligent agent obtaining action types and parameters as action space information for deep reinforcement learning; designing a reward function in the process of interaction between the intelligent agent and the environment; constructing a deep reinforcement learning neural network model; training the deep learning neural network; making autonomous decisions on the obtained perception information based on the trained deep reinforcement learning neural network model to obtain instructions for controlling the action of the circuit breaker.

[0006] However, this solution is to enable the protection device to autonomously perceive the operating status of the distribution network, and through continuous trial and error learning, adaptively adjust the protection action strategy to meet the selectivity and speed of protection actions in a highly uncertain distribution network environment; it does not involve the optimal configuration of the protection device.

[0007] Chinese invention patent publication CN114123178A describes a smart grid partitioning and network reconfiguration method based on multi-agent reinforcement learning. The method's implementation steps are: 1) Divide the power grid into N regions based on operational needs and construct the basic elements of multi-agent reinforcement learning, including the environment, agents, states, observations, actions, and reward functions; 2) Run a power system simulation environment to create an initial operational state dataset for the power system; 3) Construct a deep neural network model and train the decision-making agent using reinforcement learning between agents; and 4) Use the trained agent to develop a strategy for grid reconfiguration. This solution addresses the complex problem of post-fault grid reconfiguration, rather than simply optimizing the structure of the distribution network to achieve the optimal configuration. Summary of the invention

[0008] The technical problem to be solved by the present invention is how to achieve optimal configuration of distributed protection devices.

[0009] The present invention solves the above technical problems through the following technical means:

[0010] The present invention proposes a method for optimizing configuration of a distribution network distributed protection device, the method comprising:

[0011] Based on the configuration states of distributed protection devices in the distribution network and combined with the application theory of ultra-reliable low-latency communication, the data transmission error rates corresponding to different configuration states of the distribution network distributed protection system are calculated. The configuration states of the distributed protection devices include the number of master stations and the locations of the distributed protection devices.

[0012] Based on the data transmission error rates corresponding to different configuration states, the reliability of the distribution network distributed protection system corresponding to different configuration states is calculated;

[0013] The reliability corresponding to different configuration states is used as the reward basis for reinforcement learning. The reliability requirements of the distribution network protection service are used as constraints. The neural network model is trained to optimize the reward and obtain the optimal configuration of the distribution network distributed protection device.

[0014] First, based on the method of reinforcement learning, this invention arranges protection devices in the distribution network. The reliability under different arrangements is calculated according to the established model as the basis for reinforcement learning rewards. Constrained by the reliability requirements of the distribution network protection service, iterative learning is continuously carried out to optimize the arrangement of distribution network distributed protection devices and solve the optimal arrangement of distribution network distributed protection devices. By proposing a quantitative calculation index for the reliability of the system in the case of different numbers of master station protection devices and distributed protection device arrangement schemes, it changes the problem that traditional protection devices are manually designed and arranged and lack quantitative evaluation indicators.

[0015] Further, the calculation formula for the data transmission error rate is:

[0016]

[0017] In the formula: ε is the data transmission error rate; W is the bandwidth; SINR i represents the signal-to-interference-plus-noise ratio; V represents the channel dispersion; D tx is the data block length; C is the maximum communication rate; Q is the Gaussian integral function.

[0018] Further, based on the data transmission error rates corresponding to different configuration states, the reliability corresponding to different configuration states of the distribution network distributed protection system is calculated, and the formula is expressed as:

[0019] y = 1 - ε

[0020] In the formula: ε is the data transmission error rate; y is the reliability.

[0021] Further, when using the reliability corresponding to different configuration states as the reward basis for reinforcement learning and training the neural network model with the reliability requirements of the distribution network protection service as the constraint conditions, it includes:

[0022] According to the learning rate, at each time step t, each weight parameter θ of the neural network model i is updated using different learning rates, and the formula is expressed as:

[0023]

[0024] In the formula: η is the learning rate; g t,i is the partial derivative operation of the objective function with respect to the parameter θ at the t-th step of learning i , is the gradient operator, J() is the loss function of the parameter θ t,i ; G t,ii is a diagonal matrix; ξ is a constant term to avoid division by zero; θ t,i is the weight parameter θ at the t-th step of learning i ; θ t+1,iThe weight parameters after updating using the learning rate.

[0025] Further, when performing parameter update, the method further includes:

[0026] Perform a matrix-vector product between G t and g t for vectorization, and the formula is expressed as:

[0027]

[0028] In the formula: η is the learning rate; g t is the gradient of the t-th step of learning, G t is the matrix formed by all g t,i ; ξ is a constant term to avoid division by zero; θ t is the weight parameter at time t; θ t+1 is the weight parameter at time t + 1; ⊙ is the matrix-vector product.

[0029] Further, the neural network model includes an input layer, a first fully connected layer, and a second fully connected layer connected in sequence, and a softmax function is connected after the second fully connected layer;

[0030] The second fully connected layer includes neurons, where H is the number of distributed protection devices.

[0031] In addition, the present invention also proposes an optimization configuration system for a distribution network distributed protection device, and the system includes:

[0032] A transmission error rate calculation module, configured to calculate the data transmission error rate corresponding to different configuration states of the distribution network distributed protection system based on the respective configuration states of the distributed protection devices in the distribution network and in combination with the theory of ultra-reliable low-latency communication applications, where the configuration states of the distributed protection devices include the number of master stations arranged and the positions of the distributed protection devices;

[0033] A reliability calculation module, configured to calculate the reliability corresponding to different configuration states of the distribution network distributed protection system based on the data transmission error rate corresponding to different configuration states;

[0034] A configuration optimization module, configured to use the reliability corresponding to different configuration states as the reward basis for reinforcement learning, and train the neural network model with the reliability requirement of the distribution network protection service as the constraint condition to make the reward reach the optimal value, and solve to obtain the optimized configuration of the distribution network distributed protection device.

[0035] Further, the calculation formula of the data transmission error rate is:

[0036]

[0037] Where: ε is the data transmission error rate; W is the bandwidth; SINR i represents the signal-to-interference noise ratio; V represents the channel dispersion; D tx is the data block length; C is the maximum communication rate; Q is the Gaussian integral function.

[0038] Furthermore, the calculation formula of the reliability is:

[0039] y=1-ε

[0040] Where: ε is the data transmission error rate; y is the reliability.

[0041] Furthermore, the configuration optimization module is specifically used to:

[0042] According to the learning rate, each weight parameter θ of the neural network model is adjusted at each time step t. i Using different learning rates to update, the formula is expressed as:

[0043]

[0044] Where: η is the learning rate; g t,i The objective function learns the parameters θ at step t i The partial derivative operation of is the gradient operator, J() is the parameter θ t,i The loss function of G t,ii is a diagonal matrix; ξ is a constant term to avoid division by zero; θ t,i is the weight parameter θ at the t-th step of learning i θ t+1,i is the weight parameter after updating using the learning rate.

[0045] The advantages of the present invention are:

[0046] (1) The present invention first arranges protection devices in the distribution network based on the reinforcement learning method, calculates the reliability under different arrangements according to the established model as the basis for reinforcement learning rewards, and takes the reliability requirements of the distribution network protection business as a constraint, continuously iterates and learns to achieve the optimal arrangement of the distribution network distributed protection devices, and solves the optimal arrangement of the distribution network distributed protection devices; by proposing to establish quantitative calculation indicators of the system reliability under different numbers of master station protection devices and distributed protection device arrangement schemes, the problem that traditional protection devices are manually designed and arranged and lack quantitative evaluation indicators is changed.

[0047] (2) Using different learning rates to update each weight parameter at each time step and combining it with the gradient descent method can make the weight parameters of the neural network model reach the optimal value faster.

[0048] (3)Regarding the problem that reinforcement learning is prone to falling into local optima, an iterative adaptive parameter optimization method for the layout of distribution network devices based on reinforcement learning is studied to avoid local optima.

[0049] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a schematic flowchart of an optimization configuration method for a distribution network distributed protection device proposed in an embodiment of the present invention;

[0051] Figure 2 is an overall principle block diagram of an optimization configuration method for a distribution network distributed protection device in an embodiment of the present invention;

[0052] Figure 3 is a schematic structural diagram of an optimization configuration system for a distribution network distributed protection device proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] As Figure 1 shown, a first embodiment of the present invention proposes an optimization configuration method for a distribution network distributed protection device, and the method includes the following steps:

[0055] S10. Based on the various configuration states of the distributed protection devices in the distribution network and in combination with the theory of ultra-reliable low-latency communication applications, calculate the data transmission error rates corresponding to different configuration states of the distribution network distributed protection system, where the configuration states of the distributed protection devices include the number of master station layouts and the positions of the distributed protection devices;

[0056] It should be noted that for the distribution network distributed protection system using 5G communication, the distribution network distributed protection system needs to adopt the ultra-reliable low-latency communication (uRLLC) application of 5G communication to ensure low latency and high reliability; based on the various configuration states of the distributed protection devices in the distribution network, the data transmission error rates under different configuration states can be calculated according to the theory of the uRLLC application of 5G communication.

[0057] It should be noted that in the distribution network distributed protection system, different numbers of master station protection devices are arranged at the master station position of the distribution network, and distributed protection devices are arranged at feasible positions such as ring main units and switching stations; the number of master station protection devices arranged at the master station position (which can also be understood as the ratio of master station / slave station) and the arrangement position of the distributed protection devices are used as the configuration status of the distribution network distributed protection system.

[0058] S20. Calculate the reliability corresponding to different configuration states of the distribution network distributed protection system based on the data transmission error rates corresponding to different configuration states.

[0059] It should be noted that the data transmission error rate is an index for configuration optimization, and the data transmission error rate is mutually exclusive with reliability.

[0060] S30. Use the reliability corresponding to different configuration states as the reward basis for reinforcement learning, and train the neural network model with the reliability requirement of the distribution network protection service as the constraint condition to optimize the reward, and solve to obtain the optimal configuration of the distribution network distributed protection device.

[0061] It should be noted that according to the established reliability model, calculate the reliability under different arrangements as the reward basis for reinforcement learning, and train the neural network model with the reliability requirement of the distribution network protection service as the constraint, and continuously iterate and learn to realize the optimal arrangement of the distribution network distributed protection device. By proposing a quantitative calculation index for the reliability of the system in the case of different numbers of master station protection devices and distributed protection device arrangement schemes, the problem that traditional protection devices are manually designed and arranged and lack quantitative evaluation indexes is changed.

[0062] In one embodiment, in the step S10: Based on the various configuration states of the distributed protection devices in the distribution network, combined with the theory of ultra-reliable low-latency communication applications, calculate the data transmission error rates corresponding to different configuration states of the distribution network distributed protection system, and the formula is expressed as follows:

[0063]

[0064] In the formula: ε is the data transmission error rate; W is the bandwidth; SINR i represents the signal-to-interference-plus-noise ratio; V represents the channel dispersion; D tx is the data block length; C is the maximum communication rate; Q is the Gaussian integral function.

[0065] Furthermore, SINR i = αP t g / (I + N0W), where P t is the transmit power; α is the average channel gain for capturing path loss and shadow; g is the normalized instantaneous channel gain; N0 is the unilateral noise spectral density; I is the total interference.

[0066] Furthermore, the expression of channel dispersion V is:

[0067]

[0068] Furthermore, the expression of Gaussian integral function Q is:

[0069]

[0070] Where: e is the natural logarithm; s is the integration interval.

[0071] In one embodiment, the step S20: based on the data transmission error rates corresponding to the different configuration states, the reliability of the distribution network distributed protection system corresponding to the different configuration states is calculated, and the formula is expressed as:

[0072] y=1-ε

[0073] Where: ε is the data transmission error rate; y is the reliability.

[0074] It should be noted that, based on the reliability model constructed in this embodiment, the reliability under different numbers of master stations and distributed arrangements can be calculated.

[0075] In one embodiment, in step S30, the reliability corresponding to different configuration states is used as a reward basis for reinforcement learning, and the reliability requirement of the distribution network protection service is used as a constraint condition. When training the neural network model, the following steps are included:

[0076] According to the learning rate, each weight parameter θ of the neural network model is adjusted at each time step t. i Using different learning rates to update, the formula is expressed as:

[0077]

[0078] Where: η is the learning rate; g t,i The objective function learns the parameters θ at step t i The partial derivative operation of is the gradient operator, J() is the parameter θ t,i The loss function of G t,ii It is a diagonal matrix, each diagonal element is i, i is the gradient θ to the tth step i The sum of the squares of ξ is a constant term to avoid division by zero; θ is a constant term to avoid division by zero; t,i is the weight parameter θ at the t-th step of learning i θ t+1,i is the weight parameter after updating using the learning rate.

[0079] It should be noted that this embodiment calculates the reliability under different numbers of master station layouts and distributed layouts based on the established reliability model, which serves as the reward basis for reinforcement learning. Then, the strategy body neural network model of reinforcement learning is trained with the reliability requirements of the distribution network protection service as a constraint.

[0080] In addition, in order to solve the problem that the traditional gradient descent method is prone to fall into the local optimum during training, the learning rate η is introduced to adapt to the parameter. According to the frequently appearing feature-related parameters, an update function is established to determine the degree of execution update. Different from all existing parameters θ i The idea of updating together, the embodiment of the present invention updates each parameter θ at each time step i Update with different learning rates, and use the new learning rate Selecting weights and choosing different weights for different points at different times is more conducive to convergence and reaching the optimal value.

[0081] In one embodiment, since G t Contains the sum of the squares of the past gradients of all parameters θ along its diagonal. In the embodiment of the present invention, t and g t Perform a matrix-vector product ⊙ between them to vectorize and facilitate the operation:

[0082]

[0083] Where: η is the learning rate; g t is the gradient of learning at step t, G t For all g t,i The matrix formed; C is a constant term to avoid division by zero; θ t is the weight parameter at time t; θ t+1 is the weight parameter at time t+1; ⊙ is the matrix-vector product.

[0084] It should be noted that this embodiment can achieve the purpose of using the adaptive parameter optimization method to improve the disadvantage of local optimum. Figure 2 As shown, according to the optimization algorithm method of adaptive adjustment gradient, the optimal layout problem of the distribution network distributed protection device is solved, and the configuration of the distributed protection device is changed based on the optimal configuration scheme of the distribution network distributed protection device obtained by the solution.

[0085] In one embodiment, the neural network model includes an input layer, a first fully connected layer, and a second fully connected layer connected in sequence, wherein the second fully connected layer is connected to a softmax function;

[0086] The second fully connected layer includes neurons, where H is the number of distributed protection devices.

[0087] Among them, the input of the input layer is the state of reinforcement learning, the output of the second fully connected layer is all behavior values, and finally all behavior values output by the second fully connected layer are converted into the probability of the corresponding action through the softmax function.

[0088] In addition, if Figure 3 As shown, the second embodiment of the present invention further proposes an optimization configuration system for a distribution network distributed protection device, the system comprising:

[0089] The transmission error rate calculation module 10 is used to calculate the data transmission error rate corresponding to different configuration states of the distribution network distributed protection system based on the configuration states of the distributed protection devices in the distribution network and in combination with the application theory of ultra-reliable low-latency communication, wherein the configuration state of the distributed protection device includes the number of master stations and the location of the distributed protection device;

[0090] Reliability calculation module 20, used to calculate the reliability of the distribution network distributed protection system corresponding to different configuration states based on the data transmission error rates corresponding to different configuration states;

[0091] The configuration optimization module 30 is used to use the reliability corresponding to different configuration states as the reward basis for reinforcement learning, and to train the neural network model with the reliability requirements of the distribution network protection service as the constraint condition to optimize the reward and obtain the optimal configuration of the distribution network distributed protection device.

[0092] This embodiment uses a reinforcement learning method to deploy different numbers of master station protection devices at the distribution network master station, and to deploy distributed protection devices at feasible locations such as ring main units and switchgear. The reliability of different deployments is calculated based on the established model and used as the basis for reinforcement learning rewards. With the reliability requirements of the distribution network protection service as a constraint, continuous iterative learning is used to achieve the optimal deployment of the distribution network distributed protection devices.

[0093] In one embodiment, the calculation formula for the data transmission error rate is:

[0094]

[0095] Where: ε is the data transmission error rate; W is the bandwidth; SINR i represents the signal-to-interference noise ratio; V represents the channel dispersion; D tx is the data block length; C is the maximum communication rate; Q is the Gaussian integral function.

[0096] In one embodiment, the reliability calculation formula is:

[0097] y=1-ε

[0098] Where: ε is the data transmission error rate; y is the reliability.

[0099] In one embodiment, the configuration optimization module is specifically configured to:

[0100] According to the learning rate, each weight parameter θ of the neural network model is adjusted at each time step t. i Using different learning rates to update, the formula is expressed as:

[0101]

[0102] Where: η is the learning rate; g t,i The objective function learns the parameters θ at step t i The partial derivative operation of is the gradient operator, J() is the parameter θ t,i The loss function of G t,ii is a diagonal matrix; g t is the gradient of learning at step t, G t For all g t,i The matrix formed; ξ is a constant term to avoid division by zero; θ t,i is the weight parameter θ at the t-th step of learning i θ t+1,i is the weight parameter after updating using the learning rate.

[0103] In one embodiment, the configuration optimization module 30 is further configured to:

[0104] In G t and g t Perform matrix-vector product between them for vectorization, and the formula is expressed as:

[0105]

[0106] Where: η is the learning rate; g t is the gradient of learning at step t, G t For all g t,i The matrix formed; ξ is a constant term to avoid division by zero; θ t is the weight parameter at time t; θ t+1 is the weight parameter at time t+1; ⊙ is the matrix-vector product.

[0107] This embodiment addresses the problem that reinforcement learning is prone to falling into local optimality. By studying an iterative adaptive parameter optimization method for distribution network device layout based on reinforcement learning, local optimality is avoided and the optimal layout of distributed protection devices in the distribution network is solved.

[0108] It should be noted that other embodiments or implementation methods of the optimization configuration system of the distribution network distributed protection device of the present invention can refer to the above-mentioned method embodiments, which will not be repeated here.

[0109] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus or device and execute the instructions), or in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in combination with an instruction execution system, apparatus or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation or other suitable processing as necessary, and then stored in a computer memory.

[0110] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0111] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0112] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0113] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. An optimization configuration method for a distribution network distributed protection device, characterized in that, The method includes: Based on the configuration states of distributed protection devices in the distribution network and combined with the theory of ultra-reliable and low-latency communication applications, calculate the data transmission error rates corresponding to different configuration states of the distribution network distributed protection system, where the configuration states of the distributed protection devices include the number of master stations arranged and the positions of the distributed protection devices; Based on the data transmission error rates corresponding to different configuration states, calculate the reliabilities corresponding to different configuration states of the distribution network distributed protection system; Take the reliabilities corresponding to different configuration states as the reward basis for reinforcement learning, and use the reliability requirements of the distribution network protection service as the constraint conditions to train the neural network model to optimize the reward and solve for the optimal configuration of the distribution network distributed protection devices.

2. The optimization configuration method of the distribution network distributed protection device according to claim 1, characterized in that The calculation formula for the data transmission error rate is: Where: ε is the data transmission error rate; W is the bandwidth; SINR i represents the signal-to-interference-plus-noise ratio; V represents the channel dispersion; D tx is the data block length; C is the maximum communication rate; Q is the Gaussian integral function.

3. The optimization configuration method of the distribution network distributed protection device according to claim 1, characterized in that, Based on the data transmission error rates corresponding to different configuration states, calculate the reliabilities corresponding to different configuration states of the distribution network distributed protection system, and the formula is expressed as: y = 1 - ε In the formula: ε is the data transmission error rate; y is the reliability.

4. The optimization configuration method of the distribution network distributed protection device according to claim 1, characterized in that When taking the reliabilities corresponding to different configuration states as the reward basis for reinforcement learning and using the reliability requirements of the distribution network protection service as the constraint conditions to train the neural network model, it includes: At each time step t, each weight parameter θ of the neural network model is updated according to the learning rate i using different learning rates, which is expressed by the formula: Where: η is the learning rate; g t,i is the partial derivative operation of the objective function with respect to the parameter θ i at the t-th step of learning, is the gradient operator, J() is the loss function of the parameter θ t,i ; G t,ii is a diagonal matrix; ξ is a constant term to avoid division by zero; θ t,i is the weight parameter θ at the t-th step of learning i ; θ t+1,i is the weight parameter updated using the learning rate.

5. The optimization configuration method of the distribution network distributed protection device according to claim 4, characterized in that, When performing parameter updates, the method further includes: Between G t and g t Perform a matrix-vector product and vectorize it, which is expressed by the formula: Where: η is the learning rate; g t is the gradient of the t-th step of learning, and G t is the matrix formed by all g t,i ; ξ is a constant term to avoid division by zero; θ t is the weight parameter at time t; θ t+1 is the weight parameter at time t + 1; ⊙ is the matrix-vector product.

6. The optimization configuration method of the distribution network distributed protection device according to claim 1, characterized in that The neural network model includes an input layer, a first fully connected layer, and a second fully connected layer connected in sequence, and a softmax function is connected after the second fully connected layer; The second fully-connected layer includes neurons, where H is the number of distributed protection devices.

7. An optimized configuration system for a distribution network distributed protection device, characterized in that, The system includes: A transmission error rate calculation module, which is used to calculate the data transmission error rates corresponding to different configuration states of the distribution network distributed protection system based on the configuration states of the distributed protection devices in the distribution network and combined with the theory of ultra-reliable and low-latency communication applications, where the configuration states of the distributed protection devices include the number of master stations arranged and the positions of the distributed protection devices; A reliability calculation module, which is used to calculate the reliabilities corresponding to different configuration states of the distribution network distributed protection system based on the data transmission error rates corresponding to different configuration states; A configuration optimization module, which is used to take the reliabilities corresponding to different configuration states as the reward basis for reinforcement learning, and use the reliability requirements of the distribution network protection service as the constraint conditions to train the neural network model to optimize the reward and solve for the optimal configuration of the distribution network distributed protection devices.

8. The optimization configuration system of the distribution network distributed protection device according to claim 7, characterized in that, The calculation formula for the data transmission error rate is: Where: ε is the data transmission error rate; W is the bandwidth; SINR i represents the signal-to-interference-plus-noise ratio; V represents the channel dispersion; D tx is the data block length; C is the maximum communication rate; Q is the Gaussian integral function.

9. The optimization configuration system of the distribution network distributed protection device according to claim 7, characterized in that The calculation formula for the reliability is: y = 1 - ε In the formula: ε is the data transmission error rate; y is the reliability.

10. The optimization configuration system of the distribution network distributed protection device according to claim 7, characterized in that, The configuration optimization module is specifically used for: At each time step t, each weight parameter θ of the neural network model is updated according to the learning rate i using different learning rates, which is expressed by the formula: where: η is the learning rate; g t,i is the partial derivative operation of the objective function with respect to the parameter θ i at the t-th step of learning, is the gradient operator, J() is the loss function of the parameter θ t,i ; G t,ii is a diagonal matrix; g t is the gradient at the t-th step of learning, G t is the matrix formed by all g t,i ; ξ is a constant term to avoid division by zero; θ t,i is the weight parameter θ at the t-th step of learning i ; θ t+1,i is the weight parameter updated using the learning rate.

Citation Information

Patent Citations

  • Smart power grid partition network reconstruction method based on multi-agent reinforcement learning

    CN114123178A

  • Power distribution network protection control method and system based on deep reinforcement learning

    CN114678860A

  • Protection device proportion optimization method of 5G distribution network distributed protection system

    CN114697200A