Reinforcement learning-based edge computing smart power grid resource scheduling method and system

By applying the edge computing resource scheduling method based on reinforcement learning in the smart grid, the problems of insufficient real-time data processing capabilities and inaccurate feature information extraction are solved, efficient resource scheduling in a dynamic environment is achieved, and the stability and efficiency of the power grid are improved.

CN119944602APending Publication Date: 2025-05-06GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411702925.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing smart grid resource scheduling methods have insufficient real-time data processing capabilities, inaccurate feature information extraction, and the difficulty of achieving efficient resource scheduling in dynamic and complex environments.

Method used

Adopt edge computing smart grid resource scheduling method based on reinforcement learning. By building a smart grid environment, the agent at key nodes is initialized and the initial resource scheduling strategy is assigned, real-time data points are collected, abnormal data is filtered and key feature information is extracted, and the agent status is updated, and the scheduling strategy is optimized.

Benefits of technology

It improves the effectiveness and scalability of the scheduling strategy, enhances the agent's understanding of complex grid topology, realizes efficient data screening and processing, improves the stability and efficiency of power grid operation, and ensures efficient and reliable utilization of power grid resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944602A_ABST
    Figure CN119944602A_ABST
Patent Text Reader

Abstract

The invention discloses an edge computing intelligent power grid resource scheduling method and system based on reinforcement learning, and relates to the technical field of intelligent power grid resource scheduling, and the method comprises the steps: constructing an intelligent power grid environment, initializing an intelligent agent at a key node, and endowing an initial resource scheduling strategy; collecting real-time data points, screening abnormal data and extracting key feature information; and updating the state of the intelligent agent by using the key feature information, and generating and optimizing a scheduling strategy. According to the method, by collecting real-time data points, screening abnormal data and extracting key feature information, efficient data screening and processing are achieved, an intelligent agent can make a decision quickly and accurately when facing a complex and dynamic power grid environment, the state of the intelligent agent is updated by using the key feature information, and the intelligent agent can be quickly and accurately determined. The scheduling strategy is generated and optimized, so that the whole performance of the system can be improved while the intelligent agent continuously optimizes the scheduling strategy, and efficient and reliable utilization of power grid resources is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart grid resource scheduling, and specifically to an edge computing smart grid resource scheduling method and system based on reinforcement learning. Background Art

[0002] With the advancement of the Internet of Things, big data analysis and artificial intelligence technologies, modern power grids are undergoing a revolution and are moving towards a more flexible, efficient and reliable direction. Edge computing reduces cloud transmission delays and improves response speed by processing data at the edge of the network. At the same time, reinforcement learning, as a branch of machine learning, enables the system to learn and optimize its strategies autonomously, and is particularly suitable for dealing with decision-making problems in complex dynamic environments. Combining the two has opened up new possibilities for smart grid resource scheduling. However, although advanced technologies such as reinforcement learning and edge computing have been introduced into the field of smart grid resource scheduling, they still face many technical bottlenecks in practical applications.

[0003] When faced with massive real-time data streams, it is impossible to effectively and accurately extract key feature information, resulting in a significant reduction in the optimization efficiency, scientificity and accuracy of subsequent resource scheduling strategies. This is obviously a major challenge for smart grids that require real-time response and rapid decision-making. Since smart grids need to respond to various dynamic changes in a short period of time, such as load fluctuations, changes in renewable energy power generation, and power grid failures, if key feature information cannot be extracted and processed in a timely and accurate manner, it will be difficult for the intelligent agent to make the best decision, which may lead to untimely and inaccurate resource scheduling, and even cause problems such as grid instability or power outages. Summary of the invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by the present invention is: the existing smart grid resource scheduling method has insufficient real-time data processing capabilities, inaccurate feature information extraction, and the problem of how to achieve efficient resource scheduling in a dynamic and complex environment.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: an edge computing smart grid resource scheduling method based on reinforcement learning, comprising building a smart grid environment, initializing intelligent agents at key nodes and assigning initial resource scheduling strategies; collecting real-time data points, screening abnormal data and extracting key feature information; using key feature information to update the state of the intelligent agent, and generating and optimizing the scheduling strategy.

[0007] As a preferred solution of the edge computing smart grid resource scheduling method based on reinforcement learning described in the present invention, wherein: the construction of the smart grid environment includes defining smart grid elements and connection relationships.

[0008] As a preferred solution of the edge computing smart grid resource scheduling method based on reinforcement learning described in the present invention, wherein: the initialization of key nodes includes identifying key nodes, determining the type and location of key nodes, and assigning unique identifiers to key nodes.

[0009] As a preferred solution of the edge computing smart grid resource scheduling method based on reinforcement learning described in the present invention, wherein: the collection of real-time data points and the screening of abnormal data include collecting node operation data in real time through edge devices, calculating the Z-score value of the real-time data point, obtaining the Z-score standard value of the real-time data point according to the Z-score value of the real-time data point, and performing analysis on the difference between the Z-score value of the real-time data point and the Z-score standard value; if the current difference is less than or equal to two times the standard deviation, the current difference is a normal difference; if the current difference is greater than two times the standard deviation, the current difference is an abnormal difference; the abnormal difference is corrected, and the corrected difference is calculated, which is expressed as:

[0010]

[0011] Among them, D corrected is the corrected difference, D1 and D2 are two adjacent normal differences, Z 1( X1) is the Z-score value of the real-time data point X1, Z2 ( X2) is the Z-score value of the real-time data point X2, Z abnormal ( X abnormal ) is the real-time data point X abnormal Z-score value; aggregate the corrected difference and the normal difference to form a difference set, and perform analysis on the differences in the difference set; if the current difference in the difference set is within the threshold interval, retain the real-time data point corresponding to the current difference; if the current difference in the difference set is not within the threshold interval, remove the real-time data point corresponding to the current difference; obtain the intermediate standard threshold according to the mean of the upper and lower threshold limits, assign weights to the retained real-time data points, and obtain the retained real-time data points with priority.

[0012] As a preferred solution of the edge computing smart grid resource scheduling method based on reinforcement learning described in the present invention, the extraction of key feature information includes transferring the retained and prioritized real-time data points to the edge computing node for preprocessing, and extracting the prioritized feature information based on the preprocessing results.

[0013] As a preferred solution of the edge computing smart grid resource scheduling method based on reinforcement learning described in the present invention, wherein: the use of key feature information to update the state of the intelligent body includes updating the initial state of the initialized intelligent body through feature information with priority, generating an optimized resource scheduling strategy and executing it.

[0014] As a preferred solution of the edge computing smart grid resource scheduling method based on reinforcement learning described in the present invention, wherein: the generation and optimization of the scheduling strategy includes storing the experience of the optimized resource scheduling strategy executed by the intelligent agent to form an experience pool, and using the experience pool to continuously update the optimized resource scheduling strategy.

[0015] Another object of the present invention is to provide an edge computing smart grid resource scheduling system based on reinforcement learning, which can collect real-time data points, filter abnormal data and extract key feature information, thereby solving the problem of inaccurate feature information extraction in current smart grid resource scheduling technology.

[0016] As a preferred solution of the edge computing smart grid resource scheduling system based on reinforcement learning described in the present invention, it includes: an initial module, a data processing module, and an optimization strategy module; the initial module is used to build a smart grid environment, initialize the intelligent agent at the key nodes and assign an initial resource scheduling strategy; the data processing module is used to collect real-time data points, filter abnormal data and extract key feature information; the optimization strategy module is used to use key feature information to update the intelligent agent state, generate and optimize the scheduling strategy.

[0017] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a step of an edge computing smart grid resource scheduling method based on reinforcement learning.

[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of an edge computing smart grid resource scheduling method based on reinforcement learning.

[0019] Beneficial effects of the present invention: The edge computing smart grid resource scheduling method based on reinforcement learning provided by the present invention constructs a smart grid environment, initializes the intelligent agent at the key node and assigns the initial resource scheduling strategy, so that the intelligent agent can learn in a real simulated dynamic environment, improves the effectiveness and scalability of the scheduling strategy, and improves the intelligent agent's ability to understand complex power grid topologies, providing more reliable input conditions for optimizing scheduling. By collecting real-time data points, screening abnormal data and extracting key feature information, efficient data screening and processing are achieved, so that the intelligent agent can make decisions quickly and accurately when facing a complex and dynamic power grid environment, thereby improving the stability and efficiency of power grid operation. By using key feature information to update the intelligent agent state, generate and optimize the scheduling strategy, the intelligent agent can improve the overall performance of the system while continuously optimizing the scheduling strategy, ensuring that power grid resources are efficiently and reliably utilized, thereby ultimately achieving low-cost, low-carbon, and high-efficiency power grid management. The present invention achieves better results in terms of efficiency, accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0021] Figure 1 An overall flow chart of an edge computing smart grid resource scheduling method based on reinforcement learning provided for the first embodiment of the present invention.

[0022] Figure 2 A module diagram of an edge computing smart grid resource scheduling system based on reinforcement learning provided for the third embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0024] Example 1, reference Figure 1 , is an embodiment of the present invention, and provides an edge computing smart grid resource scheduling method based on reinforcement learning, comprising:

[0025] S1: Build a smart grid environment, initialize the intelligent agents at key nodes and assign initial resource scheduling strategies.

[0026] Furthermore, building a smart grid environment includes defining smart grid elements and their connection relationships.

[0027] It should be noted that initializing the key nodes includes identifying the key nodes, determining the type and location of the key nodes, and assigning a unique identifier to the key nodes.

[0028] It should also be noted that to build a smart grid environment, the components of the smart grid are first defined, including but not limited to power stations, energy storage systems, power transmission lines and consumer terminals; then the connection relationship between the various elements is defined, including but not limited to defining which transmission lines each power station is connected to, as well as the capacity and transmission loss of the transmission lines, defining how the transmission lines are connected to each substation, as well as the capacity and conversion efficiency of the substations, defining how the substations supply power to consumer terminals through distribution lines, defining how the energy storage system is connected to the grid, as well as its charging and discharging capabilities, etc.; then, the operating parameters of the grid are set, including but not limited to power demand, power generation capacity, the status of the energy storage system and the topology of the grid.

[0029] It should also be noted that when identifying key nodes, the key node types are first determined, including but not limited to substations, which are responsible for converting high voltage electricity into low voltage electricity and distributing it to various users; microgrid controllers, which are responsible for managing power generation, energy storage and power consumption equipment within the microgrid; large power stations, which have an important impact on the overall power supply capacity of the power grid; important load nodes, which are large industrial users or critical infrastructure that have an important impact on the stability of the power grid; then identify the location of the key nodes, including but not limited to the geographical location, and determine the geographical location of the key nodes based on the power grid topology map; functional importance, determine the key nodes based on the function of the node and its importance to the operation of the power grid; finally, assign identifiers to the key nodes, including assigning a unique identifier to each key node to facilitate subsequent data processing and communication.

[0030] It should also be noted that when initializing the agent at the key node, the structure of the agent is first defined, including state representation, which includes but is not limited to the current power demand, power generation, energy storage status, transmission line status, etc.; action space, which defines the actions that the agent can perform, such as adjusting the output power of the power station, controlling the charging and discharging of the energy storage system, adjusting the output voltage of the substation, etc.; reward function, which defines the reward function of the agent, such as reducing operating costs, improving power supply reliability, and reducing carbon emissions; then the agent is initialized for each key node, including creating an agent instance, creating an agent instance for each key node; setting the initial state, setting the initial state for each agent according to the initial state of the power grid, for example, the initial power generation of the power station, the initial storage capacity of the energy storage system, the initial output voltage of the substation, etc.; setting the initial strategy, setting the initial resource scheduling strategy for each agent, the initial strategy can be a simple rule-based strategy or a random strategy for the agent to gradually optimize during the learning process.

[0031] It should also be noted that by constructing a smart grid environment and initializing the intelligent agents at key nodes, the foundation can be laid for the subsequent grid resource scheduling. First, by defining the components of the smart grid and the connection relationship between the elements, the physical and information structure of the grid is clear, and the functions of each node in the grid are reasonably planned, the overall framework of the system is clear, and the scientific and coordinated resource scheduling is ensured. By clearly identifying key nodes such as substations and microgrid controllers and assigning unique identifiers to them, data processing and communication are smooth, thereby providing an efficient information exchange and feedback mechanism for the intelligent agent to perform optimization tasks. When initializing the intelligent agent at the key nodes, the state representation, action space and reward function of each node are set, so that the intelligent agent can make effective decisions based on the current state of the grid. By creating an intelligent agent instance for each node and assigning an initial resource scheduling strategy, the rationality of the initial scheduling is guaranteed, and a good starting point is provided for the subsequent optimization of the scheduling strategy through reinforcement learning, which helps to achieve the real-time and autonomy of grid scheduling, thereby maximizing the utilization efficiency of grid resources and reducing energy waste.

[0032] S2: Collect real-time data points, filter abnormal data and extract key feature information.

[0033] Furthermore, collecting real-time data points and filtering abnormal data includes collecting node operation data in real time through edge devices, calculating the Z-score value of the real-time data point, obtaining the Z-score standard value of the real-time data point according to the Z-score value of the real-time data point, and performing analysis on the difference between the Z-score value of the real-time data point and the Z-score standard value; if the current difference is less than or equal to two times the standard deviation, the current difference is a normal difference; if the current difference is greater than two times the standard deviation, the current difference is an abnormal difference; performing correction on the abnormal difference, first obtaining two adjacent normal differences D1 and D2, and obtaining the Z-score value of the real-time data point X1 based on the difference D1, and obtaining the Z-score value of the real-time data point X2 based on the difference D2, and then, based on the abnormal difference D abnormal Get real-time data point X abnormal Finally, according to the Z-score value of the two adjacent normal differences D1 and D2, the Z-score value of the real-time data points X1 and X2, and the real-time data point X abnormal The Z-score value of the corrected difference is calculated as:

[0034]

[0035] Among them, D corrected is the corrected difference, D1 and D2 are two adjacent normal differences, Z 1( X1) is the Z-score value of the real-time data point X1, Z2 ( X2) is the Z-score value of the real-time data point X2, Z abnormal ( X abnormal ) is the real-time data point X abnormal Z-score value; aggregate the corrected difference and the normal difference to form a difference set, and perform analysis on the differences in the difference set; if the current difference in the difference set is within the threshold interval, retain the real-time data point corresponding to the current difference; if the current difference in the difference set is not within the threshold interval, remove the real-time data point corresponding to the current difference; obtain the intermediate standard threshold according to the mean of the upper and lower threshold limits, assign weights to the retained real-time data points, and obtain the retained real-time data points with priority.

[0036] It is worth noting that, from the perspective of data concentration, the intermediate standard threshold represents the central trend of the data. The closer the data point is to the intermediate standard threshold, the closer its value is to the average value of the data, representing the typical characteristics of the data. Combined with the perspective of abnormal data exclusion, by setting the upper and lower thresholds, extremely abnormal data points have been excluded, and the remaining data points are relatively stable and reliable. The closer the data point is to the intermediate standard threshold, the higher its credibility and representativeness. Combined with the perspective of data quality, the closer the data point is to the intermediate standard threshold, the smaller its fluctuation is, and it is more in line with the normal data distribution. The quality of such data points is higher and can better reflect the true state of the smart grid.

[0037] It should be noted that extracting key feature information includes transferring the retained and prioritized real-time data points to the edge computing node for preprocessing, and extracting prioritized feature information based on the preprocessing results.

[0038] It should also be noted that feature information with priority is extracted based on the preprocessing results, where the preprocessing includes data cleaning, standardization, data verification and feature engineering. It is worth noting that since real-time data points have priority, high-priority real-time data points should be given priority during the preprocessing process, and high-priority feature information should be extracted first.

[0039] It should also be noted that by collecting real-time data points and screening abnormal data, accurate monitoring of the operating status of the power grid is achieved. In this process, the data collected by the edge device is analyzed by Z-score, which can effectively identify and correct abnormal data, ensure the accuracy of subsequent data analysis, and filter out abnormal data points that may cause scheduling errors by correcting the abnormal difference, providing more real and reliable data support for subsequent resource scheduling decisions. This step effectively solves the interference problem of abnormal data in massive data in the power grid and ensures the efficiency of intelligent scheduling decisions. Extracting key feature information ensures that only the most representative real-time data points are retained. These data points are crucial for intelligent decision-making and scheduling strategy optimization. By screening priority data points, the feature extraction process is optimized, which not only improves data processing efficiency, but also provides reliable basic data for subsequent scheduling strategy optimization.

[0040] S3: Use key feature information to update the agent state and generate and optimize the scheduling strategy.

[0041] Furthermore, using key feature information to update the state of the intelligent agent includes updating the initial state of the initialized intelligent agent through feature information with priority, generating an optimized resource scheduling strategy and executing it.

[0042] It should be noted that generating and optimizing the scheduling strategy includes storing the experience of optimizing the resource scheduling strategy executed by the intelligent agent to form an experience pool, and using the experience pool to continuously update the optimized resource scheduling strategy.

[0043] It should also be noted that the key feature information is used to update the agent state and generate and optimize the scheduling strategy, which specifically includes the following steps:

[0044] First, the feature information with priority is read, and then the state representation is updated. The state representation of the agent usually includes multiple dimensions, such as power demand, power generation, energy storage status, transmission line status, etc. When the state is updated, it includes updating the current power demand value, updating the power generation of each power station, updating the current storage capacity and charging and discharging status of the energy storage system, and updating the current load and status of the transmission line; then, based on the updated feature information, the state variables of the agent are recalculated to obtain a new state representation instead of the initial state. Then, based on the updated state representation, the agent evaluates the current operation status of the power grid, for example, judging the current power supply and demand balance, the available capacity of the energy storage system, the load of the transmission line, etc.; then, the action space of the agent should be clarified, including adjusting the output power of the power station, controlling the charging and discharging of the energy storage system, adjusting the output voltage of the substation, etc., and then the action selection is performed, that is, according to the current state and the initial strategy of the agent, the optimal action is selected, such as adjusting the output power of the power station. Finally, based on the selected action, an optimized resource scheduling strategy is generated, for example, a strategy set including the output power of the power station, the charging and discharging instructions of the energy storage system, and the output voltage adjustment of the substation is generated and executed.

[0045] It should also be noted that the experience pool is used to continuously update the optimized resource scheduling strategy. Specifically, each time the agent executes the optimized resource scheduling strategy, the experience is stored in the experience pool in the form of a tuple. The tuple contains the state, action, reward and next state. The experience pool is a circular buffer. When the maximum capacity is reached, the old experience will be overwritten by the new experience. When the experience pool is used to update the optimized resource scheduling strategy, sampling is performed first, that is, a batch of experience is randomly drawn from the experience pool, and only this batch of experience is used for training, that is, this batch of experience is used to update the agent's policy network, and the above sampling and training process is repeated every a certain number of time steps or rounds to gradually optimize the agent's strategy.

[0046] It should also be noted that by using key feature information to update the agent state and generate an optimized scheduling strategy, the agent can adjust its operation in real time according to the current state of the power grid, thereby continuously optimizing the resource scheduling strategy. Based on the updated state information of the agent, multiple elements such as power generation, energy storage and load in the power grid are dynamically adjusted. This process reflects the advantages of reinforcement learning in power grid scheduling. It can gradually accumulate optimization strategies through the historical experience pool and finally generate the optimal scheduling plan. In this process, each optimization decision of the agent will be recorded and stored in the experience pool. The update mechanism of the experience pool ensures the continuity and adaptability of the learning process, allowing the agent to continuously adjust its strategy according to changes in the power grid. This not only improves the decision-making accuracy of the agent in complex environments, but also enables the power grid to respond more flexibly to emergencies and load fluctuations.

[0047] Example 2 is an embodiment of the present invention, which provides an edge computing smart grid resource scheduling method based on reinforcement learning. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0048] This experiment is conducted in a simulated smart grid environment, which includes power stations, energy storage systems, power transmission lines, and consumer terminals. The dynamic and complex environment simulation is completed by introducing power demand fluctuations and random failures in different time periods. The specific steps of the experiment are as follows:

[0049] 1. Dynamic construction of smart grid environment:

[0050] Ten key nodes are defined, covering large power stations (2), energy storage systems (2), substations (4) and important load nodes (2).

[0051] Set up dynamic scenarios, such as randomly increasing power demand by 30% to 50% during peak hours and randomly disconnecting one transmission line to simulate a sudden failure.

[0052] A transmission line and power loss model is established, with the capacity set to 200MW and the loss dynamic range to 5%±1%.

[0053] 2. Real-time data collection and processing:

[0054] Use edge devices to collect real-time operating data of key nodes (such as power, voltage, and line load).

[0055] For abnormal data, Z-score analysis and correction are used. For example, 10% of the sampled data are simulated as abnormal values ​​and correction is performed using adjacent normal data points.

[0056] Extract features of real-time data including load distribution, voltage stability and energy storage system status.

[0057] 3. Optimization of resource scheduling strategy:

[0058] Based on the reinforcement learning method, the agent updates the state representation of the collected key feature information.

[0059] Dynamically adjust the power output of power stations, the charging and discharging of energy storage systems, and the output voltage of substations.

[0060] Continuously optimize strategies through the experience pool to ensure that resource scheduling can adapt to complex environmental changes.

[0061] Refer to Table 1 for a comparative analysis of the experimental data.

[0062] Table 1 Experimental data record table

[0063]

[0064]

[0065] It can be seen from the data in Table 1 that the present invention has shown obvious advantages in real-time data processing capability, feature information extraction accuracy and scheduling efficiency in dynamic environments. Through the efficient correction of abnormal data by edge computing devices, the data accuracy of power plants and important load nodes is significantly improved, and the final data of all nodes are successfully used for feature extraction and resource scheduling. With the support of data cleaning and correction, the load distribution efficiency has been improved. For example, the efficiency of important load 1 has been improved from 78% to 81%, which effectively reflects the importance of the high priority characteristics of feature extraction for efficient scheduling. In addition, the performance of the present invention in a dynamic environment is particularly outstanding, and the response time of power grid faults is significantly shortened. For example, the response time of power station 1 is 120 seconds, which is 30 seconds less than that of traditional scheduling. At the same time, the optimized resource scheduling strategy reduces the operating cost by an average of 12% and the transmission loss by 0.3 percentage points. In general, the present invention improves the efficiency and reliability of resource scheduling, effectively responds to dynamic load changes and sudden failures in complex environments, and fully guarantees the stable operation of smart grids.

[0066] Example 3, reference Figure 2 , as an embodiment of the present invention, provides an edge computing smart grid resource scheduling system based on reinforcement learning, including an initial module, a data processing module, and an optimization strategy module.

[0067] The initial module is used to build the smart grid environment, initialize the intelligent agents at key nodes and assign initial resource scheduling strategies; the data processing module is used to collect real-time data points, filter abnormal data and extract key feature information; the optimization strategy module is used to use key feature information to update the intelligent agent status, generate and optimize the scheduling strategy.

[0068] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0069] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0070] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0071] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not limited. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.

[0072] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for resource scheduling of edge computing smart grid based on reinforcement learning, characterized in that: include: Build a smart grid environment, initialize the intelligent agents at key nodes and assign initial resource scheduling strategies; Collect real-time data points, filter abnormal data and extract key feature information; Use key feature information to update the agent state and generate and optimize the scheduling strategy.

2. The edge computing smart grid resource scheduling method based on reinforcement learning as claimed in claim 1, characterized in that: The construction of the smart grid environment includes defining smart grid elements and connection relationships.

3. The edge computing smart grid resource scheduling method based on reinforcement learning as claimed in claim 2, characterized in that: Initializing the key nodes includes identifying the key nodes, determining the type and location of the key nodes, and assigning a unique identifier to the key nodes.

4. The edge computing smart grid resource scheduling method based on reinforcement learning as claimed in claim 3 is characterized by: The collecting of real-time data points and screening of abnormal data includes collecting node operation data in real time through edge devices, calculating Z-score values ​​of real-time data points, obtaining Z-score standard values ​​of real-time data points according to the Z-score values ​​of real-time data points, and performing analysis on the difference between the Z-score values ​​of real-time data points and the Z-score standard values; If the current difference is less than or equal to two times the standard deviation, the current difference is a normal difference; If the current difference is greater than two times the standard deviation, the current difference is an abnormal difference; Correct the abnormal difference and calculate the corrected difference, which is expressed as: Among them, D corrected is the corrected difference, D1 and D2 are two adjacent normal differences, Z1(X1) is the Z-score value of the real-time data point X1, Z2(X2) is the Z-score value of the real-time data point X2, and Z abnormal (X abnormal ) is the real-time data point X abnormal Z-score value; Aggregating the corrected differences and the normal differences to form a difference set, and performing analysis on the differences in the difference set; If the current difference in the difference set is within the threshold range, the real-time data point corresponding to the current difference is retained; If the current difference in the difference set is not within the threshold range, the real-time data point corresponding to the current difference is removed; According to the average of the upper and lower thresholds, the intermediate standard threshold is obtained, and the retained real-time data is The data points are weighted to obtain the reserved and prioritized real-time data points.

5. The edge computing smart grid resource scheduling method based on reinforcement learning as claimed in claim 4, characterized in that: The extracting of key feature information includes transferring the reserved and prioritized real-time data points to the edge computing node for preprocessing, and extracting prioritized feature information based on the preprocessing results.

6. The edge computing smart grid resource scheduling method based on reinforcement learning as claimed in claim 5, characterized in that: The method of updating the state of the intelligent agent by using the key feature information includes updating the initial state of the initialized intelligent agent by using the feature information with priority, generating and executing an optimized resource scheduling strategy.

7. The edge computing smart grid resource scheduling method based on reinforcement learning as claimed in claim 6, characterized in that: The generating and optimizing scheduling strategies includes storing the experience of optimizing resource scheduling strategies executed by the intelligent agent to form an experience pool, and continuously updating the optimizing resource scheduling strategies using the experience pool.

8. A system using the edge computing smart grid resource scheduling method based on reinforcement learning as described in any one of claims 1 to 7, characterized in that: Including initial module, data processing module, and optimization strategy module; The initial module is used to build a smart grid environment, initialize the intelligent agents at key nodes and assign initial resource scheduling strategies; The data processing module is used to collect real-time data points, filter abnormal data and extract key feature information; The optimization strategy module is used to update the agent state using key feature information and generate and optimize the scheduling strategy.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the edge computing smart grid resource scheduling method based on reinforcement learning described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the edge computing smart grid resource scheduling method based on reinforcement learning described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Equipment security measure automatic generation rule base construction method and system

    CN120338088A

  • A method and system for constructing a rule base for automatically generating equipment safety measures

    CN120338088B

  • Power grid island detection and protection method and device based on intelligent agent

    CN120582224A