In-situ data storage management method based on deep reinforcement learning assisted optimization in smart power grid scene
By adopting deep reinforcement learning optimization methods in the smart grid environment, combining the regular division mechanism and distributed queue architecture, the C51-ER algorithm is used to optimize the dequeuing management of in-situ data storage tasks, solving the complexity and inefficiency of in-situ data storage task management in the smart grid, and achieving efficient and low-latency storage task processing.
Patent Information
- Application Number
- CN202510014890.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
The management of in-situ data storage tasks in smart grid environments is complex and inefficient. Especially when dealing with different types of storage tasks, how to make efficient decisions and manage unified management becomes a challenge.
The optimization method based on deep reinforcement learning is adopted, and by establishing a regular division mechanism and distributed queue architecture, combined with the classification deep Q network (C51-ER) algorithm optimized by empirical playback optimization, the dequeue management of in-situ data storage tasks is optimized, and the minimum of stranding tasks and delays is pursued.
It realizes high-efficiency and low-latency in-situ data storage task management in smart grid environments, improves the system's response speed and processing efficiency, and reduces network bandwidth consumption and latency.
Smart Images

Figure FT_1 
Figure FT_2 
Figure QLYQS_4
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of edge computing, and specifically relates to the management of in-situ data storage tasks in power grid scenarios. Background Art
[0002] With the rapid development of Internet technology and virtualization technology, cloud computing models based on remote servers have gradually replaced traditional solutions that rely on local hardware resources, showing higher flexibility and economy. Although computing resources are concentrated in the cloud, in many industrial systems, especially smart grids, a large amount of data is still generated by distributed devices such as machines, monitoring equipment, metering instruments and various sensors. Studies have shown that a single surveillance camera can generate hundreds of GB of data per day, while smart sensors covering a large area can generate TB of data in a week. The emergence of this massive distributed data provides new possibilities for the intelligent development of industrial scenarios, but it also brings data processing and transmission problems that need to be solved urgently.
[0003] As a typical industrial scenario, smart grids are equipped with a large number of new terminal devices under the impetus of the intelligent wave, including smart meters, substation monitoring systems, smart sensors, etc. These devices are widely deployed in various links to achieve real-time monitoring of information such as power flow, equipment status and environmental conditions. The terminal devices under the smart grid are no longer limited to simple numerical data, but distributed data with more data types and larger data volumes. Its functions have also expanded from simple data recording to power and load monitoring, equipment and on-site status monitoring, etc. Even terminal devices will have storage requirements with timeliness requirements, which puts higher requirements on storage and processing, making traditional centralized data storage and processing methods face huge challenges.
[0004] In order to cope with the challenge of massive data processing in industrial scenarios, the cloud-edge architecture has gradually become an ideal solution. This architecture deploys computing and storage devices at the edge layer, sinking some data processing tasks from the cloud to edge nodes close to the data source, effectively sharing the computing load of the cloud and reducing dependence on centralized processing. This distributed computing method not only significantly reduces the amount of data transmission and network pressure, but also improves the overall efficiency of the system.
[0005] Because wind resources are usually concentrated in sparsely populated areas with high and stable wind speeds, such as mountainous areas, coastlines or desert areas, and wind power generation equipment requires a large area of land and an open environment to arrange wind turbine arrays and reduce wind interference. These conditions are difficult to meet in cities or densely populated areas, so the site selection of wind power stations is relatively remote. However, the current mature cloud-edge architecture usually deploys edge devices around cities or factories with high demand density. For smart grids, the distributed data generated by terminal devices still faces significant transmission delays and burdens due to physical distance and bandwidth limitations. If we continue to rely on edge devices to transmit data to the cloud for centralized processing, it will not only reduce the economic benefits of the system, but may also further aggravate network congestion and prolong data processing time.
[0006] In industrial systems, in-situ computing is a computing mode in which data is processed locally or nearby. The basic idea is to deploy computing resources near the source of data generation to avoid transmitting large amounts of data to remote data centers for centralized processing. This approach can reduce data transmission delays, reduce bandwidth consumption, and improve the response speed and processing efficiency of the system. In the smart grid scenario, in order to efficiently process the in-situ data generated by terminal devices, an in-situ server system consisting of multiple servers is deployed. The deployment of this system will effectively alleviate the dependence of the smart grid ring on centralized cloud computing capabilities, reduce network bandwidth consumption and latency, and improve data processing efficiency. However, the in-situ server system in the smart grid environment still faces some challenges in practical applications. First, the types and processing requirements of in-situ data are relatively complex, and different data tasks may have different storage requirements; second, the in-situ server system needs to uniformly manage the in-situ data storage tasks generated by all terminal devices, and how to make efficient decisions and processes needs to be solved urgently. Solving this problem is of great significance to promoting the storage management of in-situ data in the smart grid scenario. Summary of the invention
[0007] The purpose of the present invention is to provide an in-situ data storage task management method that can achieve high-speed and low-latency processing in a smart grid environment while taking into account the differentiated demand characteristics of in-situ data storage tasks.
[0008] The technical solution to achieve the purpose of the present invention is: an in-situ data storage management method based on deep reinforcement learning assisted optimization in a smart grid scenario, comprising the following steps:
[0009] Step 1: Establish a regular division mechanism;
[0010] Step 2: Establish a distributed queue architecture;
[0011] Step 3: Propose the optimization problem of in-situ data storage task management;
[0012] Step 4: Propose a storage task management problem based on the Markov decision process. Use the classification deep Q network (C51-ER) algorithm optimized by experience replay to solve the problem. Consider the differentiated requirements of storage tasks, make decisions on the dequeue management of in-situ data storage tasks, and minimize the number of stranded tasks and waiting and completion delays.
[0013] Furthermore, the regular partitioning mechanism described in step 1 is established as follows:
[0014] Suppose an in-place data storage task can be represented as a vector It encapsulates the various requirements of the storage task. Storage requirements i Reflects the requirements of the task on the storage medium performance and describes the read and write speed required for the storage task; update frequency f i Reflects the frequency of task data updates and determines the importance of real-time performance; storage size v i Describe the storage capacity required for the task and provide a basis for resource allocation; security requirements i Indicates the task's requirements for security features such as data encryption and access control; format constraints Describes the task's dependency on a specific data format or protocol. i Indicates the task generation device or data source, used to identify its physical ownership; data type d i Identify the specific form of the task data, such as video, audio or text. The regularity partitioning mechanism can be formalized as a classification function C(T i ). This function combines the core features of the storage task and transforms the storage task vector T through specific rules. i It is divided into three categories: high-frequency update tasks (HUT), type-sensitive tasks (TST) and large-capacity storage tasks (LCT).
[0015] High-frequency update tasks are storage tasks that have a high update frequency and require a fast response. They are widely used in real-time data processing scenarios, such as log data and real-time monitoring data. The response speed of such tasks to the storage medium is s i There are higher requirements, and real-time processing needs must be guaranteed; its update frequency is relatively high. i , data needs to be modified frequently; storage size v i Relatively small or medium. i and format constraints There are no special requirements. α1 to α5 are weight parameters. τ1 represents the update frequency threshold. If the update frequency of the task is f i If it exceeds τ1, it means that the task needs to be updated frequently. The specific discriminant function is as follows:
[0016]
[0017] Type-sensitive tasks have specific requirements for data types and storage media, which are often seen in data that needs to be encrypted or processed in a special format, such as privacy-protected data and sensitive files. i Higher, need to ensure data privacy and security; under format constraints In terms of storage requirements, specific formats or storage media support are required. i Usually higher requirements are required to support special storage operations; and the update frequency f i and storage size v i It changes according to the specific requirements of the task. β1 to β5 are weight parameters. τ2 represents the safety requirement threshold. If the safety requirement of the task is r i If the task exceeds τ2, it means that the task has high security requirements. τ3 represents the format constraint strength threshold. If it exceeds τ3, it means that the task has strict requirements on the data format. The specific discriminant function is as follows:
[0018]
[0019] Large-capacity storage tasks require large storage capacity, which is common in scenarios such as video storage and historical data backup. The storage size of such tasks is v i Large, the amount of task data is significantly higher; the update frequency f i Usually low, data modification frequency is relatively low; response speed to storage media s i The requirements are relatively low, usually focusing on stable storage; at the same time, the security requirements of such tasks are i and format constraints Generally low, the requirements for storage media and format are not high. γ1 to γ5 are weight parameters. τ4 represents the storage capacity threshold. If the storage size v of the task i If it exceeds τ4, it means that the task requires a large storage space. The specific discriminant function is as follows:
[0020]
[0021] By designing a regular division mechanism based on the differentiated demand characteristics of storage tasks, efficient classification of storage tasks can be achieved.
[0022] Furthermore, the distributed queue architecture described in step 2 is established as follows:
[0023] The distributed queue architecture consists of three parts: queue collection {Que t}、Team management rules And the team management rules The specific definitions are as follows:
[0024]
[0025] Queue Set Que t Contains three queues for storage task management. Each queue is dedicated to receiving and processing a specific type of storage task. It contains three queues: Queue for storing HUT HUT , Que for storing LCT LCT And Que for storing TST TST Each queue Que t It can be expressed as:
[0026] Que t ={M t , R t , N t , L t}, t∈{HUT, LCT, TST} (5)
[0027] M t Represents a queue Que t The maximum capacity of the queue is the maximum storage demand that the queue can accommodate. This value is subject to the constraints of physical storage media or resource configuration. t Represents a queue Que t The current remaining storage capacity, which indicates the amount of storage resources in the queue that have not been occupied. t Represents a queue Que t The number of tasks selected for processing, that is, the number of times the queue is selected to perform dequeue processing operations from the beginning of the operation to the current round. t Represents a queue Que t The cumulative delay indicates the total delay of the storage tasks that have been dequeued and processed in the queue so far. t | indicates queue Que t The number of currently stranded tasks in .
[0028] After the regular division mechanism divides the storage tasks into three categories, the queue management rules The corresponding queue will be selected for allocation according to the task type. To avoid queue congestion during the continuous generation and classification of storage tasks, Assign tasks at fixed time intervals and perform queue operations. Before tasks are queued, k task samples are randomly selected for each type of storage task, and the mean value S of the storage requirements of this type of storage task is calculated. t , the specific calculation is as follows:
[0029] n t ~Poisson(λ t ),λ t =φSt (6)
[0030] Among them, φ is the mapping factor, which controls the dynamic adjustment of the number of tasks entering the queue. In each round, the number of tasks entering the queue in different queues follows a Poisson distribution with a mean of S t , the specific calculation is as follows:
[0031]
[0032] Team management rules Specifies the number and order of queue tasks. All queue tasks follow the first-in-first-out (FIFO) strategy. Due to the properties of the media, only a fixed number of storage tasks can perform dequeue processing operations at a time. t The number of queues is δ t where t∈{HUT,LCT,TST}
[0033] When a storage task is enqueued, it is encapsulated as a queue task unit (QTU), which is specifically expressed as follows:
[0034]
[0035] in Represents the original features of storage task i before it is queued. Indicates the number of rounds of storage task i entering the queue, and is assigned to the current round number T when the task is assigned to the queue cut . Indicates the number of rounds of dequeuing task i, initialized to NULL, and assigned to the current round number T when the task is processed and dequeued cur .
[0036] The distributed queue architecture combines the aforementioned regular division mechanism to rationally divide and organize storage tasks into distributed queues, thereby increasing the orderliness and flexibility of task processing.
[0037] Furthermore, the optimization problem of in-situ data storage task management described in step 3 is proposed as follows:
[0038] The management of in-situ data storage tasks focuses on how to coordinate different types of storage tasks to perform queue and dequeue operations. Appropriate queue and dequeue rules will greatly reduce the difficulty of management and speed up the processing rate.
[0039] The queue rule is responsible for assigning the generated storage tasks to the corresponding queues and updating the queue status. At the beginning of each round, the number of tasks n that need to be queued in this round is generated according to the Poisson distribution. t , and calculate the sum of the storage requirements of these tasks Then process each task in turn and encapsulate it into a queue task unit QTU i And try to join the target queue. If the remaining capacity of the target queue is R t Enough to accommodate the current round of n t storage tasks, the queue operation is performed and the number of queue rounds for each storage task is updated at the same time. and the remaining capacity of the queue R t If the target queue capacity is insufficient, then round n t The enqueue request for a storage task was rejected.
[0040] Departure rules Manage the processing flow of queue tasks. All queue tasks follow the first-in-first-out (FIFO) strategy, and can only be taken from the queue at a time. t Take a fixed amount δ t When the dequeue operation is executed, the stored tasks are taken out from the head of the queue in sequence. t tasks, update the number of dequeue rounds for each storage task Queue t Each time a dequeue operation is performed, the corresponding number of selection responses is N t Increase by 1.
[0041] Furthermore, the storage task management problem based on the Markov decision process described in step 4 is proposed. The classification deep Q network (C51-ER) algorithm optimized by experience replay is used to solve the problem. The differentiated requirements of storage tasks are comprehensively considered to make decisions on the dequeue management of in-situ data storage tasks, and the minimization of stranded tasks and waiting and completion delays is pursued. The details are as follows:
[0042] In order to optimize the management decision of distributed task queues, improve the efficiency of storage task processing and reduce system latency, the task management process of the queue is modeled as a Markov model, and K = (Sta, Act, Rew) is used to represent the decision process. Among them, Sta is the state space, which describes the current state of the system; Act is the action space, which defines the optional actions of the system in each state; Rew is the reward function, which is used to measure the pros and cons of the action. The environment of Markov decision is the management system of in-situ data storage tasks in smart grids. In each round, all distributed queues will have storage tasks executing enqueue operations. To ensure the response rate, the system needs to select a queue to perform dequeue processing based on the state characteristics of each queue to optimize the overall storage performance.
[0043] The state is a description of the agent's environment at a specific moment. It abstracts the perceived state information and allows the agent to make decisions. The state space Sta in this study defines the current state of the system in each round, that is, the properties of different queues in the current round. n、AUT n , ACL n They represent the attributes of queue n in round t. The state in each round will record the attributes of the three queues. The state of round t can be specifically expressed as:
[0044] Sta t ={AWL n , AWL n , AWL n}, n∈{HUT, LCT, TST} (9)
[0045] Actions are decisions that the agent can make in the current state. The action selection in each state will not only affect the direction of model training, but also affect the agent's subsequent behavior. The action space Act in this study defines the actions that the system can take in any state, that is, selecting a queue to perform storage task dequeue processing. The action of round t can be specifically expressed as:
[0046] Act t = {Que HUT , Que LCT , Que TST} (10)
[0047] The reward function measures the immediate benefit of the system after selecting an action in the current state. In order to comprehensively optimize task processing efficiency and delay performance, this study defines the reward function Rew t is the weighted penalty of the three features. AWL, AUT and ACL represent the performance indicators of the entire system in round t. w1, w2 and w3 are weight coefficients used to balance the impact of the three features on the reward. The reward in round t can be specifically expressed as:
[0048] Rew t =-(w1·AWL+w2·AUT+w3·ZCL) (11)
[0049] The patent of this invention uses the classification deep Q network (C51-ER) optimized by experience replay to solve the in-situ data storage task management decision in the smart grid environment.
[0050] Compared with the prior art, the present invention has the following significant advantages: (1) A regular division mechanism is designed, which can perform regular classification according to the demand characteristics of the in-situ data storage tasks, providing convenience for subsequent management. (2) A distributed queue architecture is designed, which can record a single type of storage tasks in a targeted manner, set anti-blocking queue entry and exit rules, and ensure the robustness of the processing process. (3) Taking into account the number of stranded tasks and waiting and completion delays in the distributed queue, the classification deep Q network (C51-ER) algorithm optimized by experience replay is used to assist management decisions, and optimize the system's high-speed, low-latency processing of storage tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flowchart of the in-situ data storage task management under the smart grid in the present invention.
[0052] Figure 2 This is the algorithm framework diagram of the classification deep Q network (C51-ER) optimized for experience replay in the present invention. DETAILED DESCRIPTION
[0053] The present invention will be further described in detail below with reference to the accompanying drawings.
[0054] The present invention discloses an in-situ data storage management method based on deep reinforcement learning assisted optimization in a smart grid scenario, comprising the following steps:
[0055] Step 1: Establish a regular division mechanism, as follows:
[0056] Combination Figure 1 , assuming that an in-situ data storage task can be represented as a vector It encapsulates the various requirements of the storage task. Storage requirements i Reflects the requirements of the task on the storage medium performance and describes the read and write speed required for the storage task; update frequency f i Reflects the frequency of task data updates and determines the importance of real-time performance; storage size u i Describe the storage capacity required for the task and provide a basis for resource allocation; security requirements i Indicates the task's requirements for security features such as data encryption and access control; format constraints Describes the task's dependency on a specific data format or protocol. i Indicates the task generation device or data source, used to identify its physical ownership; data type d i Identify the specific form of the task data, such as video, audio or text. The regularity partitioning mechanism can be formalized as a classification function C(T i ). This function combines the core features of the storage task and transforms the storage task vector T through specific rules. iIt is divided into three categories: high-frequency update tasks (HUT), type-sensitive tasks (TST) and large-capacity storage tasks (LCT).
[0057] High-frequency update tasks are storage tasks that have a high update frequency and require a fast response. They are widely used in real-time data processing scenarios, such as log data and real-time monitoring data. The response speed of such tasks to the storage medium is s i There are higher requirements, and real-time processing needs must be guaranteed; its update frequency is relatively high. i , data needs to be modified frequently; storage size v i Relatively small or medium. i and format constraints There are no special requirements. α1 to α5 are weight parameters. τ1 represents the update frequency threshold. If the update frequency of the task is f i If it exceeds τ1, it means that the task needs to be updated frequently. The specific discriminant function is as follows:
[0058]
[0059] Type-sensitive tasks have specific requirements for data types and storage media, which are often seen in data that needs to be encrypted or processed in a special format, such as privacy-protected data and sensitive files. i Higher, need to ensure data privacy and security; under format constraints In terms of storage requirements, specific formats or storage media support are required. i Usually higher requirements are required to support special storage operations; and the update frequency f i and storage size v i It changes according to the specific requirements of the task. β1 to β5 are weight parameters. τ2 represents the safety requirement threshold. If the safety requirement of the task is r i If the task exceeds τ2, it means that the task has high security requirements. τ3 represents the format constraint strength threshold. If it exceeds τ3, it means that the task has strict requirements on the data format. The specific discriminant function is as follows:
[0060]
[0061] Large-capacity storage tasks require large storage capacity, which is common in scenarios such as video storage and historical data backup. The storage size of such tasks is v i Large, the amount of task data is significantly higher; the update frequency f i Usually low, data modification frequency is relatively low; response speed to storage media s i The requirements are relatively low, usually focusing on stable storage; at the same time, the security requirements of such tasks arei and format constraints Generally low, the requirements for storage media and format are not high. γ1 to γ5 are weight parameters. τ4 represents the storage capacity threshold. If the storage size v of the task i If it exceeds τ4, it means that the task requires a large storage space. The specific discriminant function is as follows:
[0062]
[0063] By designing a regular division mechanism based on the differentiated demand characteristics of storage tasks, efficient classification of storage tasks can be achieved.
[0064] Step 2: Establish a distributed queue architecture, as follows:
[0065] Combination Figure 1 ,The distributed queue architecture consists of three parts: queue set {Que t}、Team management rules And the team management rules The specific definitions are as follows:
[0066]
[0067] Queue Set Que t Contains three queues for storage task management. Each queue is dedicated to receiving and processing a specific type of storage task. It contains three queues: Queue for storing HUT HUT , Que for storing LCT LCT And Que for storing TST TST Each queue Que t It can be expressed as:
[0068] Que t ={M t , R t , N t , L t}, t∈{HUT, LCT, TST} (16)
[0069] M t Represents a queue Que t The maximum capacity of the queue is the maximum storage demand that the queue can accommodate. This value is subject to the constraints of physical storage media or resource configuration. t Represents a queue Que t The current remaining storage capacity, which indicates the amount of storage resources in the queue that have not been occupied. t Represents a queue Que t The number of tasks selected for processing, that is, the number of times the queue is selected to perform dequeue processing operations from the beginning of the operation to the current round.t Represents a queue Que t The cumulative delay indicates the total delay of the storage tasks that have been dequeued and processed in the queue so far. t | indicates queue Que t The number of currently stranded tasks in .
[0070] After the regular division mechanism divides the storage tasks into three categories, the queue management rules The corresponding queue will be selected for allocation according to the task type. To avoid queue congestion during the continuous generation and classification of storage tasks, Tasks are assigned and queued at fixed time intervals. Before tasks are queued, k task samples are randomly selected for each type of storage task, and the mean value St of the storage requirements of this type of storage task is calculated. The specific calculation is as follows:
[0071] n t ~Poisson(λ t ), λ t =φS t (17)
[0072] Among them, φ is the mapping factor, which controls the dynamic adjustment of the number of tasks entering the queue. In each round, the number of tasks entering the queue in different queues follows a Poisson distribution with a mean of S t , the specific calculation is as follows:
[0073]
[0074] Team management rules Specifies the number and order of queue tasks. All queue tasks follow the first-in-first-out (FIFO) strategy. Due to the properties of the media, only a fixed number of storage tasks can perform dequeue processing operations at a time. t The number of queues is δ t where t∈{HUT,LCT,TST}
[0075] When a storage task is enqueued, it is encapsulated as a queue task unit (QTU), which is specifically expressed as follows:
[0076]
[0077] in Represents the original features of storage task i before it is queued. Indicates the number of rounds of storage task i entering the queue, and is assigned to the current round number T when the task is assigned to the queue cur . Indicates the number of rounds of dequeuing task i, initialized to NULL, and assigned to the current round number T when the task is processed and dequeuedcur .
[0078] The distributed queue architecture combines the aforementioned regular division mechanism to rationally divide and organize storage tasks into distributed queues, thereby increasing the orderliness and flexibility of task processing.
[0079] Step 3: Propose the optimization problem of in-situ data storage task management, as follows:
[0080] The management of in-situ data storage tasks focuses on how to coordinate different types of storage tasks to perform queue and dequeue operations. Appropriate queue and dequeue rules will greatly reduce the difficulty of management and speed up the processing rate.
[0081] The queue rule is responsible for assigning the generated storage tasks to the corresponding queues and updating the queue status. At the beginning of each round, the number of tasks n that need to be queued in this round is generated according to the Poisson distribution. t , and calculate the sum of the storage requirements of these tasks Then process each task in turn and encapsulate it into a queue task unit QTU i And try to join the target queue. If the remaining capacity of the target queue is R t Enough to accommodate the current round of n t storage tasks, the queue operation is performed and the number of queue rounds for each storage task is updated at the same time. and the remaining capacity of the queue R t If the target queue capacity is insufficient, then round n t The enqueue request for a storage task was rejected.
[0082] Departure rules Manage the processing flow of queue tasks. All queue tasks follow the first-in-first-out (FIFO) strategy, and can only be taken from the queue at a time. t Take a fixed amount δ t When the dequeue operation is executed, the stored tasks are taken out from the head of the queue in sequence. t tasks, update the number of dequeue rounds for each storage task Queue t Each time a dequeue operation is performed, the corresponding number of selection responses is N t Increase by 1.
[0083] Step 4: Propose a storage task management problem based on the Markov decision process, use the classification deep Q network (C51-ER) algorithm optimized by experience replay to solve the problem, comprehensively consider the different needs of storage tasks, make decisions on the dequeue management of in-situ data storage tasks, and pursue the minimization of stranded tasks and waiting and completion delays. The details are as follows:
[0084] Combination Figure 2 In order to optimize the management decision of distributed task queues, improve the efficiency of storage task processing and reduce system latency, the task management process of the queue is modeled as a Markov model, and K = (Sta, Act, Rew) is used to represent the decision process. Among them, Sta is the state space, which describes the current state of the system; Act is the action space, which defines the optional actions of the system in each state; Rew is the reward function, which is used to measure the pros and cons of the action. The environment of Markov decision is the management system of in-situ data storage tasks in smart grids. In each round, all distributed queues will have storage tasks executing enqueue operations. To ensure the response rate, the system needs to select a queue to perform dequeue processing based on the state characteristics of each queue to optimize the overall storage performance.
[0085] The state is a description of the agent's environment at a specific moment. It abstracts the perceived state information and allows the agent to make decisions. The state space Sta in this study defines the current state of the system in each round, that is, the properties of different queues in the current round. n 、AUT n , ACL n They represent the attributes of queue n in round t. The state in each round will record the attributes of the three queues. The state of round t can be specifically expressed as:
[0086] Sta t ={AWL n , AWL n , AWL n}, n∈{HUT, LCT, TST} (20)
[0087] Actions are decisions that the agent can make in the current state. The action selection in each state will not only affect the direction of model training, but also affect the agent's subsequent behavior. The action space Act in this study defines the actions that the system can take in any state, that is, selecting a queue to perform storage task dequeue processing. The action of round t can be specifically expressed as:
[0088] Act t = {Que HUT , Que LCT , Que TST} (twenty one)
[0089] The reward function measures the immediate benefit of the system after selecting an action in the current state. In order to comprehensively optimize task processing efficiency and delay performance, this study defines the reward function Rew tis the weighted penalty of the three features. AWL, AUT and ACL represent the performance indicators of the entire system in round t. w1, w2 and w3 are weight coefficients used to balance the impact of the three features on the reward. The reward in round t can be specifically expressed as:
[0090] Rew t =-(w1·AWL+w2·AUT+w3·ACL) (22)
[0091] The present invention uses the classification deep Q network (C51-ER) optimized by experience replay to solve the in-situ data storage task management decision in the smart grid environment. The algorithm pseudo code can be defined as:
[0092]
[0093]
[0094] The above content describes the implementation process and advantages of the present invention. Those skilled in the art should understand that the present invention may have various changes and improvements without departing from the principles of the present invention, and these changes and improvements fall within the scope of the present invention claimed for protection.
Claims
1. An in-situ data storage management method based on deep reinforcement learning assisted optimization in a smart grid scenario, characterized in that: The following steps are involved: Step 1: Establish a regular division mechanism; Step 2: Establish a distributed queue architecture; Step 3: Propose the optimization problem of in-situ data storage task management; Step 4: Propose a storage task management problem based on the Markov decision process. Use the classification deep Q network (C51-ER) algorithm optimized by experience replay to solve the problem. Consider the differentiated requirements of storage tasks, make decisions on the dequeue management of in-situ data storage tasks, and minimize the number of stranded tasks and waiting and completion delays.
2. According to the in-situ data storage management method based on deep reinforcement learning assisted optimization in the smart grid scenario described in claim 1, it is characterized in that: The mechanism for establishing a regular partitioning as described in step 1 is as follows: Suppose an in-place data storage task can be represented as a vector It encapsulates the various requirements of the storage task. Storage requirements i Reflects the requirements of the task on the storage medium performance and describes the read and write speed required for the storage task; update frequency f i Reflects the frequency of task data updates and determines the importance of real-time performance; storage size v i Describe the storage capacity required for the task and provide a basis for resource allocation; security requirements i Indicates the task's requirements for security features such as data encryption and access control; format constraints Describes the task's dependency on a specific data format or protocol. i Indicates the task generation device or data source, used to identify its physical ownership; data type d i Identify the specific form of the task data, such as video, audio or text. The regularity partitioning mechanism can be formalized as a classification function C(T i ). This function combines the core features of the storage task and transforms the storage task vector T through specific rules. i It is divided into three categories: high-frequency update tasks (HUT), type-sensitive tasks (TST) and large-capacity storage tasks (LCT). High-frequency update tasks are storage tasks that have a high update frequency and require a fast response. They are widely used in real-time data processing scenarios, such as log data and real-time monitoring data. The response speed of such tasks to the storage medium is s i There are higher requirements and real-time processing needs must be guaranteed; its update frequency is relatively high. i , data needs to be modified frequently; storage size v i Relatively small or medium. i and format constraints There are no special requirements. α1 to α5 are weight parameters. τ1 represents the update frequency threshold. If the update frequency of the task is f i If it exceeds τ1, it means that the task needs to be updated frequently. The specific discriminant function is as follows: Type-sensitive tasks have specific requirements for data types and storage media, which are often seen in data that needs to be encrypted or processed in a special format, such as privacy-protected data and sensitive files. i Higher, need to ensure data privacy and security; under format constraints In terms of storage requirements, specific formats or storage media support are required. i Usually higher requirements are required to support special storage operations; and the update frequency f i and storage size v i It changes according to the specific requirements of the task. β1 to β5 are weight parameters. τ2 represents the safety requirement threshold. If the safety requirement of the task is r i If the task exceeds τ2, it means that the task has high security requirements. τ3 represents the format constraint strength threshold. If it exceeds τ3, it means that the task has strict requirements on the data format. The specific discriminant function is as follows: Large-capacity storage tasks require large storage capacity, which is common in scenarios such as video storage and historical data backup. The storage size of such tasks is v i Large, the amount of task data is significantly higher; the update frequency f i Usually low, data modification frequency is relatively low; response speed to storage media s i The requirements are relatively low, usually focusing on stable storage; at the same time, the security requirements of such tasks are i and format constraints Generally low, the requirements for storage media and format are not high. γ1 to γ5 are weight parameters. τ4 represents the storage capacity threshold. If the storage size v of the task i If it exceeds τ4, it means that the task requires a large storage space. The specific discriminant function is as follows: By designing a regular division mechanism based on the differentiated demand characteristics of storage tasks, efficient classification of storage tasks can be achieved.
3. According to the in-situ data storage management method based on deep reinforcement learning assisted optimization in the smart grid scenario described in claim 1, it is characterized in that: Step 2 describes the establishment of a distributed queue architecture, as follows: The distributed queue architecture consists of three parts: queue collection {Que t }、Team management rules And the team management rules The specific definitions are as follows: Queue Set Que t Contains three queues for storage task management. Each queue is dedicated to receiving and processing a specific type of storage task. It contains three queues: Queue for storing HUT HUT , Que for storing LCT LCT And Que for storing TST TST Each queue Que t It can be expressed as: That t ={M t ,R t ,N t ,L t },t∈{HUT,LCT,TST} (5) M t Represents a queue Que t The maximum capacity of the queue is the maximum storage demand that the queue can accommodate. This feature is constrained by the physical storage medium or resource configuration. t Represents a queue Que t The current remaining storage capacity, which indicates the amount of storage resources in the queue that have not been occupied. t Represents a queue Que t The number of tasks selected for processing, that is, the number of times the queue is selected to perform dequeue processing operations from the beginning of the operation to the current round. t Represents a queue Que t The cumulative delay indicates the total delay of the storage tasks that have been dequeued and processed in the queue so far. t | indicates queue Que t The number of currently stranded tasks in . After the regular division mechanism divides the storage tasks into three categories, the queue management rules The corresponding queue will be selected for allocation according to the task type. To avoid queue congestion during the continuous generation and classification of storage tasks, Assign tasks at fixed time intervals and perform queue operations. Before tasks are queued, k task samples are randomly selected for each type of storage task, and the mean value S of the storage requirements of this type of storage task is calculated. t , the specific calculation is as follows: n t ~Poisson(λ t ), λ t =φS t (6) Where φ is the mapping factor, which controls the dynamic adjustment of the number of tasks entering the queue. In each round, the number of tasks entering the queue in different queues follows a Poisson distribution with a mean of S t , the specific calculation is as follows: Team management rules Specifies the number and order of queue tasks. All queue tasks follow the first-in-first-out (FIFO) strategy. Due to the properties of the media, only a fixed number of storage tasks can perform dequeue processing operations at a time. t The number of queues is δ t where t∈{HUT,LCT,TST} When a storage task is enqueued, it is encapsulated as a queue task unit (QTU), which is specifically expressed as follows: in Represents the original features of the storage task o before it is queued. Indicates the number of rounds of storage task o entering the queue, and is assigned to the current round number T when the task is assigned to the queue cur . Indicates the number of rounds of dequeuing task i, initialized to NULL, and assigned to the current round number T when the task is processed and dequeued cur . The distributed queue architecture combines the aforementioned regular division mechanism to rationally divide and organize storage tasks into distributed queues, thereby increasing the orderliness and flexibility of task processing.
4. The in-situ data storage management method based on deep reinforcement learning assisted optimization in the smart grid scenario described in claim 1 is characterized in that: The optimization problem of in-situ data storage task management described in step 3 is as follows: The management of in-situ data storage tasks focuses on how to coordinate different types of storage tasks to perform queue and dequeue operations. Appropriate queue and dequeue rules will greatly reduce the difficulty of management and speed up the processing rate. The enqueue rule is responsible for assigning the generated storage tasks to the corresponding queues and updating the status of the queues. At the beginning of each round, the number of tasks n that need to be queued in this round is generated according to the Poisson distribution t , and calculate the sum of the storage requirements of these tasks Then process each task in turn and encapsulate it into a queue task unit QTU i And try to join the target queue. If the remaining capacity of the target queue is R t Enough to accommodate the current round of n t storage tasks, the queue operation is performed and the number of queue rounds for each storage task is updated at the same time. and the remaining capacity of the queue R t If the target queue capacity is insufficient, then round n t The enqueue request for a storage task was rejected. Departure rules Manage the processing flow of queue tasks. All queue tasks follow the first-in-first-out (FIFO) strategy, and can only be taken from the queue at a time. t Take a fixed amount δ t When the dequeue operation is executed, the stored tasks are taken out from the head of the queue in sequence. t tasks, update the number of dequeue rounds for each storage task Queue t Each time a dequeue operation is performed, the corresponding number of selection responses is N t Increase by 1.
5. The in-situ data storage management method based on deep reinforcement learning assisted optimization in the smart grid scenario described in claim 1 is characterized in that: Step 4 describes the storage task management problem based on the Markov decision process. The classification deep Q network (C51-ER) algorithm optimized by experience replay is used to solve the problem. The differentiated requirements of storage tasks are comprehensively considered to make decisions on the dequeue management of in-situ data storage tasks, and the minimization of stranded tasks and waiting and completion delays is pursued. The details are as follows: In order to optimize the management decision of distributed task queues, improve the efficiency of storage task processing and reduce system latency, the task management process of the queue is modeled as a Markov model, and K = (Sta, Act, Rew) is used to represent the decision process. Among them, Sta is the state space, which describes the current state of the system; Act is the action space, which defines the optional actions of the system in each state; Rew is the reward function, which is used to measure the pros and cons of the action. The environment of Markov decision is the management system of in-situ data storage tasks in smart grids. In each round, all distributed queues will have storage tasks executing enqueue operations. To ensure the response rate, the system needs to select a queue to perform dequeue processing based on the state characteristics of each queue to optimize the overall storage performance. The state is a description of the agent's environment at a specific moment. It abstracts the perceived state information and allows the agent to make decisions. The state space Sta in this study defines the current state of the system in each round, that is, the properties of different queues in the current round. n 、AUT n , ACL n They represent the attributes of queue n in round t. The state in each round will record the attributes of the three queues. The state of round t can be specifically expressed as: Sta t ={AWL n ,AWL n ,AWL n },n∈{HUT,LCT,TST} (9) Actions are decisions that the agent can make in the current state. The action selection in each state will not only affect the direction of model training, but also affect the agent's subsequent behavior. The action space Act in this study defines the actions that the system can take in any state, that is, selecting a queue to perform storage task dequeue processing. The action of round t can be specifically expressed as: Ast t ={That HUT ,That LCT ,That TST } (10) The reward function measures the immediate benefit of the system after selecting an action in the current state. In order to comprehensively optimize task processing efficiency and delay performance, this study defines the reward function Rew t is the weighted penalty of the three features. AWL, AUT and ACL represent the performance indicators of the entire system in round t. w1, w2 and w3 are weight coefficients used to balance the impact of the three features on the reward. The reward in round t can be specifically expressed as: Ice t =-(w1·AWL+w2·AUT+w3·ACL) (11) The present invention uses the classification deep Q network (C51-ER) optimized by experience replay to solve the in-situ data storage task management decision in the smart grid environment. The algorithm pseudo code can be defined as: