In-situ data storage management method combining reinforcement learning and load balancing under scene of Internet of Things

By combining reinforcement learning and load balancing technologies in the Internet of Things (IoT) environment, setting up hierarchical data queues and utilizing the Dueling DQN algorithm, the in-situ data storage management is dynamically optimized, solving the problems of server performance differences and uneven load in the IoT environment, and improving data processing efficiency and resource utilization.

CN120849110APending Publication Date: 2025-10-28CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957328.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In the Internet of Things (IoT) environment, in-situ computing faces challenges such as server performance differences, limited computing resources, uneven task allocation, and load fluctuations, resulting in low data processing efficiency. Existing static rules cannot effectively cope with sudden load changes.

Method used

By combining reinforcement learning and load balancing techniques, and by setting up hierarchical data queues and abstract storage units, the Dueling DQN algorithm is used to optimize in-situ data storage management, dynamically adjust data priority classification strategies, reduce response latency, and avoid storage task blocking.

Benefits of technology

It improves the efficiency of data access and scheduling and the utilization of system resources in the Internet of Things environment, optimizes server load balancing, and reduces data processing latency and task response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure QLYQS_5
    Figure QLYQS_5
Patent Text Reader

Abstract

The invention discloses an in-situ data storage management method in combination with reinforcement learning and load balancing in an Internet of Things scene. The method comprises the following steps of: abstracting servers with different performances into hierarchical data queues, and specifying properties of storage units; then setting a buffer area for receiving an in-situ data storage task, packaging the in-situ data storage task into a storage unit, and setting a dynamic classification strategy by considering the own priority and actual fluctuation of the storage unit and the load condition of a server queue; then proposing an optimization problem of in-situ data storage management; and finally, describing the problem by using a Markov decision process model, and solving the storage management decision of the in-situ data in combination with reinforcement learning and a load balancing mechanism. According to the method, under the environment of the Internet of Things, the performance difference and the load state of each level of server queue are comprehensively considered, a management decision is made by taking reduction of response delay and avoiding of storage task blockage as targets, the resource utilization rate and load balance of each type of server are guaranteed, and the data processing efficiency of a system in a hierarchical data queue structure is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of edge computing, specifically involving the optimization of in-situ data storage management in IoT scenarios by combining reinforcement learning and load balancing technologies. Background Technology

[0002] With the rapid development of IoT technology, various smart terminals and sensors are widely used in industries such as manufacturing, agriculture, transportation, and healthcare, generating massive amounts of real-time data. This data is diverse and widely distributed. Various devices (such as smart sensors, monitoring equipment, and automated machinery) collect real-time operational status, environmental monitoring data, behavioral data, and predictive data, while also generating a large volume of data streams. This data includes not only high-frequency data with extremely high latency requirements but also low-frequency but important historical data. With the widespread adoption of IoT devices, the scale and variety of data are exploding.

[0003] Faced with such massive and diverse data, traditional cloud-based processing architectures are under immense pressure in terms of storage and transmission. While cloud computing boasts powerful computing capabilities and elastic scalability, it still presents a series of challenges for data processing requiring low latency and high real-time performance, including insufficient bandwidth, excessive latency, and privacy and security concerns.

[0004] To address the aforementioned issues, edge computing has emerged as a widely adopted technology. Edge computing migrates data storage and processing capabilities to the network edge, deploying computing resources closer to the data source. This reduces data transmission distance, latency, and bandwidth pressure, making it particularly suitable for scenarios with high real-time requirements, such as the Internet of Things (IoT). By bringing computing resources closer to end devices, edge computing effectively solves the real-time processing problem of high-frequency data generated by IoT devices.

[0005] However, with the increasing prevalence of edge computing, the pure edge computing model still faces some bottlenecks, especially in IoT environments where devices are widely deployed and computing resources are limited. Some terminal devices are often located in remote areas, such as mines and oil fields, where infrastructure is relatively weak and network connectivity is unstable. Although cloud-edge collaborative architectures alleviate transmission latency issues to some extent, they still cannot avoid the physical distance and bandwidth limitations between remote devices and edge servers. Over-reliance on centralized processing solutions may affect the system's ability to process real-time data and could even lead to network congestion.

[0006] To further improve the real-time performance and efficiency of data processing, the concept of in-situ computing has emerged and has been widely applied in the Internet of Things (IoT) environment. In-situ computing is a model that deploys computing resources directly near the data source, enabling localized data processing. In in-situ computing, data processing no longer relies on remote data centers but is handled directly by servers near the devices. In this way, in-situ computing reduces data transmission latency and improves real-time performance.

[0007] In IoT environments, in-situ computing typically involves deploying in-situ servers with varying performance levels. These servers provide data storage and computing capabilities, and schedule in-situ data access. While this approach effectively reduces data transmission latency, it still faces several challenges in practical applications: First, the performance differences among in-situ servers and their limited computing resources lead to uneven task allocation and potential overload issues. Second, existing data often relies on static rules or simple priorities, making it unable to cope with sudden load fluctuations and impacting data processing efficiency. Finally, the response of storage tasks across different servers requires coordinated management and optimization, resulting in resource contention and response latency problems. Solving these problems is of great significance for optimizing load-balanced in-situ data storage management in IoT environments. Summary of the Invention

[0008] This invention aims to optimize in-situ data storage management in the Internet of Things (IoT) environment by combining reinforcement learning and load balancing technologies, and to provide an optimization scheme that can improve data access scheduling efficiency and system resource utilization.

[0009] The technical solution for achieving the purpose of this invention is: an in-situ data storage management method combining reinforcement learning and load balancing in an Internet of Things (IoT) scenario, characterized by the following steps:

[0010] Step 1: Set up hierarchical data queues and abstract storage units

[0011] Step 2: Develop a load balancing data priority classification strategy.

[0012] Step 3: Propose optimization issues for in-situ data storage management and server load balancing.

[0013] Step 4: Propose a Markov decision process model to describe the in-situ data storage management problem. Based on this model, use the Dueling DQN algorithm to comprehensively consider the load differences of hierarchical data queues and dynamically optimize the processing response decision of the in-situ data storage unit to reduce response latency and avoid blocking of storage tasks.

[0014] Furthermore, the setting up of the hierarchical data queue and abstracting the storage unit described in step 1 is as follows:

[0015] In an in-situ server system within an IoT environment, servers are deployed near distributed terminals to receive and process in-situ data. By combining servers with varying performance and functionalities, and configuring them differently in terms of data processing capabilities and storage response speeds, the diverse data storage needs of the Industrial IoT can be met. Furthermore, these servers with different capabilities can be abstracted into corresponding hierarchical data queues: Q HPD Used to store high-priority data (HPD), this queue is abstracted from the highest-performing server and is suitable for storing critical data with high real-time requirements and high access frequency. Q MPD Used for storing Medium Priority Data (MPD) type data, this queue consists of servers with moderate performance, balancing access efficiency and storage cost, and is suitable for storing periodic task data. Q LPD Used for storing low-priority data (LPD), this queue is abstracted from cost-optimized servers and is suitable for data with low access frequency and long lifecycles, emphasizing storage economy. The hierarchical data queue (HDQ) can be specifically represented as:

[0016] HDQ={Q HPD Q MPD Q LPD} (1)

[0017] Each hierarchical data queue includes the following characteristics: Queue Q represents i Maximum data capacity Queue Q represents i The number of storage units processed. Indicates the latency of processed memory units. Res i Queue Q represents i When selected for a single scheduler, the maximum amount of computational resources available per round is used to constrain dequeue operations: M represents the CPU computing power that can be allocated per round. i B represents the available memory capacity per round. i Network bandwidth available per round.

[0018]

[0019]

[0020] To receive in-situ data generated by all distributed terminals, a unified data buffer is set up, encapsulating the main characteristics of each in-situ data in vector form: F represents the frequency of data access or reading, measuring the number of times the data is accessed within a certain period. C represents the storage space occupied by the data, usually in bytes. T represents the duration the data needs to be stored, usually measured in time. P represents the initial priority of the data in the system. These four characteristics are the key basis for implementing data classification and scheduling.

[0021] D = {F, C, T, P} (4)

[0022] Based on the characteristics of the data encapsulated in the data buffer, a scheduling classification strategy is designed to classify all preprocessed in-situ data into three categories: HPD, MPD, and LPD. Upon receiving in-situ data of the corresponding category, different storage queues will encapsulate it into a storage unit (SU), defined in detail below:

[0023] SU = {D, T} in T out ,κ},κ={k CPU κ mem , κ bw} (5)

[0024] Where T in Used to record the enqueue round number of a storage unit, initialized to the current round number T when the storage unit was added to the queue. current T out This is used to record the dequeue round number of a storage unit. It is initialized to null, indicating that the storage unit has not yet been processed or removed from the queue. When the storage unit is dequeued, it is updated to the current round number T at the time of dequeue. current κ represents the resources required to complete the processing of each storage unit in this round: κ CPU Indicates the required CPU cycles; κ mem Indicates the required memory; κ bw This indicates the required bandwidth.

[0025] Furthermore, the data priority classification strategy for load balancing described in step 2 is as follows:

[0026] To adaptively store storage units of different priorities into appropriate hierarchical data queues, a load-balancing data priority classification strategy is introduced. Taking into account the static characteristics of the storage units themselves, a static priority score G is calculated:

[0027] G = w F ·F+w C ·C+w T ·T+(1+μ)·P (6)

[0028] Among them, w F w S and w T These represent the priority weights of each static feature. To describe the suddenness and unpredictability of in-situ data priority in the IoT environment, a coefficient of variation μ is introduced to describe the probabilistic abrupt changes in data priority. The coefficient of variation μ follows a Bernoulli distribution:

[0029]

[0030] Where ∈ is a small constant representing the probability of triggering a priority mutation; μhig h This is a relatively large constant representing the increase in priority during sudden changes. Data generally maintains its own priority, but there is a certain probability that it will suddenly change to extremely high priority.

[0031] Design a static classification function Categorystati based on the static priority score G. c (D) Perform preliminary classification and discrimination on the in-situ data, where G high G low The score threshold used for classifying categories:

[0032]

[0033] Each in-situ data point will be initially divided into three categories—HPD, MPD, and LPD—based on a static priority G score. The number of data points in each category after classification is as follows: and At this point, the load on each queue after static priority classification is:

[0034]

[0035] To further optimize the classification results based on static priority, a dynamic priority adjustment mechanism with adaptive threshold adjustment is introduced. This mechanism dynamically adjusts the classification threshold by sensing the load of each queue in the hierarchical data queue architecture.

[0036]

[0037] Where λ is used as an adjustment factor to control the threshold update rate, and the updated threshold G′ is used to control the threshold update rate. high and G′ low Considering the load on each queue in a hierarchical data queue architecture, the data is reclassified:

[0038]

[0039] The specific quantities for each category after classification are as follows: and At this point, the result of dynamic classification makes the load ratio of each queue close to its computing power ratio, satisfying the following constraint:

[0040]

[0041] To standardize the weights of different features in the static priority score, each time a storage unit is classified, log data is stored, recording all data features and the dynamic priority classification threshold used for classification. Correlation analysis is then used to calibrate the weights in the priority score G using the log data. For any feature X (the static features F, C, T, and P of the storage unit), its correlation with the dynamic classification threshold G′ can be calculated. high and G′ low Correlation coefficient:

[0042]

[0043] in Used to calculate the expectation. To ensure the weights sum to 1, a normalization method is used to calculate the weights based on two thresholds:

[0044]

[0045] Each feature yields two normalized weight values ​​based on a classification threshold. To better calibrate the weights in the static priority score G, a balancing coefficient θ is introduced to control the bias of the final weights. The final weight calculation formula is as follows:

[0046]

[0047] Furthermore, step 3 describes the optimization problem of in-situ data storage management and server load balancing, as follows:

[0048] In IoT scenarios, optimizing in-situ data storage management and server load balancing primarily addresses how to load-balance storage units with different characteristics to appropriate storage queues and implement strategic management responses. Reinforcement learning algorithm-driven management strategies significantly impact the effectiveness of system data processing.

[0049] In the in-situ data storage management process within an IoT environment, the data generation rate is typically affected by the system's current load and operating status. To ensure the stability of the hierarchical data queue structure during management, a Markov-modulated Poisson process is introduced. v The Modified Poisson Process (MMPP) is used to determine the total amount of in-situ data generated in each round. At round t, the load X of each queue needs to be calculated first. i (t) Case:

[0050]

[0051] To model using MMPP, the load value X needs to be... i The load (t) is discretized into a finite number of state categories. The load value is transformed into three levels: high (H), medium (M), and low (L) using a mapping function g(·). Furthermore, the overall load state of the system at round t consists of the states of all queues.

[0052] state i (t)=g(X i (t)), state i (t)∈{H,M,L} (19)

[0053] state(t) = (state HPD (t), state MPD (t), state LOD (t)) (20)

[0054] Historical data from the in-situ server system is time-series, recording the load of each queue across multiple scheduling rounds, and the actual amount of in-situ data N(t) generated in round t. All rounds T with the same state stat(t) are then processed. state(t) Furthermore, by calculating the average data generation rate under the same state stat(t), the Poisson process parameter corresponding to that state, i.e., the data generation rate, is estimated.

[0055] T state(t) =t′|state(t′)=state(t) (21)

[0056]

[0057] The above method can establish a mapping relationship f between load state and generation rate, which can then be used for subsequent simulations to generate specific amounts of data.

[0058] λ(state(t))=f(state HPD (t), state MPD (t), state LPD (t)) (23)

[0059] Considering the suddenness and unpredictability of in-situ data, the amount of data N(t) generated in round t can be obtained using MMPP modeling. The mutation factor M(t) is used to increase the suddenness of data generation, and will mutate to a larger constant value M with a relatively low probability p. high This causes the number of units generated in that round to suddenly reach a higher peak.

[0060]

[0061] N(t)=Poisson(M(t)·λ(state(t))) (25)

[0062] The changes in the system state follow the transition rules of a Markov chain, meaning the state in the next round depends only on the state in the current round and is independent of earlier historical states. Through discretization of the historical data load, the finite state space S of the system can be defined as:

[0063] s={state1, state2,..., state |S|} (26)

[0064] We can statistically analyze the transitions of each state from historical data and construct a state transition counting matrix M. The elements of M are... j Indicates the state i Transition to state j The number of times. 1(·) is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise.

[0065]

[0066] The state transition probability matrix P is obtained by normalizing the state transition count matrix. i,j This indicates the current state of the system. i In the next round, the state transitions to that state. j The probability of the state transition is given. Due to the properties of Markov processes, the state transition probability matrix P satisfies the constraint that the sum of the probabilities is 1:

[0067]

[0068] Using the state transition matrix P, the states of future rounds can be deduced sequentially from the known initial state state(t). The system's state state(t) in round t determines the state state(t+1) in round t+1, unaffected by earlier states. Based on the deduced state state(t+1), the amount of data generated in round t+1 can be further determined.

[0069] state(t+1)~P(state(t)) (30)

[0070] N(t+1)~Poisson(λ(state(t+1)))(31)

[0071] After the MMPP model generates the total amount of data in round t, the data buffer and classification strategy will classify the data into three specific categories. The number of data entries in each queue corresponds to the number of data entries after classification. and In each round, each queue will perform an enqueue operation based on the quantity of data categorized by priority. First, calculate... The system calculates the storage requirements of each piece of data and compares them with the remaining available capacity of the queue. If the storage requirements are less than the remaining capacity of the queue, the enqueue operation is discarded and the data is returned to the data buffer; otherwise, the enqueue operation is performed. Each data point is added to the queue, and the corresponding enqueue round number is updated.

[0072] In a hierarchical data queue architecture, if data units from multiple server queues are processed simultaneously, the actual resources available to each queue will be significantly lower than the theoretical limit due to system-level contention. Therefore, the system comprehensively considers the states of different queues across the entire system and selects only one queue for optimal scheduling in each round. The data buffer and classification algorithm divide the in-situ data generated by IoT terminals into three categories, which are then stored in storage queues abstracted from servers of different performance levels. When a single queue is selected for response in each round, the available resources are Res. i The number of dequeues from a specific storage unit, ξ i This can be expressed as:

[0073]

[0074] The dequeue method of the storage queue follows the First-In-First-Out (FIFO) principle, meaning that storage units in the queue are processed sequentially according to the enqueue order. In each round, the system selects a queue Q. i Perform scheduling processing and retrieve the corresponding number of storage units ξ from the head of the queue. i After dequeueing is complete, the queue's characteristics need to be updated, including the cumulative number of scheduled processes. and total delay in completion At the same time, update the number of rounds for leaving the team. out .

[0075] Furthermore, step 4 describes the proposal of a Markov decision process model to describe the in-situ data storage management problem. Based on this model, the Dueling DQN algorithm is used to comprehensively consider the load differences of hierarchical data queues and dynamically optimize the processing response decisions of in-situ data storage units to reduce response latency and avoid blocking of storage tasks. Specifically:

[0076] In the process of in-situ data storage management in IoT scenarios, the amount of in-situ data in the hierarchical data queue structure changes dynamically according to the enqueue and dequeue strategies. Furthermore, each round requires selecting a storage queue for response. Due to the dynamic nature and decision complexity of this problem, reinforcement learning methods are well-suited for optimal decision-making based on the system state. Therefore, this problem needs to be abstracted into a Markov model:

[0077]

[0078] To evaluate the response performance of in-situ data storage units driven by different algorithms, three metrics were set: Average Waiting Rounds (AWR), which calculates the average number of waiting rounds for unprocessed data in all queues, directly reflects the timeliness of the system in data scheduling and processing; Average Unprocessed Data (AUD), which calculates the average amount of unprocessed data in all queues, reflects the adaptability of the system in terms of task processing capacity; and Average Processing Latency (APL), which calculates the average time taken for all scheduled data, reflects the efficiency of the system in data scheduling and processing. These are expressed as follows:

[0079]

[0080] Because these indicators have different dimensions, to avoid any one indicator having too large an impact on the reward function, they will be normalized first.

[0081]

[0082] In addition, the upper bound of the delay for all queues in each round needs to be calculated. Up to round T... current Queue Q i Minimum number of teams dispatched per round in historical data and the global expected size of the storage unit It can be represented as:

[0083]

[0084] For any queue Q i The maximum latency a storage unit may experience It can be represented as:

[0085]

[0086] Based on the above metrics, in the in-situ data storage management system, state is used to represent the key attributes of each storage queue in each round. These attributes accurately describe the current status of the queue, thus providing a basis for scheduling and processing decisions. In round t, state S t It can be represented as:

[0087] S t ={AWR i AUD i APL i}, i∈{HPD, MPD, LPD} (44)

[0088] In an in-situ data storage management system, an action represents the selected storage queue, that is, selecting a server with corresponding performance to perform scheduling processing. In round t, action A... t The storage queue to be selected for the operation can be represented as follows:

[0089] A t ∈{HPD, MPD, LPD} (45)

[0090] In in-situ data storage management systems, reward functions are used to quantify the effectiveness of queue scheduling strategies, enabling the system to adjust its decision-making process based on reward values, thereby optimizing overall performance. To measure the system's performance across different rounds, three optimization metrics are selected to construct the reward function. Since these metrics have different dimensions, normalization is prioritized to prevent any single metric from having an excessive impact on the reward function. Specifically, the normalized value of each metric at round t is calculated using the following formula:

[0091]

[0092] Where X(t) represents the original index value at round t, X max (t) and X min (t) represents the maximum and minimum values ​​of the indicator up to round t, respectively. In this way, the normalized indicator values ​​are all in the range [0, 1], making different indicators comparable within the same numerical range. and Let φ1, φ2, and φ3 represent the normalized index values ​​at round t. After normalization, the three indices are weighted and summed according to their weight coefficients φ1, φ2, and φ3 to obtain the final reward function.

[0093]

[0094] Among them, 1 {·} For characteristic functions, only if The value is 1 if it is positive and 0 otherwise. The weighting coefficient λ is used to adjust the selection queue Q. i When the delay limit is exceeded The penalty weights are determined by the reward function. The goal of this reward function is to minimize the three optimization metrics of the system under the constraint of a delay upper limit. Therefore, its value is always a non-positive number, and the larger the value (i.e., the closer to 0), the better the system performance. Through this design, the agent can optimize the scheduling strategy by maximizing the reward, thereby improving the overall system efficiency.

[0095] This invention patent utilizes an adversarial dual deep Q-network algorithm to solve the in-situ data storage management decision for load balancing in an IoT environment.

[0096] Compared with existing technologies, the significant advantages of this invention are: (1) The in-situ servers deployed near IoT terminals are abstracted into a hierarchical data queue structure. Servers with different performance levels in the system are abstracted into server queues with different performance levels to process storage units with different priorities. (2) A buffer is set up to receive in-situ data generated by distributed terminals and encapsulate the attributes of the in-situ data. Then, through a priority classification mechanism, the data categories are divided considering queue load and priority mutation. (3) The Dueling DQN algorithm is used to make processing decisions based on the state of the hierarchical data queue structure, ensuring high efficiency of system data access. Attached Figure Description

[0097] Figure 1 This is a flowchart of the in-situ data storage management process based on Dueling DQN and load balancing in the Internet of Things scenario of this invention.

[0098] Figure 2 This is a diagram of the Dueling DQN algorithm framework in this invention. Detailed Implementation

[0099] The present invention will now be described in further detail with reference to the accompanying drawings.

[0100] This invention discloses an in-situ data storage management method combining reinforcement learning and load balancing in an Internet of Things (IoT) scenario, characterized by the following steps:

[0101] Step 1: Set up a hierarchical data queue and abstract the storage unit, as follows:

[0102] Combination Figure 1 In an in-situ server system within an IoT environment, servers are deployed near distributed terminals to receive and process in-situ data. By combining servers with varying performance and functionalities, and configuring them differently in terms of data processing capabilities and storage response speeds, the diverse data storage needs of the Industrial IoT can be met. Furthermore, these servers with different capabilities can be abstracted into corresponding hierarchical data queues: Q HPDUsed to store high-priority data (HPD), this queue is abstracted from a high-performance server and is suitable for storing critical data with high real-time requirements and high access frequency. Q MPD Used for storing Medium Priority Data (MPD) type data, this queue consists of servers with moderate performance, balancing access efficiency and storage cost, and is suitable for storing periodic task data. Q LPD Used for storing low-priority data (LPD), this queue is abstracted from cost-optimized servers and is suitable for data with low access frequency and long lifecycles, emphasizing storage economy. The hierarchical data queue (HDQ) can be specifically represented as:

[0103] HDQ={Q HPD Q MPD Q LPD} (1)

[0104] Each hierarchical data queue includes the following characteristics: Queue Q represents i Maximum data capacity Queue Q represents i The number of storage units processed. Indicates the latency of processed memory units. Res i Queue Q represents i When selected for a single scheduler, the maximum amount of computational resources available per round is used to constrain dequeue operations: M represents the CPU computing power that can be allocated per round. i B represents the available memory capacity per round. i Network bandwidth available per round.

[0105]

[0106] To receive in-situ data generated by all distributed terminals, a unified data buffer is set up, encapsulating the main characteristics of each in-situ data in vector form: F represents the frequency of data access or reading, measuring the number of times the data is accessed within a certain period. C represents the storage space occupied by the data, usually in bytes. T represents the duration the data needs to be stored, usually measured in time. P represents the initial priority of the data in the system. These four characteristics are the key basis for implementing data classification and scheduling.

[0107] D = {F, C, T, P} (4)

[0108] Based on the characteristics of the data encapsulated in the data buffer, a scheduling classification strategy is designed to classify all preprocessed in-situ data into three categories: HPD, MPD, and LPD. Upon receiving in-situ data of the corresponding category, different storage queues will encapsulate it into a storage unit (SU), defined in detail below:

[0109] SU = {D, T} in T out ,κ},κ={κ CPU κ mem , κ bw} (5)

[0110] Where T in Used to record the enqueue round number of a storage unit, initialized to the current round number T when the storage unit was added to the queue. current T out This is used to record the dequeue round number of a storage unit. It is initialized to null, indicating that the storage unit has not yet been processed or removed from the queue. When the storage unit is dequeued, it is updated to the current round number T at the time of dequeue. current κ represents the resources required to complete the processing of each storage unit in this round: κ CPU Indicates the required CPU cycles; κ mem Indicates the required memory; κ bw This indicates the required bandwidth.

[0111] Step 2: Develop a data priority classification strategy for load balancing, as follows:

[0112] Combination Figure 1 To adaptively store storage units of different priorities into appropriate hierarchical data queues, a load-balancing data priority classification strategy is introduced. Taking into account the static characteristics of the storage units themselves, a static priority score G is calculated:

[0113] G = w F ·F+w C ·C+w T ·T+(1+μ)·P (6)

[0114] Among them, w F w S and w T These represent the priority weights of each static feature. To describe the suddenness and unpredictability of in-situ data priority in the IoT environment, a coefficient of variation μ is introduced to describe the probabilistic abrupt changes in data priority. The coefficient of variation μ follows a Bernoulli distribution:

[0115]

[0116] Where ∈ is a small constant representing the probability of triggering a priority mutation; μ high This is a relatively large constant representing the increase in priority during sudden changes. Data generally maintains its own priority, but there is a certain probability that it will suddenly change to extremely high priority.

[0117] Design a static classification function Category based on the static priority score G. static (D) Perform preliminary classification and discrimination on the in-situ data, where G high G low The score threshold used for classifying categories:

[0118]

[0119] Each in-situ data point will be initially divided into three categories—HPD, MPD, and LPD—based on a static priority G score. The number of data points in each category after classification is as follows: and At this point, the load on each queue after static priority classification is:

[0120]

[0121] To further optimize the classification results based on static priority, a dynamic priority adjustment mechanism with adaptive threshold adjustment is introduced. This mechanism dynamically adjusts the classification threshold by sensing the load of each queue in the hierarchical data queue architecture.

[0122]

[0123] Where λ is used as an adjustment factor to control the threshold update rate, and the updated threshold G′ is used to control the threshold update rate. high and G′ low Considering the load on each queue in a hierarchical data queue architecture, the data is reclassified:

[0124]

[0125] The specific quantities for each category after classification are as follows: and At this point, the result of dynamic classification makes the load ratio of each queue close to its computing power ratio, satisfying the following constraint:

[0126]

[0127] To standardize the weights of different features in the static priority score, each time a storage unit is classified, log data is stored, recording all data features and the dynamic priority classification threshold used for classification. Correlation analysis is then used to calibrate the weights in the priority score G using the log data. For any feature X (the static features F, C, T, and P of the storage unit), its correlation with the dynamic classification threshold G′ can be calculated. high and G′ low Correlation coefficient:

[0128]

[0129] in Used to calculate the expectation. To ensure the weights sum to 1, a normalization method is used to calculate the weights based on two thresholds:

[0130]

[0131] Each feature yields two normalized weight values ​​based on a classification threshold. To better calibrate the weights in the static priority score G, a balancing coefficient θ is introduced to control the bias of the final weights. The final weight calculation formula is as follows:

[0132]

[0133] Step 3: Propose optimization issues for in-situ data storage management and server load balancing, as follows:

[0134] Combination Figure 1 In IoT scenarios, the optimization of in-situ data storage management and server load balancing primarily addresses how to load-balance storage units with different characteristics to appropriate storage queues and implement strategic management responses. Reinforcement learning algorithm-driven management strategies significantly impact the effectiveness of system data processing.

[0135] In the in-situ data storage management process within an IoT environment, the data generation rate is typically affected by the system's current load and operating status. To ensure the stability of the hierarchical data queue structure during management, a Markov-modulated Poisson process is introduced. v The Modified Poisson Process (MMPP) is used to determine the total amount of in-situ data generated in each round. At round t, the load X of each queue needs to be calculated first. i (t) Case:

[0136]

[0137] To model using MMPP, the load value X needs to be... iThe load (t) is discretized into a finite number of state categories. The load value is transformed into three levels—high H, medium M, and low L—using a mapping function g(·). Furthermore, the overall load state of the system at round t consists of the states of all queues.

[0138] state i (t)=g(X i (t)), state i (t)∈{H,M,L} (19)

[0139] state(t) = (state HPD (t), state MPD (t), state LPD (t)) (20)

[0140] Historical data from the in-situ server system is time-series, recording the load of each queue across multiple scheduling rounds, and the actual amount of in-situ data N(t) generated in round t. All rounds T with the same state stat(t) are then processed. state (t), and by calculating the average data generation rate under the same state stat(t), the Poisson process parameter corresponding to that state, i.e., the data generation rate, is estimated.

[0141] T state(t) =t′|state(t′)=state(t) (21)

[0142]

[0143] The above method can establish a mapping relationship f between load state and generation rate, which can then be used for subsequent simulation to generate specific amounts of data.

[0144] λ(state(t))=f(state HPD (t), state MPD (t), state LPD (t)) (23)

[0145] Considering the suddenness and unpredictability of in-situ data, the amount of data N(t) generated in round t can be obtained using MMPP modeling. The mutation factor M(t) is used to increase the suddenness of data generation, and will mutate to a larger constant value M with a relatively low probability p. high This causes the number of units generated in that round to suddenly reach a higher peak.

[0146]

[0147] N(t)=Poisson(M(t)·λ(state(t))) (25)

[0148] The changes in the system state follow the transition rules of a Markov chain, meaning the state in the next round depends only on the state in the current round and is independent of earlier historical states. Through discretization of the historical data load, the finite state space S of the system can be defined as:

[0149] S={state1, state2,..., state |S|} (26)

[0150] We can statistically analyze the transitions of each state from historical data and construct a state transition counting matrix M. The elements of M are... j Indicates the state i Transition to state j The number of times. 1(·) is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise.

[0151]

[0152] The state transition probability matrix P is obtained by normalizing the state transition count matrix. i,j This indicates the current state of the system. i In the next round, the state transitions to that state. j The probability of the state transition is given. Due to the properties of Markov processes, the state transition probability matrix P satisfies the constraint that the sum of the probabilities is 1:

[0153]

[0154] Using the state transition matrix P, the states of future rounds can be deduced sequentially from the known initial state state(t). The system's state state(t) in round t determines the state state(t+1) in round t+1, unaffected by earlier states. Based on the deduced state state(t+1), the amount of data generated in round t+1 can be further determined.

[0155] state(t+1)~P(state(t)) (30)

[0156] N(t+1)~Poisson(λ(state(t+1)))(31)

[0157] After the MMPP model generates the total amount of data in round t, the data buffer and classification strategy will classify the data into three specific categories. The number of data entries in each queue corresponds to the number of data entries after classification. and In each round, each queue will perform an enqueue operation based on the quantity of data categorized by priority. First, calculate... The system calculates the storage requirements of each piece of data and compares them with the remaining available capacity of the queue. If the storage requirements are less than the remaining capacity of the queue, the enqueue operation is discarded and the data is returned to the data buffer; otherwise, the enqueue operation is performed. Each data point is added to the queue, and the corresponding enqueue round number is updated.

[0158] In a hierarchical data queue architecture, if data units from multiple server queues are processed simultaneously, the actual resources available to each queue will be significantly lower than the theoretical limit due to system-level contention. Therefore, the system comprehensively considers the states of different queues across the entire system and selects only a single queue for optimal scheduling in each round. The data buffer and classification algorithm categorize the in-situ data generated by IoT terminals into three types, which are then stored in storage queues abstracted from servers of different performance levels. When a single queue is selected for response in each round, the available resources are Res. i The number of dequeues from a specific storage unit, ξ i This can be expressed as:

[0159]

[0160] The dequeue method of the storage queue follows the First-In-First-Out (FIFO) principle, meaning that storage units in the queue are processed sequentially according to the enqueue order. In each round, the system selects a queue Q. i Perform scheduling processing and retrieve the corresponding number of storage units ξ from the head of the queue. i After dequeueing is complete, the queue's characteristics need to be updated, including the cumulative number of scheduled processes. and total delay in completion At the same time, update the number of rounds for leaving the team. out .

[0161] Step 4: A Markov decision process model is proposed to describe the in-situ data storage management problem. Based on this model, the Dueling DQN algorithm is used to comprehensively consider the load differences of hierarchical data queues and dynamically optimize the processing response decisions of in-situ data storage units to reduce response latency and avoid blocking of storage tasks. Specifically:

[0162] Combination Figure 2 In the process of in-situ data storage management in IoT scenarios, the amount of in-situ data in the hierarchical data queue structure changes dynamically according to the enqueue and dequeue strategies. Furthermore, each round requires selecting a storage queue for response. Due to the dynamic nature and decision complexity of this problem, it is well-suited for reinforcement learning methods to make optimal decisions based on the system state. Therefore, this problem needs to be abstracted into a Markov model:

[0163]

[0164] To evaluate the response performance of in-situ data storage units driven by different algorithms, three metrics were set: Average Waiting Rounds (AWR), which calculates the average number of waiting rounds for unprocessed data in all queues, directly reflects the timeliness of the system in data scheduling and processing; Average Unprocessed Data (AUD), which calculates the average amount of unprocessed data in all queues, reflects the adaptability of the system in terms of task processing capacity; and Average Processing Latency (APL), which calculates the average time taken for all scheduled data, reflects the efficiency of the system in data scheduling and processing. These are expressed as follows:

[0165]

[0166] Because these indicators have different dimensions, to avoid any one indicator having too large an impact on the reward function, they will be normalized first.

[0167]

[0168] In addition, the upper bound of the delay for all queues in each round needs to be calculated. Up to round T... current Queue Q i Minimum number of teams dispatched per round in historical data and the global expected size of the storage unit It can be represented as:

[0169]

[0170] For any queue Q i The maximum latency a storage unit may experience It can be represented as:

[0171]

[0172] Based on the above metrics, in the in-situ data storage management system, state is used to represent the key attributes of each storage queue in each round. These attributes accurately describe the current status of the queue, thus providing a basis for scheduling and processing decisions. In round t, state S t It can be represented as:

[0173] S t ={AWR i AUD i APL i}, i∈{HPD, MPD, LPD} (44)

[0174] In an in-situ data storage management system, an action represents the selected storage queue, that is, selecting a server with corresponding performance to perform scheduling processing. In round t, action A... tThe storage queue to be selected for the operation can be represented as follows:

[0175] A t ∈{HPD, MPD, LPD} (45)

[0176] In in-situ data storage management systems, reward functions are used to quantify the effectiveness of queue scheduling strategies, enabling the system to adjust its decision-making process based on reward values, thereby optimizing overall performance. To measure the system's performance across different rounds, three optimization metrics are selected to construct the reward function. Since these metrics have different dimensions, normalization is prioritized to prevent any single metric from having an excessive impact on the reward function. Specifically, the normalized value of each metric at round t is calculated using the following formula:

[0177]

[0178] Where X(t) represents the original index value at round t, X max (t) and X min (t) represents the maximum and minimum values ​​of the indicator up to round t, respectively. In this way, the normalized indicator values ​​are all in the range [0, 1], making different indicators comparable within the same numerical range. and Let φ1, φ2, and φ3 represent the normalized index values ​​at round t. After normalization, the three indices are weighted and summed according to their weight coefficients φ1, φ2, and φ3 to obtain the final reward function.

[0179]

[0180] Among them, 1 {·} For characteristic functions, only if The value is 1 if it is positive and 0 otherwise. The weighting coefficient λ is used to adjust the selection queue Q. i When the delay limit is exceeded The penalty weights are determined by the reward function. The goal of this reward function is to minimize the three optimization metrics of the system under the constraint of a delay upper limit. Therefore, its value is always a non-positive number, and the larger the value (i.e., the closer to 0), the better the system performance. Through this design, the agent can optimize the scheduling strategy by maximizing the reward, thereby improving the overall system efficiency.

[0181] This invention patent utilizes an adversarial dual deep Q-network algorithm to solve the in-situ data storage management decision for load balancing in an IoT environment. The algorithm pseudocode can be defined as follows:

[0182]

[0183]

[0184] The above description outlines the implementation process and advantages of this invention. Those skilled in the art should understand that various changes and modifications can be made to this invention without departing from its principles, and all such changes and modifications fall within the scope of the invention as claimed.

Claims

1. A method for in-situ data storage management combining reinforcement learning and load balancing in an IoT scenario, characterized in that, The following steps are involved: Step 1: Set up hierarchical data queues and abstract storage units Step 2: Develop a load balancing data priority classification strategy. Step 3: Propose optimization issues for in-situ data storage management and server load balancing. Step 4: Propose a Markov decision process model to describe the in-situ data storage management problem. Based on this model, use the Dueling DQN algorithm to comprehensively consider the load differences of hierarchical data queues and dynamically optimize the processing response decision of the in-situ data storage unit to reduce response latency and avoid blocking of storage tasks.

2. The in-situ data storage management method combining reinforcement learning and load balancing in the Internet of Things scenario described in claim 1, characterized in that, Step 1 describes setting up a hierarchical data queue and abstracting storage units, as follows: In an in-situ server system within an IoT environment, servers are deployed near distributed terminals to receive and process in-situ data. By combining servers with varying performance and functionalities, and configuring them differently in terms of data processing capabilities and storage response speeds, the diverse data storage needs of the Industrial IoT can be met. Furthermore, these servers with different capabilities can be abstracted into corresponding hierarchical data queues: Q HPD Used to store high-priority data (HPD), this queue is abstracted from the highest-performing server and is suitable for storing critical data with high real-time requirements and high access frequency. Q MPD Used for storing Medium Priority Data (MPD) type data, this queue consists of servers with moderate performance, balancing access efficiency and storage cost, and is suitable for storing periodic task data. Q LPD Used to store low-priority data (LPD), this queue is abstracted from cost-optimized servers and is suitable for data with low access frequency and long lifecycles, emphasizing storage economy. The hierarchical data queue (HDQ) can be specifically represented as: HDQ={Q HPD ,Q MPD ,Q LPD } (1) Each hierarchical data queue includes the following characteristics: Queue Q represents i Maximum data capacity Queue Q represents i The number of storage units processed. Indicates the latency of processed memory units. Res i Queue Q represents i When selected for a single scheduler, the maximum amount of computational resources available per round is used to constrain dequeue operations: M represents the CPU computing power that can be allocated per round. i B represents the available memory capacity per round. i Network bandwidth available per round. To receive in-situ data generated by all distributed terminals, a unified data buffer is set up, encapsulating the main characteristics of each in-situ data in vector form: F represents the frequency of data access or reading, measuring the number of times the data is accessed within a certain period. C represents the storage space occupied by the data, usually in bytes. T represents the duration the data needs to be stored, usually measured in time. P represents the initial priority of the data in the system. These four characteristics are the key basis for implementing data classification and scheduling. D = {F, C, T, P} (4) Based on the characteristics of the data encapsulated in the data buffer, a scheduling classification strategy is designed to classify all preprocessed in-situ data into three categories: HPD, MPD, and LPD. Upon receiving in-situ data of the corresponding category, different storage queues will encapsulate it into a storage unit (SU), defined in detail below: SU={D,T in ,T out ,k},k={k CPU ,k mem ,k bw } (5) Where T in Used to record the enqueue round number of a storage unit, initialized to the current round number T when the storage unit was added to the queue. current T out This is used to record the dequeue round number of a storage unit. It is initialized to null, indicating that the storage unit has not yet been processed or removed from the queue. When the storage unit is dequeued, it is updated to the current round number T at the time of dequeue. current κ represents the resources required to complete the processing of each storage unit in this round: κ CPU Indicates the required CPU cycles; κ mem Indicates the required memory; κ bw This indicates the required bandwidth.

3. The in-situ data storage management method combining reinforcement learning and load balancing in the Internet of Things scenario described in claim 1, characterized in that, Step 2 describes the data priority classification strategy for load balancing, as follows: To adaptively store storage units of different priorities into appropriate hierarchical data queues, a load-balancing data priority classification strategy is introduced. Taking into account the static characteristics of the storage units themselves, a static priority score G is calculated: G=w F ·F+w C ·C+w T ·T+(1+μ)·P (6) Among them, w F w S and w T These represent the priority weights of each static feature. To describe the suddenness and unpredictability of in-situ data priority in the IoT environment, a coefficient of variation μ is introduced to describe the probabilistic abrupt changes in data priority. The coefficient of variation μ follows a Bernoulli distribution: Where ∈ is a small constant representing the probability of triggering a priority mutation; μ high This is a relatively large constant representing the increase in priority during sudden changes. Data generally maintains its own priority, but there is a certain probability that it will suddenly change to extremely high priority. Design a static classification function Category based on the static priority score G. static (D) Perform preliminary classification and discrimination on the in-situ data, where G high G low The score threshold used for classifying categories: Each in-situ data point will be initially divided into three categories—HPD, MPD, and LPD—based on a static priority G score. The number of data points in each category after classification is as follows: and At this point, the load on each queue after static priority classification is: To further optimize the classification results based on static priority, a dynamic priority adjustment mechanism with adaptive threshold adjustment is introduced. This mechanism dynamically adjusts the classification threshold by sensing the load of each queue in the hierarchical data queue architecture. Where λ is used as an adjustment factor to control the update rate of the threshold, and the updated threshold G′ is used to control the update rate of the threshold. high and G′ low Considering the load on each queue in a hierarchical data queue architecture, the data is reclassified: The specific quantities for each category after classification are as follows: and At this point, the result of dynamic classification makes the load ratio of each queue close to its computing power ratio, satisfying the following constraint: To standardize the weights of different features in the static priority score, each time a storage unit is classified, log data is stored, recording all data features and the dynamic priority classification threshold used for classification. Correlation analysis is then used to calibrate the weights in the priority score G using the log data. For any feature X (the static features F, C, T, and P of the storage unit), its correlation with the dynamic classification threshold G′ can be calculated. high and G′ low Correlation coefficient: in Used to calculate the expectation. To ensure the weights sum to 1, a normalization method is used to calculate the weights based on two thresholds: Each feature yields two normalized weight values ​​based on a classification threshold. To better calibrate the weights in the static priority score G, a balancing coefficient θ is introduced to control the bias of the final weights. The final weight calculation formula is as follows: 。 4. The in-situ data storage management method combining reinforcement learning and load balancing in the Internet of Things scenario described in claim 1, characterized in that, Step 3 describes the optimization problems of in-situ data storage management and server load balancing, as follows: The optimization of in-situ data storage management and server load balancing in IoT scenarios mainly addresses how to load balance storage units with different characteristics to appropriate storage queues and implement strategic management responses. Management strategies driven by reinforcement learning algorithms can significantly impact the effectiveness of system data processing. In the process of in-situ data storage management in an IoT environment, the data generation rate is typically affected by the current system load and operating status. To ensure the stability of the hierarchical data queue structure during management, a Markov-Modulated Poisson Process (MMPP) is introduced to determine the total amount of in-situ data generated in each round. At round t, the load X of each queue needs to be calculated first. i (t) Case: To model using MMPP, the load value X needs to be... i (t) is discretized into a finite number of state categories. The load value is transformed into three levels: high H, medium M, and low L by the mapping function g(·). Furthermore, the overall load state of the system at round t consists of the states of all queues. state i (t)=g(X i (t)),state i (t)∈{H,M,L} (19) state(t)=(state HPD (t),state MPD (t),state LPD (t)) (20) Historical data from the in-situ server system is time-series, recording the load of each queue across multiple scheduling rounds, and the actual amount of in-situ data N(t) generated in round t. All rounds T with the same state stat(t) are then processed. state (t), and by calculating the average data generation rate under the same state stat(t), the Poisson process parameter corresponding to that state, i.e., the data generation rate, is estimated. T state (t)=t′|state(t′)=state(t) (21) The above method can establish a mapping relationship f between load state and generation rate, which can then be used for subsequent simulations to generate specific amounts of data. λ(state(t))=f(state HPD (t),state MPD (t),state LPD (t)) (23) Considering the suddenness and unpredictability of in-situ data, the amount of data N(t) generated in round t can be obtained using MMPP modeling. The mutation factor M(t) is used to increase the suddenness of data generation, and will mutate to a larger constant value M with a relatively low probability p. high This causes the number of units generated in that round to suddenly reach a higher peak. N(t)=Poisson(M(t)·λ(state(t))) (25) The changes in the system state follow the transition rules of a Markov chain, meaning the state in the next round depends only on the state in the current round and is independent of earlier historical states. Through the discretization of historical data load, the system's finite state space... It can be defined as: We can statistically analyze the transitions of each state from historical data and construct a state transition counting matrix M. The elements of M... i,j Indicates the state i Transition to state j The number of times. 1(·) is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise. The state transition probability matrix P is obtained by normalizing the state transition count matrix. i,j This indicates the current state of the system. i In the next round, the state transitions to that state. j The probability of the state transition is given. Due to the properties of Markov processes, the state transition probability matrix P satisfies the constraint that the sum of the probabilities is 1: Using the state transition matrix P, the states of future rounds can be deduced sequentially from the known initial state state(t). The system's state state(t) in round t determines the state state(t+1) in round t+1, unaffected by earlier states. Based on the deduced state state(t+1), the amount of data generated in round t+1 can be further determined. state(t+1)~P(state(t)) (30) N(t+1)~Poisson(λ(state(t+1))) (31) After the MMPP model generates the total amount of data in round t, the data buffer and classification strategy will classify the data into three specific categories. The number of data entries in each queue corresponds to the number of data entries after classification. and In each round, each queue will perform an enqueue operation based on the quantity of data categorized by priority. First, calculate... The system calculates the storage requirements of each piece of data and compares them with the remaining available capacity of the queue. If the storage requirements are less than the remaining capacity of the queue, the enqueue operation is discarded and the data is returned to the data buffer. Otherwise, perform the enqueue operation, and Each data point is added to the queue, and the corresponding enqueue round number is updated. In a hierarchical data queue architecture, if data units from multiple server queues are processed simultaneously, the actual resources available to each queue will be significantly lower than the theoretical limit due to system-level contention. Therefore, the system comprehensively considers the states of different queues across the entire system and selects only a single queue for optimal scheduling in each round. The data buffer and classification algorithm categorize the in-situ data generated by IoT terminals into three types, which are then stored in storage queues abstracted from servers of different performance levels. When a single queue is selected for response in each round, the available resources are Res. i The number of dequeueed units ξ of specific storage units i This can be expressed as: The dequeue method of the storage queue follows the First-In-First-Out (FIFO) principle, meaning that storage units in the queue are processed sequentially according to the enqueue order. In each round, the system selects a queue Q. i Perform scheduling processing and retrieve the corresponding number of storage units ξ from the head of the queue. i After dequeueing is complete, the queue's characteristics need to be updated, including the cumulative number of scheduled processes. and total delay in completion At the same time, update the number of rounds for leaving the team. out .

5. The in-situ data storage management method combining reinforcement learning and load balancing in the Internet of Things scenario described in claim 1, characterized in that, Step 4 describes the use of a Markov decision process model to describe the in-situ data storage management problem. Based on this model, the Dueling DQN algorithm is used to dynamically optimize the processing response decisions of in-situ data storage units by comprehensively considering the load differences of hierarchical data queues, thereby reducing response latency and avoiding blocking of storage tasks. Specifically: In the process of in-situ data storage management in IoT scenarios, the amount of in-situ data in the hierarchical data queue structure changes dynamically according to the enqueue and dequeue strategies. Furthermore, each round requires selecting a storage queue for response. Due to the dynamic nature and decision complexity of this problem, reinforcement learning methods are well-suited for optimal decision-making based on the system state. Therefore, this problem needs to be abstracted into a Markov model: To evaluate the response performance of in-situ data storage units driven by different algorithms, three metrics were set: Average Waiting Rounds (AWR), which calculates the average number of waiting rounds for unprocessed data in all queues, directly reflects the timeliness of the system in data scheduling and processing; Average Unprocessed Data (AUD), which calculates the average amount of unprocessed data in all queues, reflects the adaptability of the system in terms of task processing capacity; and Average Processing Latency (APL), which calculates the average time taken for all scheduled data, reflects the efficiency of the system in data scheduling and processing. These are expressed as follows: Because these indicators have different dimensions, to avoid any one indicator having too large an impact on the reward function, they will be normalized first: In addition, the upper bound of the delay for all queues in each round needs to be calculated. Up to round T... current Queue Q i Minimum number of teams dispatched per round in historical data and the global expected size of the storage unit It can be represented as: For any queue Q i The maximum latency a storage unit may experience It can be represented as: Based on the above metrics, in the in-situ data storage management system, state is used to represent the key attributes of each storage queue in each round. These attributes accurately describe the current status of the queue, thus providing a basis for scheduling and processing decisions. In round t, state S t It can be represented as: S t ={AWR i ,AUD i ,APL i },i∈{HPD,MPD,LPD} (44) In an in-situ data storage management system, an action represents the selected storage queue, that is, selecting a server with corresponding performance to perform scheduling processing. In round t, action A... t The storage queue to be selected for the operation can be represented as follows: A t ∈{HPD,MPD,LPD} (45) In in-situ data storage management systems, reward functions are used to quantify the effectiveness of queue scheduling strategies, enabling the system to adjust its decision-making process based on reward values, thereby optimizing overall performance. To measure the system's performance across different rounds, three optimization metrics are selected to construct the reward function. Since these metrics have different dimensions, normalization is prioritized to prevent any single metric from having an excessive impact on the reward function. Specifically, the normalized value of each metric at round t is calculated using the following formula: Where X(t) represents the original index value at round t, X max (t) and X min (t) represents the maximum and minimum values ​​of the indicator up to round t, respectively. In this way, the normalized indicator values ​​are all in the range [0, 1], making different indicators comparable within the same numerical range. and Let φ1, φ2, and φ3 represent the normalized index values ​​at round t. After normalization, the three indices are weighted and summed according to their weight coefficients φ1, φ2, and φ3 to obtain the final reward function. Among them, 1 {·} For characteristic functions, only if The value is 1 if it is positive and 0 otherwise. The weighting coefficient λ is used to adjust the selection queue Q. i When the latency limit is exceeded The penalty weights are determined by the reward function. The goal of this reward function is to minimize the three optimization metrics of the system under the constraint of a delay upper limit. Therefore, its value is always a non-positive number, and the larger the value (i.e., the closer to 0), the better the system performance. Through this design, the agent can optimize the scheduling strategy by maximizing the reward, thereby improving the overall system efficiency. This invention patent utilizes an adversarial dual deep Q-network algorithm to solve the in-situ data storage management decision for load balancing in an IoT environment. The algorithm pseudocode can be defined as follows: