Mass monitoring data-oriented edge computing storage optimization algorithm and system

By using edge computing storage optimization algorithms for massive monitoring data, the data migration path between edge devices, computing centers and the cloud is dynamically adjusted, solving the problem of insufficient flexibility in resource allocation and path optimization in existing technologies, and achieving efficient utilization of system resources and performance improvement.

CN120610657APending Publication Date: 2025-09-09NANJING NANDA SIWEI TECHNOLOGY DEVELOPMENT CO LTD

Patent Information

Application Number
CN202510610910.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing data storage systems lack flexibility in storage resource configuration and path optimization, and are unable to dynamically respond to changes in storage requirements, resulting in resource waste and performance degradation.

Method used

It adopts edge computing storage optimization algorithms for massive monitoring data, dynamically adjusts the data migration path between edge devices, computing centers and the cloud through layered storage, data migration mechanism, state construction, deep reinforcement learning-driven dynamic scheduling and time series prediction optimization, and realizes adaptive optimization of resources.

Benefits of technology

It achieves dynamic optimization of system resources, ensures efficient operation when storage requirements change, avoids over- or under-allocation of resources, and improves data access speed and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610657A_ABST
    Figure CN120610657A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data storage management, and discloses an edge computing storage optimization algorithm and system for mass monitoring data, and the algorithm comprises the following steps: S1, a hierarchical storage step: dividing data into real-time data, processed data and long-term storage data based on the real-time performance, importance and access frequency of the monitoring data; s2, a data storage migration mechanism step; s3, a state construction step; s4, an action space definition step; s5, a dynamic scheduling step driven by deep reinforcement learning; S6, a time sequence prediction optimization step; and S7, a cross-domain collaborative optimization step. By adopting the technical scheme based on long-term performance evaluation and an adaptive optimization algorithm, dynamic optimal configuration of system resources is realized. According to the technology, the storage path and resource allocation can be adjusted in real time according to the system operation state, and compared with a fixed configuration scheme in the prior art, the problems of non-uniform utilization of storage resources and performance fluctuation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data storage management technology, and in particular to an edge computing storage optimization algorithm for massive monitoring data. Background Art

[0002] Currently, many data storage systems employ fixed resource allocation and path optimization strategies. These systems often determine the amount of storage resources and the selection of storage paths during design. During system operation, resource allocation and path selection are not dynamically adjusted to actual load and demand changes. The main problem with this approach is that storage requirements and data access patterns change over time, but the system cannot respond to these changes in a timely manner, resulting in wasted or insufficient storage resources.

[0003] Existing systems for the aforementioned technologies typically rely on pre-set resource configurations, lacking flexibility. This static configuration approach cannot effectively address sudden increases or changes in storage demand. For example, when a storage path becomes overloaded, the system lacks a mechanism to automatically migrate data to a less-loaded path, resulting in increased access latency and decreased performance.

[0004] Furthermore, traditional storage resource management methods often overlook long-term trends in storage demand, resulting in the system's inability to continuously optimize under varying load conditions. Existing storage path optimization techniques are typically set up and then not adjusted, making them unable to dynamically adapt to changing demand. Summary of the Invention

[0005] The purpose of this invention is to provide an edge computing storage optimization algorithm for massive monitoring data, which solves the problem that the existing data storage system cannot dynamically respond to demand changes during storage resource configuration, storage path optimization and load balancing, resulting in waste of storage resources and performance degradation.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an edge computing storage optimization algorithm for massive monitoring data, comprising the following steps: S1. Hierarchical storage step: Based on the real-time nature, importance, and access frequency of the monitoring data, the data is divided into real-time data, processed data, and long-term storage data. The data type characteristics are obtained and stored in the edge device, edge computing center, and cloud storage respectively. S2. Data storage migration mechanism steps: Dynamically adjust the data migration path between edge devices, computing centers, and the cloud based on the access patterns, lifecycles, and resource constraints of various monitoring data types; S3, state construction step: Based on the hierarchical storage and data migration mechanism, the current state of the system is extracted, including the load status of various storage nodes, the distribution of monitoring data, the waiting time of computing tasks, and access frequency information, which is used to define the state space of the deep reinforcement learning model; S4, action space definition step: define a set of scheduling operations based on storage resource utilization, data type characteristics, and dynamically adjust data migration paths to form the action space of the deep reinforcement learning model; S5. Dynamic scheduling driven by deep reinforcement learning: Based on the state space and action space, a deep reinforcement learning algorithm is used to train the scheduling agent to dynamically optimize the data storage location and the allocation strategy of computing tasks. S6, Time Series Prediction Optimization Step: Based on the historical change trends of the monitored data, a time series prediction model is constructed to predict future data access and storage requirements, and data compression and deduplication are performed accordingly. S7, cross-domain collaboration optimization steps: Build a graph structure between edge nodes and the cloud, and use graph optimization algorithms to optimize cross-domain resource allocation and data migration paths.

[0007] Preferably, in the tiered storage step, real-time data is stored in a high-performance cache of edge devices, processed data is stored in a solid-state memory of an edge computing center, and long-term storage data is regularly archived to an object-based cloud storage service.

[0008] Preferably, the hierarchical storage step specifically includes the following steps: Collect the generation time, access times and access intervals of monitoring data; Calculate the access frequency and combine it with the real-time weight to derive a comprehensive priority score; Match the corresponding storage tier based on the scoring results and store the data in edge device cache, edge computing center storage nodes, or cloud object storage resources; Records the current level tag of the data when storing it.

[0009] Preferably, the data storage migration mechanism step specifically includes the following steps: Real-time monitoring of resource usage of each storage node, including CPU, memory, and bandwidth; Calculate the load index of each storage node based on resource utilization and data access frequency; Determine the priority of data migration and select nodes with lower load as migration targets; Execute data migration tasks to migrate data from high-load nodes to low-load nodes, and adjust network bandwidth allocation to ensure data transmission efficiency during the migration process.

[0010] Preferably, the state building step specifically includes the following steps: Collect real-time resource usage of each storage node, including CPU utilization, memory usage, storage space, and network bandwidth; Record the storage location and access frequency of monitoring data at each node and establish a data distribution model; Get the queue information of computing tasks, including the waiting time, priority and resource requirements of the tasks; Based on the above information, a state vector is constructed and input into the deep reinforcement learning model for decision-making and optimization.

[0011] Preferably, the action space definition step specifically includes the following steps: Define data migration operations, task scheduling operations, and resource allocation operations based on the available resources of storage nodes; Define data storage level change operations based on different data type characteristics, including real-time data migration from edge devices to computing centers and processed data migration from computing centers to the cloud. Define task scheduling operations based on the computing requirements and priorities of the tasks; The final action set is formed as the action space of the deep reinforcement learning model.

[0012] Preferably, the deep reinforcement learning-driven dynamic scheduling step includes training a scheduling agent based on a state space and an action space using a deep reinforcement learning algorithm to optimize the data storage location and the allocation strategy of computing tasks in real time.

[0013] Preferably, the time series prediction optimization step specifically includes the following steps: Collect historical access records of monitoring data, including access time, frequency, and data volume; Use the long short-term memory network time series prediction model to predict future access popularity and storage requirements; Adjust storage resource allocation based on prediction results and perform data compression, deduplication, or data migration in advance. Regularly update the time series model to ensure the accuracy of prediction results and the optimal configuration of storage resources.

[0014] Preferably, the cross-domain collaborative optimization step specifically includes the following steps: Build a graph structure between edge nodes and the cloud, where a node represents a storage resource or computing resource, and an edge represents the connection bandwidth between resources; Assign resource weights to edges in the graph based on the load, bandwidth constraints, and task priorities of each node; Use graph optimization algorithms to calculate the optimal cross-domain resource allocation plan and determine the data migration path; Perform data migration tasks to ensure that data is migrated from nodes with higher loads to nodes with lower loads, while optimizing bandwidth usage during the migration process.

[0015] Edge computing storage optimization system for massive monitoring data, including: Data acquisition module, used to collect monitoring data; Data hierarchical storage module, used for hierarchical storage according to the characteristics of monitoring data; Data migration module, used to dynamically adjust the storage location of data based on system load and data access patterns; Scheduling optimization module, which is used for dynamic scheduling based on deep reinforcement learning algorithms to optimize data storage and task allocation; Time series prediction module, used to predict future storage needs based on historical data and adjust resource allocation; The cross-domain collaboration module is used to optimize resource allocation and data migration paths between edge nodes and the cloud.

[0016] In summary, the present invention includes at least one of the following beneficial technical effects: 1. This invention achieves dynamic optimization of system resources by employing a technical solution based on long-term performance evaluation and an adaptive optimization algorithm. This technology can adjust storage paths and resource allocation in real time based on system operating status. Compared to the fixed configuration solutions used in existing technologies, it solves the problems of uneven storage resource utilization and performance fluctuations.

[0017] 2. By introducing a storage demand prediction model, this invention enables the system to predict future storage needs based on historical and real-time data and automatically adjust resource allocation. Compared with traditional static resource management methods, this avoids over- or under-allocation of resources and ensures efficient system operation as demand changes.

[0018] 3. This invention continuously optimizes the system's storage paths and load distribution by implementing a regular evaluation and feedback mechanism. This differs from existing single-shot optimization schemes, ensuring stable performance over the long term and effectively preventing the impact of outdated optimization decisions.

[0019] 4. This invention improves data access speed by optimizing storage path selection and load balancing strategies during dynamic resource scheduling and data migration. Compared to existing solutions that under-optimize path selection and load distribution, this solution solves the problem of long response times and improves data access efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 Schematic diagram of the system framework of the present invention. DETAILED DESCRIPTION

[0021] The following is combined with Figure 1 , the present invention is described in further detail.

[0022] The present invention provides an edge computing storage optimization algorithm and system for massive monitoring data. like Figure 1 As shown, the edge computing storage optimization algorithm for massive monitoring data may include the following steps: S1. Hierarchical storage step: Based on the real-time nature, importance, and access frequency of the monitoring data, the data is divided into real-time data, processed data, and long-term storage data. The data type characteristics are obtained and stored in the edge device, edge computing center, and cloud storage respectively. Specifically, the goal of step S1 in this embodiment is to divide the data into real-time data, processed data, and long-term storage data based on the real-time nature, importance, and access frequency of the monitoring data, and store them in different storage media to optimize system performance. The specific implementation method is as follows: First, the system collects and analyzes monitoring data to determine its characteristics, including key information such as data generation time, access frequency, and access patterns. Based on this information, the system divides the data into three categories: real-time data, processed data, and long-term storage data.

[0023] Data classification and tiered storage Real-time data storage; Real-time data refers to the immediate data collected from monitoring devices. This data is highly time-sensitive and typically requires rapid processing and access. This type of data is preferably stored using high-performance cache storage on edge devices. Using high-performance storage media such as flash memory or solid-state drives (SSDs) enables low-latency data access, ensuring the system can respond to changes in monitoring data in real time.

[0024] Storage of processed data; Processed data refers to data that has been processed by the edge computing center and needs to be retained long-term. This data typically involves specific data analysis or calculation results and no longer changes frequently, but still requires on-demand access. This data is preferably stored in the edge computing center's solid-state storage (SSD), which has large storage space and high read and write speeds to support frequent access needs.

[0025] long-term storage of data; Long-term storage data refers to data that is no longer frequently accessed and updated. This type of data is typically historical records of monitoring equipment or archived data that needs to be retained for a long time. To reduce storage costs, this data is stored in object-based storage in the cloud. Cloud storage offers higher storage capacity and lower storage costs, making it suitable for long-term data storage needs that don't require fast access.

[0026] To achieve this tiered storage, the system employs a comprehensive scoring method that combines data access frequency, real-time weighting, and importance scores. First, the system calculates data access frequency by collecting access records of monitored data (including generation time, number of accesses, and access intervals). Next, the system weights the data based on real-time requirements and importance, ultimately deriving a comprehensive priority score for each data type. This score determines the tier in which the data should be stored.

[0027] The calculation formula for this comprehensive priority score is as follows; S=w1·F+w2·T+w3·P; Among them, S is the comprehensive priority score, F is the access frequency of the data, T is the real-time weight of the data, P is the importance score of the data, w1, w2, and w3 are the access frequency, real-time and importance weight coefficients respectively, satisfying w1+w2+w3=1.

[0028] Based on this scoring result, data is classified as real-time data, processed data, or long-term storage data. Real-time data has a higher priority and is stored in the high-performance cache of edge devices. Processed data is stored in the edge computing center. Long-term storage data has a lower priority and is stored in cloud object storage.

[0029] During data storage, the system also records the storage path and hierarchical label for each type of data. The storage path indicates the data's current storage location (e.g., local to the edge device, in an edge computing center, or in the cloud). This label provides a reference for subsequent data scheduling and migration modules, enabling appropriate storage migration when data access patterns change.

[0030] The storage level label is closely related to the life cycle of the data. The data is labeled when it is initially stored. During the data migration process, the label will be updated based on changes in the real-time nature of the data, access frequency, etc.

[0031] Data migration mechanism and scheduling Over time, the access frequency and storage requirements of monitoring data may change. To ensure optimal system performance, the system should have a data migration mechanism that can dynamically adjust the storage location of data based on data access patterns and storage requirements. This mechanism monitors data access patterns, lifecycles, and storage resource constraints, and regularly migrates data, migrating frequently accessed data from the cloud to the edge computing center and migrating important but less frequently accessed data to cloud storage, thereby reducing storage costs and improving system response speed. S2. Data storage migration mechanism steps: Dynamically adjust the data migration path between edge devices, computing centers, and the cloud based on the access patterns, lifecycles, and resource constraints of various monitoring data types; Specifically, step S2 of this embodiment involves dynamically adjusting and optimizing storage resources based on real-time monitoring data. The core technical feature of this step is to analyze monitoring data using a mathematical model to calculate and generate storage resource optimization decisions, allowing the system to automatically adjust storage allocation based on real-time data changes. The specific implementation is as follows: First, the system receives real-time data from monitoring devices or edge computing centers and processes it through a data analysis module. The core task of data analysis is to generate optimization decisions based on the current state of storage resources and data access patterns. Key to this process is the use of mathematical models to predict data access and storage demand, thereby guiding the allocation of storage resources.

[0032] The analysis process first involves extracting features from real-time data, including metrics such as access frequency, storage duration, and data importance. Based on these features, the system then uses a multidimensional regression model to predict future storage needs. Specifically, the variables are: F(t) represents the data access frequency at time t; D(t) represents the data storage requirement at time t; P(t) represents the importance of data at time t; C(t) represents the data storage cost at time t.

[0033] By analyzing historical data, a regression model is established. The model formula is as follows: D(t)=α·F(t)+β·P(t)+γ·C(t); Where: D(t) represents the total demand at time t, F(t) represents the data access frequency at time t, P(t) represents the data importance at time t, C(t) represents the storage cost at time t, and α, β, and γ represent the weight coefficients of data access frequency, data importance, and storage cost, respectively.

[0034] Optimizing decision making After establishing a data analysis model, the system uses an optimization algorithm to generate optimal storage resource decisions. This decision is based on the following goals: minimizing storage costs while ensuring fast access to high-frequency data and long-term retention of important data. To this end, the system uses dynamic programming (DP) to make optimization decisions based on the storage costs and access times of different storage paths. The optimization objective function is: Among them, C i is the cost of the i-th storage path, x i is the allocation ratio of the i-th storage path, and n is the total number of storage paths.

[0035] Through this optimization decision, the system can dynamically adjust the data distribution ratio among different storage paths such as edge computing centers and cloud storage to minimize storage costs and meet access requirements.

[0036] After data analysis and optimization decisions are made, the system allocates and schedules data storage resources. Specifically, the system uses a control module to translate optimization decisions into actual operations and allocates data storage paths using a resource scheduling algorithm. These storage paths include edge device storage, edge computing center storage, and cloud storage. Each type of data is allocated to the most appropriate storage medium based on its timeliness, access frequency, and importance.

[0037] The specific scheduling operations are as follows: High-frequency access data is preferentially allocated to the SSD cache area of ​​the edge device to ensure low-latency data access; Data that is more important but less frequently accessed is stored in the edge computing center to balance storage costs and access requirements; Long-term storage data is transferred to cloud storage to further reduce storage costs.

[0038] During the scheduling process, the system automatically completes the storage path allocation of data through the interface between the storage scheduling module and the storage device.

[0039] Data migration is another key task in this step. Based on real-time data changes, the system dynamically adjusts data storage locations using a prediction-based migration decision algorithm. Specifically, when the system detects a change in the access frequency of certain data, it re-evaluates the data's storage path based on the aforementioned analysis model and performs data migration. Suppose there is a type of data whose data storage path at time t1 is S1. However, at time t2, the model predicts that its storage requirements have changed, and the storage path should be updated to S2. At this point, the system makes data migration decisions based on the following formula: ΔC=C(S2)-C(S1); Where ΔC is the storage cost difference of the storage path after migration, and C(S2) and C(S1) are the storage costs of the new and old storage paths, respectively. If the storage cost difference ΔC is negative, the data migration operation is performed.

[0040] S3, state construction step: Based on the hierarchical storage and data migration mechanism, the current state of the system is extracted, including the load status of various storage nodes, the distribution of monitoring data, the waiting time of computing tasks, and access frequency information, which is used to define the state space of the deep reinforcement learning model; Specifically, the goal of step S3 in this embodiment is to perform data migration based on the analysis and optimization decisions made in the previous two steps, ensuring that the data storage path remains consistent with its access frequency, importance, and storage cost. This step is key to achieving intelligent storage resource management. Through intelligent data migration and storage path updates, data storage efficiency can be further improved and dynamic optimization can be achieved. The specific implementation method is as follows: First, the system receives the data analysis and optimization decision from step S2. Based on the storage path strategy determined in the optimization decision, the system begins data migration. The key to this process is to execute data migration operations based on the actual data storage requirements and the priority of the storage paths. Each data migration operation involves migrating from the current storage path to a new one. Data migration decisions are based on the optimal storage path determined by the predictive model and optimization algorithm.

[0041] To ensure efficient data migration between different storage paths, the system first establishes a data migration decision model. This model makes decisions based on factors such as data access frequency, importance score, current storage cost, and storage path read and write speed. The following variables are set to describe the key characteristics of data migration: M(t) represents the data migration demand at time t; S1 indicates the current storage path; S2 represents the target storage path; C(S1) represents the cost of the current storage path; C(S2) represents the cost of the target storage path; R(S1) represents the read and write speed of the current storage path; R(S2) represents the read and write speed of the target storage path.

[0042] According to the data access pattern prediction model, the data migration decision formula is as follows: M(t)=δ·(C(S1)-C(S2))+φ·(R(S1)-R(S2)); Where δ and φ are weight coefficients, representing the impact of storage cost and read / write speed on the data migration decision, respectively. If M(t)>0, the data migration operation is performed.

[0043] The data migration trigger system calculates the data migration requirements based on the migration decision model and triggers the data migration operation if M(t) > 0. The data migration trigger conditions are based on changes in storage path costs and read / write speeds, ensuring timely data migration based on storage path optimization requirements.

[0044] Once data migration is triggered, the system transfers data from its current storage path (e.g., an edge computing center) to its target storage path (e.g., cloud storage). This transfer is performed using efficient data transfer protocols to ensure data integrity and speed. For larger datasets, the system can transfer data in batches to avoid network congestion or latency caused by transferring too much data all at once.

[0045] After data migration is complete, the system updates the data's storage path tag. This tag records the data's new storage location for subsequent scheduling and access. The storage path update process is performed through the database management system, ensuring that all data's storage paths are consistent with the latest optimization decisions in the system.

[0046] During the data migration process, the system ensures data consistency and integrity. To prevent data loss or corruption during the migration process, the system uses a transactional database for data management. All migration operations are performed under transaction control, ensuring data consistency before and after migration.

[0047] In addition, the system also includes a data verification mechanism. After the data migration is complete, the migrated data will be verified for integrity and availability on the target storage path. If the verification fails, the system will automatically roll back the migration operation and re-execute the data migration.

[0048] After data migration is complete, the system reschedules storage paths to further optimize data storage and access. The system uses an intelligent scheduling module to monitor and process data access requests in real time, ensuring that data is quickly responded to on the most appropriate storage path.

[0049] Key tasks in the scheduling process include: Load balancing: Ensures that the load on each storage path is evenly distributed to avoid response delays caused by overloading of a storage path.

[0050] Cache management: Provides cache support for frequently used data to reduce access latency.

[0051] Data compression and deduplication: Optimize storage space and ensure efficient use of storage resources.

[0052] S4, action space definition step: define a set of scheduling operations based on storage resource utilization, data type characteristics, and dynamically adjust data migration paths to form the action space of the deep reinforcement learning model; Specifically, the core purpose of step S4 in this embodiment is to provide dynamic feedback, adjustment, and optimization based on the monitored data storage status and resource usage during system operation, ensuring that the system can be promptly optimized based on actual changes and continuously improving storage performance. By monitoring storage status in real time and combining it with storage resource usage data, the system can continuously optimize storage paths, scheduling strategies, and resource allocation methods.

[0053] First, the system continuously receives real-time monitoring data from storage devices, compute nodes, and other devices. This monitoring data includes, but is not limited to, storage space utilization, data access frequency, storage costs, and storage path read and write speeds. The system analyzes this data in real time and compares it with predefined optimization targets. If any significant deviations in the monitoring data occur, the system activates a feedback mechanism to further optimize storage resources.

[0054] To support real-time storage resource adjustments, the system first establishes a monitoring data analysis model to analyze and evaluate storage status and resource usage. This model comprehensively analyzes monitoring data and makes storage optimization decisions based on learning from historical data. The following variables are set to describe the system monitoring and feedback mechanism: U(t) represents the usage rate of storage resources at time t; F(t) represents the access frequency of data at time t; S(t) represents the space requirement for data storage at time t; T(t) represents the response time of the storage resource at time t; C(t) represents the cost of the storage path at time t.

[0055] Through these variables, the system establishes a storage status feedback model, which is used to measure the difference between the actual usage of storage resources and the ideal storage status. The formula is as follows: ΔU(t)=λ1·(U(t)-U ideal )+λ2·(F(t)-F ideal )+λ3·(S(t)-S ideal ); Among them, λ1, λ2, and λ3 are weight coefficients, which respectively represent the influence of storage resource utilization, access frequency, and storage demand on feedback adjustment. ideal ,F ideal ,S idealare the storage utilization rate, data access frequency, and storage demand under ideal conditions, respectively. When the value of ΔU(t) is greater than the set threshold, the feedback adjustment mechanism is triggered.

[0056] Feedback adjustment mechanism Feedback Signal Generation: When the system calculates that ΔU(t) exceeds a set threshold, it generates a feedback signal, indicating that storage resources need to be adjusted. Specifically, feedback signals include the need to increase storage space, adjust data storage paths, or modify storage path access policies.

[0057] The adjustment decision generation system uses feedback signals and optimization algorithms (such as dynamic programming or greedy algorithms) to calculate optimal decisions. These decisions involve operations such as updating storage paths and expanding or contracting storage resources. The system adjusts resource allocation and storage path configuration based on real-time data access and storage requirements.

[0058] Set the following optimization goals: minimize storage costs, maximize the efficiency of storage resource utilization, and optimize the access speed of storage paths.

[0059] After generating an adjustment decision, the system uses the scheduling module to make actual adjustments to storage resources. The adjustment process includes: Increase or decrease storage space to ensure that the capacity of each storage path adapts to actual storage needs; Adjust data storage paths to ensure that frequently accessed data is migrated to storage paths with faster response times. Optimize storage path access strategies to ensure load balancing and response speed among different storage paths.

[0060] The system builds a feedback adjustment optimization model based on factors such as storage resource usage data, storage demand data, and storage path costs to ensure that feedback adjustment has the greatest effect. The optimization goals are: Among them, C i represents the cost of the i-th storage path, T i represents the response time of the i-th storage path, x i represents the ratio of resources allocated to the i-th storage path, β i Represents the weight of the storage path response time. Through this optimization decision, the system can dynamically adjust the resource allocation of the storage path to minimize storage costs and maximize storage access speed.

[0061] Through this optimization decision, the system can dynamically adjust the resource allocation of storage paths, minimize storage costs and maximize storage access speed.

[0062] Once the feedback adjustment mechanism is activated, the system dynamically allocates and schedules resources through the resource scheduling module. Specifically, the system updates the load of each storage path in real time based on storage path usage and adjustment decisions to ensure balanced resource allocation.

[0063] Key tasks in the scheduling process include: Dynamically adjust storage capacity: Increase or decrease the capacity of storage paths based on changes in storage requirements to avoid overload or resource waste; Storage path optimization: Adjust the storage location of high-frequency data to ensure fast access to important data and reduce latency; Load balancing and optimization: Intelligently adjust the load of each storage path based on the load conditions of different storage paths to ensure data access efficiency.

[0064] S5. Dynamic scheduling driven by deep reinforcement learning: Based on the state space and action space, a deep reinforcement learning algorithm is used to train the scheduling agent to dynamically optimize the data storage location and the allocation strategy of computing tasks. Specifically, the core purpose of step S5 of this embodiment is to perform self-optimization based on the monitoring and feedback mechanism of the previous steps, combined with the system operation data and feedback information, to ensure that the system can always maintain optimal performance during long-term operation. The implementation of this step relies on a long-term performance evaluation model and an adaptive optimization algorithm, which ensures the continuous and efficient utilization of system resources through real-time adjustment and optimization. The specific implementation method is as follows: First, the system continuously receives feedback from step S4, including storage path optimization, resource allocation strategy, load balancing, and storage space usage. The system evaluates this feedback in real time and performs self-optimization based on the performance evaluation model.

[0065] To ensure the system maintains optimal performance over the long term, the system first establishes a long-term performance evaluation model. This model evaluates the overall system performance based on long-term monitoring data and generates self-optimization decisions based on the evaluation results. The following key variables are set to describe the performance evaluation process: E(t) represents the overall performance score of the system at time t; R util (t) represents the resource utilization at time t; T response (t) represents the system response time at time t; L load (t) represents the load of the storage path at time t; C cost (t) represents the storage cost at time t.

[0066] The system has established the following performance evaluation model by analyzing long-term operation data: E(t)=α1·R util (t)+α2·T response (t)+α3·L load (t)+α4·C cost (t); Among them, α1, α2, α3, and α4 are weight coefficients, which respectively represent the impact of resource utilization, response time, load, and storage cost on the overall performance. This evaluation model provides data support for subsequent self-optimization by continuously tracking the operating status of the system.

[0067] After obtaining the performance evaluation results, the system will use an adaptive optimization algorithm to make adjustments. This algorithm dynamically adjusts storage paths, resource allocation, load balancing strategies, and other aspects based on the system's long-term performance evaluation results. The specific optimization process includes: The adaptive optimization decision generation system generates adaptive optimization decisions based on the evaluation model output, combined with current storage requirements, resource utilization, and performance goals. This decision is based on a dynamic programming algorithm based on the optimization goal. The optimization goal is set as: Among them, C i is the cost of the i-th storage path, T i is the response time of the i-th storage path, x i is the ratio of resources allocated to the i-th storage path, β i is the weight coefficient of response time.

[0068] Through this optimization decision, the system can balance storage costs, access efficiency, and response time, ensuring that storage resources are reasonably allocated in the long term.

[0069] After generating an optimization decision, the system will adjust the storage path, resource allocation, and load balancing strategies based on the decision. Specific adjustments include: Adjust the storage path configuration to ensure that frequently accessed data is stored on a storage path with a faster response speed.

[0070] Dynamically allocate resources to ensure load balancing of storage resources and avoid storage path overload or resource waste.

[0071] Optimize storage path access strategies, reduce system response time, and improve data access efficiency.

[0072] After the adaptive optimization algorithm executes, the system re-evaluates the optimized performance and generates new evaluation results. This process forms a closed loop of self-optimization and feedback. After each optimization, the system evaluates the results and further adjusts the system configuration based on the evaluation results to ensure optimal long-term performance. If system performance still does not meet ideal standards after optimization, the system will re-activate the feedback mechanism to adjust the adaptive optimization algorithm parameters to further optimize storage paths, resource allocation, and load balancing strategies. The system continuously monitors key indicators such as storage paths, resource utilization, storage costs, and access efficiency, conducts regular performance evaluations, and makes self-optimization adjustments based on the evaluation results. The system can adaptively respond to changes in storage demand, changes in data access patterns, and fluctuations in resource usage.

[0073] S6, Time Series Prediction Optimization Step: Based on the historical change trends of the monitored data, a time series prediction model is constructed to predict future data access and storage requirements, and data compression and deduplication are performed accordingly. Specifically, the purpose of step S6 of this embodiment is to regularly evaluate and update the system to ensure that the configuration of storage resources and the optimization strategy of storage paths are always kept in the best state. This step performs regular resource configuration and strategy updates based on the feedback results of long-term performance evaluation and adaptive optimization, thereby achieving continuous improvement of system performance. This step ensures that the system can cope with changing storage requirements and access patterns through real-time monitoring and evaluation of key performance indicators. The specific implementation method is as follows: First, the system regularly evaluates system performance based on the long-term performance evaluation model and optimization decision results after adaptive optimization in step S5, and examines the configuration of storage resources, the optimization strategy of storage paths, and the efficiency of data access. The system uses the performance evaluation results to analyze the efficiency of storage resource utilization, and adjusts resource configuration and updates path optimization according to demand.

[0074] To regularly evaluate the system, a performance evaluation model was first established to analyze the system's performance over different time periods and guide updates to resource allocation and storage path optimization strategies. The following key variables were set to describe the regular evaluation process: P(t) represents the total performance score of the system at time t S capacity (t) represents the capacity usage of storage resources at time t; A access (t) represents the data access frequency at time t; D demand (t) represents the demand for storage resources at time t; C cost (t) represents the cost of the storage path at time t.

[0075] The evaluation model is: P(t)=β1·S capacity (t)+β2·A access (t)+β3·D demand (t)+β4·C cost (t); Where β1, β2, β3, and β4 are weight coefficients, representing the degree to which storage resource capacity, data access frequency, storage requirements, and storage path costs affect overall system performance. The system periodically evaluates and calculates P(t) and, based on the results, determines whether adjustments to storage paths or resource allocation strategies are necessary.

[0076] Based on the results of regular evaluations, the system generates new adjustment recommendations based on the optimization decision model. This process first determines whether there are insufficient resource allocations or suboptimal storage paths. If so, an optimization decision is generated. The optimization decision is based on: Among them, C i represents the cost of the i-th storage path, T i represents the response time of the i-th storage path, x i represents the ratio of resources allocated to the i-th storage path, γ i is the weight coefficient of storage path response time.

[0077] The system generates new resource configuration and storage path optimization decisions based on the evaluation results.

[0078] After obtaining a new optimization decision, the system will update the storage path and adjust the resource configuration. During this process, the system performs the following operations: The storage path adjustment system migrates data from inefficient storage paths to more suitable ones based on optimization decisions. Migration decisions take into account data access frequency, storage requirements, and storage costs, prioritizing storage paths with faster response times and lower costs.

[0079] The resource allocation adjustment system dynamically expands or contracts storage resources to ensure that resources on each storage path can meet real-time needs. If the load on a storage path is too high, the system will migrate some data to other storage paths to avoid storage resource overload.

[0080] The load balancing and access optimization system ensures efficient use of storage resources by adjusting the load and access policies of storage paths, preventing overloaded storage paths and resulting in response delays. The system also optimizes storage path access policies to increase data access speed.

[0081] Regular monitoring and evaluation feedback mechanism After each optimization, the system will ensure the effectiveness of the optimization decision through continuous monitoring and evaluation feedback mechanism. The specific process is as follows: The monitoring data collection system regularly collects key performance data, including storage resource utilization, data access frequency, storage path response time, etc. This data will serve as input for the next evaluation.

[0082] The optimization evaluation system will recalculate the performance score based on the collected monitoring data to verify whether the optimized system meets the performance goals. If the system performance does not meet the expectations, the system will re-execute the optimization decision and further adjust the resource allocation.

[0083] S7, cross-domain collaboration optimization steps: Build a graph structure between edge nodes and the cloud, and use graph optimization algorithms to optimize cross-domain resource allocation and data migration paths.

[0084] Specifically, the purpose of step S7 in this embodiment is to enable the system to dynamically respond to changes in storage requirements during operation, including adjustments to storage resource scheduling and data migration strategies. This step monitors changes in system storage requirements in real time, combines optimization algorithms with resource management models, adjusts storage paths, and optimizes storage resource configuration. This strategy enables the system to efficiently respond to changes in storage requirements, maintaining efficient system operation and rational resource allocation. Specific implementations are as follows: First, the system builds a prediction model for storage demand changes based on the periodic evaluation results from step S6 and storage resource usage, combined with real-time data access and storage demand changes. This model analyzes the changing trends of storage demand and predicts storage resource demand over the next period of time.

[0085] To predict changes in storage demand, the system first establishes a storage demand forecasting model. This model uses historical data and real-time monitoring data to predict future storage demand. The following key variables are set to describe the storage demand forecasting process.

[0086] N(t) represents the storage requirement at time t; D access (t) represents the amount of data accessed at time t; P growth (t) represents the data storage growth rate at time t; R capacity (t) represents the available capacity of storage resources at time t.

[0087] The storage demand forecasting model is: N(t+1)=α1·D access (t)+α2·P growth (t)+α3·Rcapacity (t); Where N(t+1) represents the storage demand predicted at time t+1, and α1, α2, and α3 are the weight coefficients of the prediction model, which respectively represent the influence of data access volume, storage growth rate, and available resource capacity on the storage demand prediction.

[0088] Through this predictive model, the system can estimate future storage needs and provide a basis for resource scheduling and data migration decisions.

[0089] Based on the prediction results and real-time changes in storage demand, the system will use dynamic resource scheduling algorithms to reallocate storage resources and make data migration decisions. The dynamic resource scheduling process includes the following steps: The resource scheduling decision system is based on the predicted storage demand N(t+1) and the current storage resource availability R capacity (t) determines whether storage resources need to be expanded or contracted. If the predicted demand exceeds the current resource capacity, the system dynamically expands storage resources. Conversely, if demand decreases, the system reduces storage resource allocation.

[0090] Resource scheduling decisions are made based on the following optimization objectives: Among them, C i is the cost of the i-th storage path, T i is the response time of the i-th storage path, x i is the resource ratio allocated to the i-th storage path, δ i is the response time weight coefficient of the storage path, and there are n total choices or decision items. The system dynamically adjusts resource allocation based on this optimization goal to ensure that storage resources are allocated reasonably and meet predicted demand.

[0091] Based on resource scheduling decisions, the system also generates data migration strategies based on the storage path load and data access frequency. Data migration decisions are based on the following: Among them, M(t+1) is the amount of migration data at time t+1, θ i is the migration weight coefficient of storage path i, DataLoad i is the data load on storage path i, and n is the total number of data migration tasks.

[0092] This formula reflects the load of each storage path at different points in time. Based on this load, the system decides whether to migrate data from a storage path with a higher load to a path with a lower load.

[0093] After dynamic scheduling and data migration decisions are made, the system will optimize storage paths and resource configuration accordingly. This optimization process includes: Storage path selection: Based on optimization goals and data migration decisions, the system migrates data from slower, more heavily loaded storage paths to faster, less loaded storage paths. This helps reduce data access latency and optimizes overall system storage performance.

[0094] The resource allocation adjustment system adjusts the allocation of storage resources based on the results of dynamic scheduling. If storage resources are insufficient, the system will automatically increase them; if storage resources are excessive, the system will reduce them to avoid waste.

[0095] The load balancing optimization system also optimizes the load on each storage path to prevent a single path from being overloaded and affecting overall system performance. By balancing the load, the system can improve data access speed and storage efficiency.

[0096] The edge computing storage optimization system for massive monitoring data described below and the edge computing storage optimization algorithm for massive monitoring data described above can be referenced to each other.

[0097] Please see the attached Figure 2 ,The present invention also provides an edge computing storage optimization system for massive monitoring data, including; Data acquisition module, used to collect monitoring data; Data hierarchical storage module, used for hierarchical storage according to the characteristics of monitoring data; Data migration module, used to dynamically adjust the storage location of data based on system load and data access patterns; Scheduling optimization module, which is used for dynamic scheduling based on deep reinforcement learning algorithms to optimize data storage and task allocation; Time series prediction module, used to predict future storage needs based on historical data and adjust resource allocation; The cross-domain collaboration module is used to optimize resource allocation and data migration paths between edge nodes and the cloud.

[0098] The system of this embodiment can be used to execute the above method embodiments, and its principles and technical effects are similar, so they will not be repeated here.

[0099] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. Edge computing storage optimization algorithm for massive monitoring data, characterized by: The steps include: S1. Hierarchical storage step: Based on the real-time nature, importance, and access frequency of the monitoring data, the data is divided into real-time data, processed data, and long-term storage data. The data type characteristics are obtained and stored in the edge device, edge computing center, and cloud storage respectively. S2. Data storage migration mechanism steps: Dynamically adjust the data migration path between edge devices, computing centers, and the cloud based on the access patterns, lifecycles, and resource constraints of various monitoring data types; S3, state construction step: Based on the hierarchical storage and data migration mechanism, the current state of the system is extracted, including the load status of various storage nodes, the distribution of monitoring data, the waiting time of computing tasks, and access frequency information, which is used to define the state space of the deep reinforcement learning model; S4, action space definition step: define a set of scheduling operations based on storage resource utilization, data type characteristics, and dynamically adjust data migration paths to form the action space of the deep reinforcement learning model; S5. Dynamic scheduling driven by deep reinforcement learning: Based on the state space and action space, a deep reinforcement learning algorithm is used to train the scheduling agent to dynamically optimize the data storage location and the allocation strategy of computing tasks; S6, Time Series Prediction Optimization Step: Based on the historical change trends of the monitored data, a time series prediction model is constructed to predict future data access and storage requirements, and data compression and deduplication are performed accordingly. S7, cross-domain collaboration optimization steps: Build a graph structure between edge nodes and the cloud, and use graph optimization algorithms to optimize cross-domain resource allocation and data migration paths.

2. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that: In the tiered storage step, real-time data is stored in the high-performance cache of edge devices, processed data is stored in the solid-state memory of the edge computing center, and long-term storage data is regularly archived to the object-based cloud storage service.

3. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that The hierarchical storage step specifically includes the following steps: Collect the generation time, access times and access intervals of monitoring data; Calculate the access frequency and combine it with the real-time weight to derive a comprehensive priority score; Match the corresponding storage tier based on the scoring results and store the data in edge device cache, edge computing center storage nodes, or cloud object storage resources; Records the current level tag of the data when storing it.

4. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that The data storage migration mechanism step specifically includes the following steps: Real-time monitoring of resource usage of each storage node, including CPU, memory, and bandwidth; Calculate the load index of each storage node based on resource utilization and data access frequency; Determine the priority of data migration and select nodes with lower load as migration targets; Execute data migration tasks to migrate data from high-load nodes to low-load nodes, and adjust network bandwidth allocation to ensure data transmission efficiency during the migration process.

5. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that: The state building step specifically includes the following steps: Collect real-time resource usage of each storage node, including CPU utilization, memory usage, storage space, and network bandwidth; Record the storage location and access frequency of monitoring data at each node and establish a data distribution model; Get the queue information of computing tasks, including the waiting time, priority and resource requirements of the tasks; Based on the above information, a state vector is constructed and input into the deep reinforcement learning model for decision-making and optimization.

6. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that: The action space definition step specifically includes the following steps: Define data migration operations, task scheduling operations, and resource allocation operations based on the available resources of storage nodes; Define data storage level change operations based on different data type characteristics, including real-time data migration from edge devices to computing centers and processed data migration from computing centers to the cloud. Define task scheduling operations based on the computing requirements and priorities of the tasks; The final action set is formed as the action space of the deep reinforcement learning model.

7. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that: The deep reinforcement learning-driven dynamic scheduling step includes using a deep reinforcement learning algorithm to train a scheduling agent based on state space and action space, and optimizing the storage location of data and the allocation strategy of computing tasks in real time.

8. The edge computing storage optimization algorithm for massive monitoring data according to claim 1 is characterized in that: The timing prediction optimization step specifically includes the following steps: Collect historical access records of monitoring data, including access time, frequency, and data volume; Use the long short-term memory network time series prediction model to predict future access popularity and storage requirements; Adjust storage resource allocation based on prediction results and perform data compression, deduplication, or data migration in advance. Regularly update the time series model to ensure the accuracy of prediction results and the optimal configuration of storage resources.

9. The edge computing storage optimization algorithm for massive monitoring data according to claim 1, characterized in that: The cross-domain collaborative optimization step specifically includes the following steps: Build a graph structure between edge nodes and the cloud, where a node represents a storage resource or computing resource, and an edge represents the connection bandwidth between resources; Assign resource weights to edges in the graph based on the load, bandwidth constraints, and task priorities of each node; Use graph optimization algorithms to calculate the optimal cross-domain resource allocation plan and determine the data migration path; Perform data migration tasks to ensure that data is migrated from nodes with higher loads to nodes with lower loads, while optimizing bandwidth usage during the migration process.

10. An edge computing storage optimization system for massive monitoring data, according to the edge computing storage optimization algorithm for massive monitoring data according to any one of claims 1-9, characterized in that: include; Data acquisition module, used to collect monitoring data; Data hierarchical storage module, used for hierarchical storage according to the characteristics of monitoring data; Data migration module, used to dynamically adjust the storage location of data based on system load and data access patterns; Scheduling optimization module, which is used for dynamic scheduling based on deep reinforcement learning algorithms to optimize data storage and task allocation; Time series prediction module, used to predict future storage needs based on historical data and adjust resource allocation; The cross-domain collaboration module is used to optimize resource allocation and data migration paths between edge nodes and the cloud.

Citation Information

Patent Citations

  • Data classified storage method and device based on reinforcement learning

    CN117453123A

  • SaaS-based cloud platform data storage method

    CN118075293A

  • Power distribution network data storage method suitable for cloud computing

    CN118244987A

  • Storage strategy optimization method based on data life cycle

    CN118466858A

  • Cloud computing resource allocation method based on dynamic optimization, computer device and storage medium

    CN119166354A

Cited By

  • Distributed data storage dynamic optimization method based on edge computing

    CN121193759A

  • Solid state disk integrated edge computing cache updating method

    CN121704779A