Dynamic monitoring and early warning method for storage space of data center
Through technologies such as deep learning and reinforcement learning, accurate prediction and resource scheduling of data center storage requirements are achieved, and the shortcomings of dynamic changes in storage requirements and equipment failure prediction in the existing technology are solved, and the performance and resource utilization of the storage system are improved.
Patent Information
- Application Number
- CN202510534686.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing data center storage space monitoring and early warning methods are difficult to predict the dynamic changes in storage demand and the optimized configuration of different storage devices in real time and accurately, and there is a lag in the early warning of storage failures.
The deep learning model is used to predict historical stored data, combine reinforcement learning algorithms to dynamically schedule storage resources, use the fault prediction model to provide early warning, and perform automatic data migration or redundant backup through the self-healing mechanism to implement cross-level storage optimization.
It realizes efficient scheduling of data center storage resources, improves the performance and stability of the storage system under high load conditions, optimizes the allocation strategy of storage resources, reduces storage costs, and improves the utilization rate of storage resources.
Smart Images

Figure CN120045421A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data center storage management, and specifically provides a method for dynamically monitoring and warning the storage space of a data center. Background Art
[0002] With the advent of the information age, data centers have become an important part of the core IT architecture of modern enterprises. To ensure the efficient and stable operation of data centers, the dynamic monitoring and warning of storage space have become crucial. The monitoring of storage space not only involves the usage status and performance evaluation of storage devices, but also requires predicting the future demand for storage resources in order to adjust the configuration in a timely manner to avoid system performance degradation or resource waste caused by insufficient or excessive allocation of storage resources. Therefore, the method for dynamically monitoring and warning the storage space of a data center plays a key role in ensuring the optimal allocation of storage resources, improving system reliability, and reducing costs.
[0003] Existing methods for monitoring and warning the storage space of data centers usually collect basic operation status data of storage devices (such as temperature, storage capacity, I / O load, etc.) and perform predictive analysis in combination with historical data. These methods generally rely on rule engines or simple statistical analysis and can achieve the monitoring and warning of device anomalies. For example, by setting thresholds to monitor indicators such as the temperature and response time of storage devices, an alarm is triggered once the predetermined range is exceeded. These methods have the relatively intuitive advantage of being able to simply monitor and warn the system and respond to anomalies of storage devices in real time through a rule engine.
[0004] However, there are some obvious deficiencies in the existing technology, especially in dealing with the rapid changes in the storage requirements of data centers and the optimal configuration of the performance of different storage devices. Existing monitoring and warning systems mostly rely on static threshold settings, unable to predict the dynamic changes in storage requirements in real time and accurately, and also failing to effectively perform cross-level storage optimization and automatic resource scheduling. In addition, due to the lack of intelligent data analysis and prediction capabilities, the existing systems have a lag in warning of storage failures and it is difficult to effectively intervene before storage device failures occur. Summary of the Invention
[0005] In view of the deficiencies of the existing technology, the present invention provides a method for dynamically monitoring and warning the storage space of a data center, which solves the problems of the existing methods for monitoring and warning the storage space of data centers in dealing with the dynamic changes in storage requirements and predicting device failures.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for dynamically monitoring and warning the storage space of a data center, the method comprising the following steps: S1. Collect and monitor the operation status data of the storage devices in the data center in real time; S2. Based on the collected historical storage data, predict the future storage requirements through a deep learning model; S3. Based on the storage requirement prediction results, adopt a reinforcement learning algorithm to dynamically schedule the storage resources, automatically adjust the storage device configuration to optimize the use of storage resources; S4. When potential faults occur in the storage devices, use a fault prediction model to give early warnings, and execute automatic data migration or redundant backup through a self-healing mechanism; S5. Based on the data access frequency and importance, implement cross-level storage optimization, automatically migrate the data to different storage levels, and optimize the storage cost and performance.
[0007] Preferably, the S1 includes the following steps: Collect the data of the temperature, I / O load, response time, and storage capacity utilization rate of the storage devices in real time; Input the above monitoring data into a preset data storage system for subsequent analysis and prediction.
[0008] Preferably, the S2 includes the following steps: Use a long short-term memory network (LSTM) model to train the historical storage requirement data to predict the data storage requirements within a future period of time; Evaluate the trained LSTM model, select the optimal prediction model for deployment, and use the mean squared error (MSE) as the loss function to optimize the model.
[0009] Preferably, the step of using the mean squared error (MSE) as the loss function to optimize the model includes the following steps: Calculate the error between the actual storage requirements and the model prediction requirements, and optimize the weights of the LSTM network through the backpropagation algorithm; Minimize the following loss function: ; where, is the actual storage requirement, is the predicted value, is the number of samples.
[0010] Preferably, the S3 includes the following steps: Set the state space, action space, and reward function of the reinforcement learning model, and adopt the Q-learning algorithm for storage resource scheduling; Define the state space as the current usage state of the storage devices, and define the action space as the scheduling operations of the storage resources; Define the reward function as the optimized utilization efficiency of storage resources, and adjust the storage resource allocation strategy through the Q-learning algorithm.
[0011] Preferably, the Q-learning algorithm for storage resource scheduling includes the following steps: Initialize the Q-table, set the initial Q-value to zero, and update the Q-value according to the feedback of the storage resource usage situation; Use the following update formula to update the Q-value: ; Among them, is the reward value, indicating the improvement of storage resource utilization rate, is the discount factor, is the learning rate.
[0012] Preferably, the S4 includes the following steps: Collect the temperature, error log, and I / O latency information of the storage device, and input it into the fault prediction model; Use a convolutional neural network (CNN) to analyze the state of the storage device and predict the possible fault risks.
[0013] Preferably, the using a convolutional neural network (CNN) to analyze the state of the storage device includes the following steps: Extract the health state features of the storage device and input them into the CNN network for fault identification; Calculate the fault occurrence probability, use the softmax function to output the prediction result of the fault, and trigger an early warning when the prediction probability exceeds the preset threshold.
[0014] Preferably, the S5 includes the following steps: Classify the data into hot data, cold data, and expired data according to the access frequency and importance of the data; For hot data, store it preferentially on high-performance storage devices. For cold data, store it on low-cost storage devices, and expired data is cleared or archived regularly.
[0015] Preferably, the S5 further includes the following steps: According to the access pattern of the data, adopt a hybrid storage architecture for data storage and dynamically migrate the data between different storage levels; Use linear programming to optimize the data migration strategy to minimize the storage cost and meet the maximum capacity constraint of each storage level.
[0016] The present invention provides a method for dynamically monitoring and warning the storage space of a data center. It has the following beneficial effects: 1. By adopting the dynamic storage space monitoring and prediction method, the present invention realizes the efficient scheduling of storage resources in the data center. By real-time monitoring the status of storage devices and combining with a deep learning model to predict future storage requirements, it can accurately identify storage bottlenecks and make responses in advance, ensuring that the storage system can still maintain good performance and stability under high load conditions.
[0017] 2. Based on the reinforcement learning algorithm for dynamic scheduling of storage resources, the present invention optimizes the allocation strategy of storage resources. By adaptively adjusting the storage configuration, it improves the storage efficiency, reduces the risk of over-allocation, ensures the optimal storage performance under resource constraints, and reduces resource waste.
[0018] 3. The present invention introduces a cross-level storage optimization and data migration mechanism. According to the access frequency and importance of data, it automatically allocates data to appropriate storage device levels, and combines linear programming to optimize the migration strategy, which can effectively reduce storage costs and improve the utilization rate of storage resources, meeting performance requirements at different levels. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the specification of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] Please refer to the appended Figure 1 , the embodiment of the present invention provides a method for dynamic monitoring and early warning of the storage space in a data center. The method includes the following steps: S1. Collect and real-time monitor the operation status data of the storage devices in the data center. By real-time monitoring the operation status of the storage devices, it can timely detect the bottlenecks or potential fault points of the storage resources, providing an accurate data basis, laying a foundation for subsequent storage requirement prediction and resource scheduling, thereby improving the monitoring ability of the data center for the status of storage resources, reducing storage resource waste, and enhancing the availability of the devices; S2. Based on the collected historical storage data, predict the future storage requirements through a deep learning model. This step can accurately predict the changing trend of storage requirements, help the data center understand in advance the shortage or surplus of storage resources, thereby avoiding waste or shortage of resources. At the same time, through accurate prediction of storage requirements, it can improve the utilization efficiency of storage resources and avoid performance bottlenecks caused by fluctuations in storage requirements in the data center; S3. Based on the storage demand prediction results, a reinforcement learning algorithm is used to dynamically schedule storage resources, automatically adjust the storage device configuration to optimize the use of storage resources. By introducing the reinforcement learning algorithm, it can automatically adjust the configuration of storage resources according to the predicted demand, avoiding the complexity and errors of manual operations, optimizing the utilization efficiency of storage resources. The application of reinforcement learning ensures that the dynamic scheduling of storage resources can adapt to demand fluctuations in real time, improving the intelligent level of resource scheduling in the data center; S4. When potential failures occur in storage devices, a fault prediction model is used for early warning, and an automatic data migration or redundant backup is performed through a self-healing mechanism. This step can identify potential device failure risks in advance, reducing the suddenness of failures. Through the automated self-healing mechanism, it can ensure that when a storage device fails, data will not be lost and the system can continue to run stably. Overall, it reduces manual intervention and improves the reliability and service availability of the data center; S5. Based on data access frequency and importance, cross-layer storage optimization is implemented to automatically migrate data to different storage levels, optimizing storage cost and performance. Through cross-layer storage optimization, data can be reasonably allocated to different storage levels, ensuring that the high-performance storage requirements of hot data are met while reducing the storage cost of cold data. This method not only improves the utilization efficiency of storage resources but also effectively reduces the overall storage cost of the data center.
[0022] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, S1 includes the following steps: Collect data on the temperature, I / O load, response time, and storage capacity utilization rate of storage devices in real time. Real-time monitoring of temperature can help the data center detect temperature anomalies that may cause device failures in advance. By real-time monitoring of the I / O load, the performance bottleneck of storage devices can be identified in a timely manner, avoiding performance degradation and system crashes under high load conditions. By real-time monitoring of the response time of storage devices, problems can be detected in the initial stage of performance decline, avoiding bottlenecks that affect the overall system performance; Input the above monitoring data into a preset data storage system for subsequent analysis and prediction. Inputting the monitoring data into an efficient storage system can provide rich and accurate data support for subsequent prediction and scheduling algorithms. Through long-term accumulated data, the data storage system can support tasks such as efficient trend analysis, anomaly detection, and predictive analysis, improving the intelligent management level of the data center.
[0023] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, S2 includes the following steps: The historical storage demand data is trained using a Long Short-Term Memory (LSTM) model to predict the data storage demand for a period of time in the future. By using the LSTM model, the time dependence of the storage demand data can be effectively captured, and then the future storage demand can be accurately predicted. Compared with traditional prediction methods, LSTM can handle more complex time series data and identify long-term trends and periodic fluctuations. Therefore, adopting the LSTM model can improve the accuracy of prediction and avoid resource waste or shortage problems caused by inaccurate storage demand prediction; The trained LSTM model is evaluated, and the optimal prediction model is selected for deployment. The Mean Squared Error (MSE) is used as the loss function to optimize the model. By evaluating and optimizing the trained LSTM model in detail, the efficiency and accuracy of the deployed prediction model in practical applications can be ensured. Selecting the optimal prediction model for deployment can maximize the accuracy of storage demand prediction and provide reliable data support for subsequent storage resource scheduling and optimization.
[0024] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, using the Mean Squared Error (MSE) as the loss function to optimize the model includes the following steps: Calculate the error between the actual storage demand and the model prediction demand, and optimize the weights of the LSTM network through the backpropagation algorithm. Error calculation is not only the basis of the model learning process but also helps to identify potential problems in the prediction process, such as bias, overfitting, etc. Therefore, error calculation is a prerequisite for deep learning model optimization, which helps to improve the accuracy of model prediction. The backpropagation algorithm can optimize network parameters by calculating gradients, ensuring that the weights of the LSTM model are gradually adjusted to the optimal values. Through backpropagation optimization, the prediction error of the model will be minimized, improving the accuracy of storage demand prediction and providing more reliable decision support for the storage resource management of the data center; Minimize the following loss function: ; where is the actual storage demand, is the predicted value, and is the number of samples. Using MSE as the loss function can effectively guide the model to optimize its parameters and minimize the difference between the actual demand and the predicted demand. By repeatedly optimizing MSE, the LSTM model will be able to achieve higher prediction accuracy, thus providing more accurate support for the storage demand prediction of the data center and avoiding problems of resource waste or insufficiency.
[0025] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, S3 includes the following steps: Set the state space, action space, and reward function of the reinforcement learning model, and use the Q-learning algorithm for storage resource scheduling. The Q-learning algorithm is a model-free reinforcement learning algorithm that learns the strategy of taking the best action in different states by continuously updating the Q-value table. The Q-value (quality value) measures the expected return of taking a certain action in a certain state. The Q-learning algorithm can continuously improve the strategy through the balance of exploration and exploitation, and finally find the optimal resource scheduling scheme; Define the state space as the current usage state of the storage device, and define the action space as the scheduling operation of the storage resources; Define the reward function as the optimized usage efficiency of the storage resources, and adjust the storage resource allocation strategy through the Q-learning algorithm. By clearly defining the state space, action space, and reward function, sufficient information can be provided for the Q-learning algorithm to help the model select the optimal resource scheduling strategy among multiple operations. A reasonable reward function can guide the model to optimize the configuration of the storage resources through the reward mechanism, thereby improving the usage efficiency of the storage resources, reducing resource waste, and ensuring the efficient operation of the storage space in the data center.
[0026] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, the storage resource scheduling by the Q-learning algorithm includes the following steps: Initialize the Q-table, set the initial Q-value to zero, and update the Q-value according to the feedback of the storage resource usage situation; Use the following update formula for Q-value update: ; Among them, is the reward value, indicating the improvement of the storage resource utilization rate, is the discount factor, is the learning rate. Through repeated learning and optimization, Q-learning can provide a gradually improved scheduling strategy, avoid manual intervention, and improve the usage efficiency of the storage resources. In addition, the algorithm can handle complex changes in storage requirements, adapt to changing workloads and usage environments, and thus achieve the optimal configuration of the storage resources.
[0027] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, S4 includes the following steps: Collect the temperature, error logs, and I / O latency information of the storage device and input them into the fault prediction model. By collecting the temperature, error logs, and I / O latency information of the storage device, the operating status of the storage device can be comprehensively monitored. The real-time collection of this information provides rich data support for subsequent fault prediction, helps to detect potential fault risks in advance, and reduces the probability of system failures. Use a Convolutional Neural Network (CNN) to analyze the status of the storage device and predict possible fault risks. A Convolutional Neural Network (CNN) is a deep learning algorithm mainly used for image processing and time series data analysis. By using CNN for fault prediction, features can be effectively automatically extracted from multiple data sources (such as temperature, error logs, I / O latency), potential fault risks can be identified. CNN can learn complex patterns in the data and predict the possibility of device failures in advance. Through timely fault warnings, the losses caused by device failures can be reduced, and the stability and reliability of the data center can be improved.
[0028] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, using a Convolutional Neural Network (CNN) to analyze the status of the storage device includes the following steps: Extract the health status features of the storage device and input them into the CNN network for fault identification. By extracting health status features from multiple dimensions (such as temperature, latency, error rate, etc.), the operating status of the storage device can be comprehensively understood. CNN can automatically identify complex patterns in these features, providing high-quality input data for subsequent fault prediction. Automatic feature extraction avoids the cumbersome process of manual feature extraction, improving the accuracy and efficiency of fault prediction. Calculate the probability of a fault occurring, use the softmax function to output the prediction result of the fault, and trigger an alarm when the predicted probability exceeds a preset threshold. The definition of the Softmax function is: ; where is the original score output by the model (i.e., the output of CNN), is the probability of the corresponding category, is the base of the natural logarithm. By using the softmax function to calculate the probability of a fault occurring, the system can quantify the fault risk of the storage device and make decisions based on the actual fault probability. The preset threshold enables the warning system to be flexibly adjusted according to different fault tolerance requirements, ensuring the accuracy and timeliness of warning triggering. This probability-based prediction and warning mechanism can improve the intelligence level of the system, help operation and maintenance personnel identify potential faults in advance, and reduce the risk of device downtime or data loss.
[0029] Please refer to the appendix Figure 1, in a preferred embodiment of the present invention, S5 includes the following steps: According to the access frequency and importance of the data, the data is divided into hot data, cold data, and expired data. Among them, hot data refers to the data that is frequently accessed or modified, and cold data refers to the data that is infrequently accessed or has not been modified for a long time. By classifying the data into hot data, cold data, and expired data according to the access frequency and importance, efficient management of storage resources can be achieved. Hot data is preferentially stored on high-performance devices to ensure access speed; cold data is stored on low-cost devices to save storage expenses; the clearing or archiving of expired data effectively releases storage space and avoids valuable resources being occupied by useless data; For hot data, it is preferentially stored on high-performance storage devices. For cold data, it is stored on low-cost storage devices. Expired data is periodically cleared or archived. Storing hot data in high-performance storage devices can significantly improve the access efficiency of the data, reduce latency, and improve the response speed of the application program. By storing cold data on low-cost storage devices, the storage cost can be effectively reduced. This strategy can maximize the utilization of the storage resources in the data center, ensure that hot data can be quickly accessed, while cold data remains in the system for future reference without consuming high storage costs. Finally, by periodically clearing or archiving expired data, a large amount of storage space can be released, and useless data is prevented from occupying storage resources.
[0030] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, S5 further includes the following steps: According to the access pattern of the data, a hybrid storage architecture is adopted for data storage, and data is dynamically migrated between different storage levels. By adopting the hybrid storage architecture, the storage system can be optimized according to the data access pattern and storage requirements. While hot data can be quickly accessed, cold data can obtain a more economical storage solution, thereby achieving the optimization of storage cost and the improvement of performance. Through intelligent management of the storage levels, hot data always remains in high-performance storage devices, while cold data is stored at a lower cost. The dynamic migration mechanism enhances the flexibility and adaptability of the storage system, can quickly respond to changes in the data access pattern, and improves the utilization efficiency of storage resources; Use linear programming to optimize the data migration strategy to minimize the storage cost and meet the maximum capacity constraints of each storage level. The linear programming formula: Suppose there are storage levels (such as SSD, HDD, etc.), the storage capacity of each storage level is , the storage cost is , the amount of data stored is , and the goal of linear programming is to minimize the storage cost: ; At the same time, it is necessary to meet the capacity constraints of each storage tier: ; By solving this linear programming problem, an optimal data migration strategy can be obtained, which can maximize the reduction of storage costs while the storage configuration after data migration meets the performance requirements; By considering the capacity limitations and performance constraints of the storage tiers, linear programming can provide an optimal solution, avoiding over - or under - allocation of storage resources and improving the economic efficiency and operation efficiency of the data center storage system.
[0031] To better understand the present invention, the above content will be described in detail below in conjunction with specific embodiments. Embodiment
[0032] Embodiment 1: Dynamic Monitoring and Early Warning Method for Storage Space Based on Deep Learning Prediction and Reinforcement Learning Scheduling (Basic Embodiment)
[0033] In this embodiment, the system first realizes the dynamic monitoring of storage devices by collecting the health data of storage devices in real - time (such as temperature, I / O load, response time, and storage capacity utilization rate). On this basis, historical data and the LSTM (Long Short - Term Memory Network) model are used to predict future storage requirements, so as to evaluate storage pressure and resource bottlenecks in advance. Next, the Q - learning algorithm is used for dynamic scheduling of storage resources, and adaptive allocation of resources is achieved through reinforcement learning. Finally, based on the collected device data, a convolutional neural network (CNN) model is used to predict device failures and trigger early warnings before the failures occur. The optimization operations of storage devices do not involve cross - tier storage or data migration strategies.
[0034] Beneficial effects: Accurate storage requirement prediction and resource scheduling: Through prediction by deep learning models, storage resource configuration is planned in advance to avoid system overload or resource waste.
[0035] Enhanced fault early warning ability: By predicting faults through CNN, the accuracy and timeliness of fault detection can be improved, reducing the risk of data loss.
[0036] Adaptive resource scheduling: The Q - learning algorithm can adaptively adjust storage resources, optimize the configuration of storage devices, and improve the reliability and performance of the system.
[0037] Embodiment 2: Dynamic Monitoring and Early Warning Method for Storage Space with Cross - Tier Storage Optimization Added
[0038] This embodiment adds a cross - level storage optimization function on the basis of Embodiment 1. In addition to monitoring the real - time data of storage devices and using the LSTM model to predict storage requirements, the system also classifies data into hot data, cold data, and expired data based on the access frequency and importance of the data. Hot data is preferentially stored on high - performance storage devices, while cold data is stored on low - cost storage devices, and expired data is periodically cleared or archived. In addition, the system uses a linear programming algorithm to optimize the data migration strategy, thereby performing dynamic data migration between different storage levels and optimizing storage costs and performance.
[0039] Beneficial effects: Optimized storage resource allocation: Through cross - level storage optimization, the system can intelligently allocate storage resources according to data access frequency, improving storage performance and reducing costs.
[0040] More refined data storage management: Hot data, high - frequency data, and low - frequency data are stored at different levels respectively, effectively avoiding data storage redundancy and improving the overall system performance.
[0041] Reduced storage costs: By storing cold data on low - cost storage devices, the system significantly reduces storage costs while meeting performance requirements.
[0042] Embodiment 3: A method for dynamic monitoring and early warning of storage space combining self - healing mechanism and fault prediction
[0043] In this embodiment, in addition to performing storage requirement prediction and resource scheduling, the system also introduces a self - healing mechanism. When the storage device fault prediction system (a fault prediction model based on CNN) triggers a fault warning, the system not only sends an alarm message but also automatically performs data migration or redundant backup through the self - healing mechanism, migrating data from the faulty device to a healthy device to ensure the continuous operation of the system and data security. The cooperation between the self - healing mechanism and the fault prediction model can achieve an automated response before a fault occurs.
[0044] Beneficial effects: Enhanced fault tolerance: The addition of the self - healing mechanism enables the storage system to automatically perform data migration and redundant backup when a device fails, avoiding data loss caused by hardware failures.
[0045] Real - time and automated response: The system can monitor fault risks in real - time and respond automatically, reducing manual intervention after a fault occurs and shortening the fault recovery time.
[0046] Improved system reliability: Through the automated self - healing mechanism, the data center can maintain stable operation when storage devices fail, improving the overall system reliability.
[0047] Example 4: Dynamic Storage Space Monitoring and Warning Method Based on Global Resource Optimization
[0048] This example combines the storage resource scheduling and optimization between multiple data centers. Among multiple data centers, the system first monitors the status of storage devices in each data center in real time, uses deep learning and reinforcement learning algorithms to predict the storage requirements of each data center, and performs cross-data center storage resource scheduling through a global optimization algorithm (such as a global scheduling algorithm based on linear programming). This cross-data center resource optimization can dynamically adjust resources according to load and demand between different data centers, thus achieving more efficient resource utilization and cost savings.
[0049] Beneficial effects: Cross-data center resource optimization: Through global resource scheduling optimization, the system can intelligently allocate storage resources among multiple data centers, improving resource utilization and reducing overall costs.
[0050] Improve storage efficiency and scalability: The system can dynamically adjust resource allocation to meet the storage requirements of different data centers, improving the collaborative efficiency and overall scalability between data centers.
[0051] Reduce resource waste: Cross-data center scheduling and optimization effectively avoid over-allocation or idleness of storage resources, reducing system operation and maintenance costs.
[0052] Example 5: Fault Prediction and Warning Method with Fault Tolerance Mechanism
[0053] In this example, the system combines a fault tolerance mechanism. During the fault prediction and warning process, in addition to using the CNN model for storage device fault prediction, further fault tolerance processing is also performed on the prediction results. For example, when the probability of a certain device failure predicted by the system reaches a preset threshold, data is replicated in other devices through distributed storage to ensure data security. And when multiple devices fail simultaneously, the system automatically switches to backup devices to avoid affecting the stability of the entire system.
[0054] Beneficial effects: Enhanced data protection: The fault tolerance mechanism can back up or migrate data in a timely manner before a fault occurs to ensure data security.
[0055] Improve system stability: When multiple devices fail, the system can automatically switch to backup devices to maintain stable operation of the system and avoid service interruption.
[0056] Reduce losses caused by faults: Through the fault tolerance mechanism, even if a device fails, the system can maintain high availability and reduce losses caused by faults.
[0057] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data center storage space dynamic monitoring and early warning method, characterized in that: The method comprises the following steps: S1. Collect and monitor the operating status data of data center storage devices in real time; S2, based on the collected historical storage data, predict future storage needs through deep learning models; S3. Based on the storage demand prediction results, a reinforcement learning algorithm is used to dynamically schedule storage resources and automatically adjust storage device configuration to optimize the use of storage resources. S4. When a potential failure occurs in a storage device, the failure prediction model is used to provide early warning, and automatic data migration or redundant backup is performed through the self-healing mechanism; S5. Implement cross-level storage optimization based on data access frequency and importance, automatically migrate data to different storage tiers, and optimize storage costs and performance.
2. A data center storage space dynamic monitoring and early warning method according to claim 1, characterized in that: The S1 comprises the following steps: Collect data on storage device temperature, I / O load, response time, and storage capacity utilization in real time; The above monitoring data is input into the preset data storage system for subsequent analysis and prediction.
3. A data center storage space dynamic monitoring and early warning method according to claim 1, characterized in that: The S2 comprises the following steps: Use the long short-term memory network (LSTM) model to train historical storage demand data to predict data storage demand in the future; The trained LSTM model is evaluated, the optimal prediction model is selected for deployment, and the mean square error (MSE) is used as the loss function to optimize the model.
4. A data center storage space dynamic monitoring and early warning method according to claim 3, characterized in that: The optimization of the model using mean square error (MSE) as the loss function includes the following steps: Calculate the error between the actual storage demand and the model prediction demand, and optimize the weights of the LSTM network through the back-propagation algorithm; Minimize the following loss function: ; in, For actual storage needs, is the predicted value, is the sample size.
5. The method for dynamic monitoring and early warning of storage space in a data center according to claim 1, characterized in that: The S3 comprises the following steps: Set the state space, action space and reward function of the reinforcement learning model, and use the Q-learning algorithm to schedule storage resources; The state space is defined as the current usage status of the storage device, and the action space is defined as the scheduling operation of the storage resource; The reward function is defined as the optimal utilization efficiency of storage resources, and the storage resource allocation strategy is adjusted through the Q-learning algorithm.
6. A data center storage space dynamic monitoring and early warning method according to claim 5, characterized in that: The Q-learning algorithm for storage resource scheduling includes the following steps: Initialize the Q table, set the initial Q value to zero, and update the Q value based on the feedback of storage resource usage; Use the following update formula to update the Q value: ; in, is the reward value, indicating the improvement of storage resource utilization. is the discount factor, is the learning rate.
7. A data center storage space dynamic monitoring and early warning method according to claim 1, characterized in that: The S4 comprises the following steps: Collect temperature, error logs, and I / O latency information of storage devices and input them into the failure prediction model; A convolutional neural network (CNN) is used to analyze the status of storage devices and predict possible failure risks.
8. A data center storage space dynamic monitoring and early warning method according to claim 7, characterized in that: The use of a convolutional neural network (CNN) to analyze the state of a storage device includes the following steps: Extract the health status features of the storage device and input them into the CNN network for fault identification; Calculate the probability of failure and use the softmax function to output the prediction result of the failure. When the prediction probability exceeds the preset threshold, an early warning is triggered.
9. A data center storage space dynamic monitoring and early warning method according to claim 1, characterized in that: The S5 comprises the following steps: According to the access frequency and importance of data, data is divided into hot data, cold data and expired data; Hot data is stored on high-performance storage devices first, cold data is stored on low-cost storage devices, and expired data is regularly cleared or archived.
10. A data center storage space dynamic monitoring and early warning method according to claim 1, characterized in that: The S5 further comprises the following steps: Based on the data access pattern, a hybrid storage architecture is used for data storage and data is dynamically migrated between different storage tiers; Use linear programming to optimize data migration strategies to minimize storage costs and meet the maximum capacity constraints of each storage tier.
Citation Information
Patent Citations
Data center resource prediction and scheduling method and system based on machine learning
CN117453409A
Data storage space dynamic adjustment method, storage subsystem and intelligent computing platform
CN118069073A
Data interaction method, device, equipment and storage medium for multi-layer stacked memory
CN119781696A
Multivariable real-time measurement and state test method for data center
CN119782187A
A system for optimizing the allocation of computing resources
DE202024107285U1
Cited By
Storage resource dynamic allocation method and device, computer equipment and storage medium
CN120371549A
Storage resource dynamic allocation method and device, computer device and storage medium
CN120371549B
Self-adaptive storage capacity adjusting method and system for industrial-grade solid state disk
CN120596037A
Resource scheduling method and device for storage equipment, equipment and storage medium
CN120670177A
Resource scheduling method and device for storage device, equipment and storage medium
CN120670177B