A capacity management method, device and equipment of a distributed storage system and a medium
Patent Information
- Application Number
- CN202611280965.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明实施例的目的是提供一种分布式存储系统的容量管理方法、装置、设备及介质,可以解决如何对分布式存储系统的容量进行趋势预测、主动预警和影响评估,从而提高系统管理的前瞻性、智能化以及可靠性的问题
[0014]可见,本发明中,对容量性能统计数据进行采集,以得到当前容量数据采集结果;所述容量性能统计数据为分布式存储系统中对象存储守护进程层面与存储池层面的容量性能统计数据;基于当前容量数据采集结果和预测模型,进行容量的静态耗尽时间预测和未来使用情况预测,以确定当前容量预测结果;基于当前容量数据采集结果和预设哈希数据分布算法,模拟所述预设哈希数据分布算法与集群拓扑的变化对容量的影响,以确定当前算法模拟结果;基于当前容量预测结果和当前算法模拟结果确定当前容量数据采集结果对应的目标容量预警信息,并根据所述目标容量预警信息触发信息展示操作。
Smart Images

Figure CN122816552A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for capacity management of a distributed storage system. Background Technology
[0002] As distributed storage systems continue to expand in scale and become increasingly complex in architecture, the fundamental aspect of system capacity management is facing unprecedented challenges. Currently, mainstream distributed storage systems generally remain in a static, passive, and lagging mode in terms of capacity display and management, mainly reflected in the following aspects: (1) Single monitoring dimension and lack of trend insight: Existing solutions only show the real-time capacity utilization of the storage cluster, but the growth of data is not linear. Based on the current utilization rate, the administrator cannot predict the future capacity bottleneck. (2) Passive alarm mechanism: The alarms in the existing solutions are only passively triggered when the capacity reaches a preset fixed threshold, which cannot effectively cope with the scenario of drastic fluctuations in data write speed. (3) Difficulty in capacity impact assessment: The existing solutions do not consider the complex impact of cluster changes on the available capacity and data balance of the remaining cluster, which makes the administrator's planning decisions lack data support.
[0003] It is evident that how to predict the capacity trends of distributed storage systems, provide proactive early warnings, and conduct impact assessments to improve the foresight, intelligence, and reliability of system management is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of this invention is to provide a capacity management method, apparatus, device, and medium for distributed storage systems. This addresses the challenges of predicting capacity trends, providing proactive early warnings, and assessing impacts in distributed storage systems, thereby improving the foresight, intelligence, and reliability of system management. The specific solution is as follows: In a first aspect, the present invention provides a capacity management method for a distributed storage system, comprising: Capacity performance statistics are collected to obtain the current capacity data collection results; the capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system. Based on the current capacity data collection results and prediction model, the static exhaustion time and future usage of capacity are predicted to determine the current capacity prediction results. Based on the current capacity data collection results and the preset hash data distribution algorithm, the impact of the preset hash data distribution algorithm and changes in cluster topology on capacity is simulated to determine the simulation results of the current algorithm. Based on the current capacity prediction results and the current algorithm simulation results, the target capacity early warning information corresponding to the current capacity data collection results is determined, and the information display operation is triggered according to the target capacity early warning information.
[0005] Optionally, capacity performance statistics can be collected to obtain current capacity data collection results, including: The current capacity data collection results are obtained by collecting capacity performance statistics at the object storage daemon level and storage pool level in the distributed storage system through a data collection plugin. The data collection plugin is located in the manager of the distributed storage system. The current capacity data collection results include data read and write performance indicators corresponding to the object storage daemon, capacity usage information of each storage pool, and data distribution information of the storage pool on each object storage device. The data collection plugin stores the current capacity data collection results into a time-series database.
[0006] Optionally, based on the current capacity data collection results and prediction models, static exhaustion time and future usage predictions are performed, including: Based on a preset capacity threshold, the capacity usage information of each storage pool, and the data distribution information, a static exhaustion time prediction of the capacity is performed to determine the static prediction result; The future usage of capacity is predicted based on the prediction model, the capacity usage information, and the historical data write speed, to determine the capacity usage prediction result; wherein, the prediction model is a model built based on a long short-term memory network, a regularization layer, a fully connected layer, and an output layer; the historical data write speed is an indicator in the data read and write performance metrics; The target weight ratio is determined based on the capacity prediction results. Based on the target weight ratio, the static prediction result and the capacity usage prediction result are weighted and fused to determine the current capacity prediction result.
[0007] Optionally, static exhaustion time prediction of capacity is performed based on a preset capacity threshold, the capacity usage information of each storage pool, and the data distribution information, including: For any of the aforementioned storage pools, the data write speed of each placement group on the current storage pool is determined based on the corresponding capacity usage information and data distribution information; Based on the data write speed, the capacity usage information, and the data distribution information, predict the time it will take for the remaining capacity of the current storage pool to reach a first preset capacity threshold, so as to determine the first static prediction result; Based on the data write speed, the capacity usage information, and the data distribution information, predict the time it will take for the remaining capacity of the current storage pool to reach the second preset capacity threshold, so as to determine the second static prediction result; The static prediction result corresponding to the current storage pool is determined based on the first static prediction result and the second static prediction result.
[0008] Optionally, based on the prediction model, the capacity usage information, and historical data write speed, future capacity usage is predicted to determine the capacity usage prediction result, including: For any of the aforementioned storage pools, historical capacity information is obtained from the corresponding capacity usage information, and historical data write speed is obtained from the corresponding data read / write performance metrics. The historical capacity information and the historical data write speed are standardized to determine the processed historical capacity and the processed historical write speed. The processed historical capacity and the processed historical write speed are processed using a sliding window method to obtain time series data; The presence of missing values in the time series data is detected to determine the sequence detection result. If the sequence detection result indicates the presence of missing values, the time series data is processed based on the sequence detection result, forward imputation algorithm, and / or seasonal interpolation algorithm to determine the processed sequence data. The processed sequence data is input into the prediction model to predict future capacity usage and dynamic exhaustion time, so as to determine the current capacity usage prediction result of the storage pool.
[0009] Optionally, based on the current capacity data collection results and the preset hash data distribution algorithm, the impact of the preset hash data distribution algorithm and changes in cluster topology on capacity is simulated to determine the simulation results of the current algorithm, including: For any fault domain in the distributed storage system, record the amount of data corresponding to each placement group in the storage pool within the current fault domain based on the current capacity data acquisition results; Based on the current capacity data collection results, record the distribution information of the storage pool in the current fault domain on each object storage device. Based on the data volume, the placement group distribution information, the preset hash data distribution algorithm, and the simulated faulty node or simulated faulty hard disk, the impact of the preset hash data distribution algorithm and the change in cluster topology on the capacity under fault conditions is simulated to determine the fault simulation results; the fault simulation results include the difference in placement group distribution before and after the fault. The amount of data to be migrated in each of the object storage devices is determined based on the distribution difference of the placement groups and the amount of data. Based on the amount of data to be migrated and the current capacity data collection results, determine the target water level information of each object storage device in the current fault domain after the fault. Based on the target water level information, the target placement group distribution information of the storage pool in the current fault domain after the fault is determined on each of the object storage devices.
[0010] Optionally, based on the current capacity prediction results and the current algorithm simulation results, the target capacity early warning information corresponding to the current capacity data collection results is determined, and an information display operation is triggered according to the target capacity early warning information, including: Based on the preset early warning mechanism, the current capacity prediction results, and the current algorithm simulation results, the target capacity early warning information corresponding to the current capacity data collection results is determined; the target capacity early warning information includes a capacity early warning report. The target capacity warning information, current capacity prediction results, and current algorithm simulation results are displayed through a human-computer interaction interface.
[0011] In a second aspect, the present invention provides a capacity management device for a distributed storage system, comprising: The capacity data acquisition module is used to collect capacity performance statistics to obtain the current capacity data acquisition results; the capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system. The capacity prediction module is used to predict the static exhaustion time and future usage of capacity based on the current capacity data acquisition results and prediction models, so as to determine the current capacity prediction results. The simulation analysis module is used to simulate the impact of changes in the cluster topology on the capacity based on the current capacity data collection results and the preset hash data distribution algorithm, so as to determine the simulation results of the current algorithm. The early warning module is used to determine the target capacity early warning information corresponding to the current capacity data collection results based on the current capacity prediction results and the current algorithm simulation results, and to trigger the information display operation according to the target capacity early warning information.
[0012] Thirdly, the present invention provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute computer programs to implement the steps of the aforementioned capacity management method for a distributed storage system.
[0013] Fourthly, the present invention provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the aforementioned capacity management method for a distributed storage system.
[0014] As can be seen, in this invention, capacity performance statistics are collected to obtain the current capacity data collection result; the capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system; based on the current capacity data collection result and the prediction model, the static exhaustion time and future usage of the capacity are predicted to determine the current capacity prediction result; based on the current capacity data collection result and the preset hash data distribution algorithm, the impact of the preset hash data distribution algorithm and changes in cluster topology on the capacity is simulated to determine the current algorithm simulation result; based on the current capacity prediction result and the current algorithm simulation result, the target capacity warning information corresponding to the current capacity data collection result is determined, and the information display operation is triggered according to the target capacity warning information.
[0015] As can be seen from the above technical solution, this invention first collects capacity performance statistics at the object storage daemon level and storage pool level in the distributed storage system; then, based on the current capacity data collection results and prediction model, it predicts the static exhaustion time and future usage of the capacity, obtaining the current capacity prediction result; and based on the current capacity data collection results and a preset hash data distribution algorithm, it simulates the impact of changes in the algorithm and cluster topology on the capacity, obtaining the current algorithm simulation result; finally, based on the current capacity prediction result and the current algorithm simulation result, it determines and displays the target capacity warning information corresponding to the current capacity data collection result. The beneficial effect of this invention is that it can solve the problems existing in existing related solutions, thereby improving the automation level, efficiency, foresight, and intelligence of distributed storage system operation and maintenance, and further enhancing the reliability and business continuity of the distributed storage system. Attached Figure Description
[0016] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart of a capacity management method for a distributed storage system provided by the present invention; Figure 2 A flowchart of a specific capacity management method for a distributed storage system provided by the present invention; Figure 3 A schematic diagram of the capacity management device structure of a distributed storage system provided by the present invention; Figure 4 This invention provides a structural diagram of an electronic device. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0019] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.
[0020] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] Currently, mainstream distributed storage systems generally remain in a static, passive, and lagging mode in terms of capacity display and management, mainly reflected in the following aspects: (1) Single monitoring dimension and lack of trend insight: Existing solutions only display the real-time capacity utilization rate of the storage cluster, but the growth of data is not linear, and the administrator cannot predict the future capacity bottleneck based on the current utilization rate alone; (2) Passive alarm mechanism: Alarms in existing solutions are only passively triggered when the capacity reaches a preset fixed threshold, which cannot effectively cope with scenarios where the data write speed fluctuates drastically; (3) Difficulty in capacity impact assessment: Existing solutions do not consider the complex impact of cluster changes on the available capacity and data balance of the remaining cluster, which makes the administrator's planning decisions lack data support. To this end, the present invention provides a capacity management solution for distributed storage systems, which can solve the problems existing in existing solutions, thereby improving the automation level, efficiency, foresight and intelligence of distributed storage system operation and maintenance, and further improving the reliability and business continuity of distributed storage systems.
[0022] See Figure 1 As shown in the figure, an embodiment of the present invention discloses a capacity management method for a distributed storage system, comprising: Step S11: Collect capacity performance statistics to obtain the current capacity data collection results; the capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system.
[0023] In this embodiment, firstly, the capacity, read / write performance metrics, historical data, and data distribution of storage pools on each OSD (Object-based Storage Device) of the storage cluster are collected in real time. Specifically, a data collection plugin collects capacity performance statistics at the object storage daemon level and storage pool level in the distributed storage system to obtain the current capacity data collection result. The data collection plugin is located within the manager of the distributed storage system. The current capacity data collection result includes data read / write performance metrics corresponding to the object storage daemon, capacity usage information for each storage pool, and data distribution information of the storage pools on each object storage device. The data collection plugin then stores the current capacity data collection result in a time-series database. It is understood that a storage pool is a logical storage unit in the distributed storage system.
[0024] It's important to understand that regarding data collection, the MGR (Manager, Monitor) process aggregates capacity and performance statistics from all OSDs (Object Storage Daemon, the process responsible for data read / write in a distributed storage system) and storage pools in its memory. However, the MGR itself doesn't have a database. Therefore, this embodiment adds a plugin to write the data into the database for use by subsequent predictive models. The MGR process is the process responsible for capacity statistics in the distributed storage system. In other words, this embodiment uses the MGR module in the distributed storage system for data collection. A data collection plugin is added to the MGR module to collect real-time monitoring data, including the current used and remaining capacity of each storage unit, write speed, historical write speed data, and the data distribution of the storage pool on each object storage device. This collected data is then stored in a time-series database.
[0025] Step S12: Based on the current capacity data collection results and prediction model, perform static capacity depletion time prediction and future usage prediction to determine the current capacity prediction results.
[0026] In this embodiment, combined with Figure 2As shown, after data collection, capacity prediction can be initiated based on the current capacity data collection results. Capacity prediction is divided into static capacity exhaustion prediction and time-series-based capacity prediction. Specifically: static exhaustion time prediction is performed based on a preset capacity threshold, the capacity usage information of each storage pool, and the data distribution information to determine the static prediction result; future capacity usage is predicted based on a prediction model, the capacity usage information, and historical data write speed to determine the capacity usage prediction result. The prediction model is a model constructed based on a Long Short-Term Memory network, a regularization layer, a fully connected layer, and an output layer; the historical data write speed is an indicator in the data read / write performance metrics. A target weight ratio is determined based on the capacity usage prediction result; based on the target weight ratio, the static prediction result and the capacity usage prediction result are weighted and fused to determine the current capacity prediction result. In other words, this embodiment uses current data for static capacity exhaustion prediction and historical data write speed for time-series-based capacity prediction.
[0027] Furthermore, regarding the prediction based on static capacity exhaustion, in this embodiment: for any of the storage pools, the data write speed of each placement group on the current storage pool is determined based on the corresponding capacity usage information and the data distribution information; based on the data write speed, the capacity usage information, and the data distribution information, the time required for the remaining capacity of the current storage pool to reach a first preset capacity threshold is predicted to determine a first static prediction result; based on the data write speed, the capacity usage information, and the data distribution information, the time required for the remaining capacity of the current storage pool to reach a second preset capacity threshold is predicted to determine a second static prediction result; based on the first static prediction result and the second static prediction result, the static prediction result corresponding to the current storage pool is determined.
[0028] It's important to understand the formula for static prediction: t = (w / pool_avail) / v, where w represents the amount of space remaining when the storage pool reaches the nearfull threshold (in bytes); pool_avail represents the amount of space remaining when the storage pool reaches the full threshold (in bytes); v is the total amount of data changed within 40 seconds across each PG (Placement Group) in the storage pool / 40 seconds, representing the amount of data written per second; t represents the time required for the storage pool to reach either the nearfull or full threshold. By default, a warning is given when the predicted nearfull t of the storage pool is less than a certain time length (configured or updated based on actual needs), and a warning is given when the predicted full t of the storage pool is less than a certain time length (configured or updated based on actual needs).
[0029] Thus, this embodiment proposes a storage pool capacity prediction scheme based on current data. This scheme can provide early warning of insufficient storage pool capacity and indicate the time when the storage pool will be alarmed, so as to avoid the problem of business being suspended due to the storage pool being full.
[0030] Meanwhile, regarding time-series-based capacity prediction, in this embodiment: for any storage pool, historical capacity information is obtained from the corresponding capacity usage information, and historical data write speed is obtained from the corresponding data read / write performance indicators. The historical capacity information and the historical data write speed are standardized to determine the processed historical capacity and processed historical write speed. The processed historical capacity and the processed historical write speed are processed using a sliding window method to obtain time-series data. Missing values are detected in the time-series data to determine the sequence detection result. If the sequence detection result indicates the presence of missing values, the time-series data is processed based on the sequence detection result, a forward filling algorithm, and / or a seasonal interpolation algorithm to determine the processed sequence data. The processed sequence data is input into the prediction model to predict future capacity usage and dynamic exhaustion time points to determine the current storage pool's capacity usage prediction result.
[0031] It is important to understand that regarding dynamic prediction, in this embodiment, based on historical data of write speed and capacity usage data, the future capacity usage is predicted using the LSTM (Long Short-Term Memory) time series prediction algorithm, and the exhaustion time prediction is dynamically adjusted. The above dynamic prediction is implemented by a prediction model, with the following specific steps: 1) Standardizing the historical capacity and write speed data of the storage pool to eliminate the influence of unit dimensions; 2) Constructing time series samples and generating a training dataset using a sliding window method; 3) Addressing the data missing problem by using forward imputation and seasonal interpolation methods to ensure data continuity; 4) The LSTM layer in the prediction model: a two-layer LSTM network with 128 neurons per layer, using the tanh activation function to process the samples; 5) The Dropout layer (i.e., regularization layer) in the prediction model: a dropout rate of 0.2 to prevent overfitting; 6) The fully connected layer in the prediction model: mapping the LSTM output to the predicted value; 7) The output layer in the prediction model: the capacity prediction value for the next several days. Furthermore, in this embodiment, the storage pool capacity is dynamically predicted based on the amount of new data added each day in the near future, thereby dynamically adjusting the prediction of the capacity exhaustion time.
[0032] Thus, this embodiment proposes a storage pool capacity prediction scheme based on historical data time series. This scheme uses a long short-term memory network to predict future capacity usage and dynamically adjusts the exhaustion time prediction, which can effectively avoid the problem of the storage pool being full and causing services to be suspended due to periodic surges in business.
[0033] Furthermore, after completing the static capacity depletion prediction and the time series-based capacity prediction, the obtained static prediction results and capacity are weighted and fused together. The weight ratio is dynamically adjusted according to the success probability of the prediction results output by the prediction model. The weight ratio formula is w=s g–b. Where s is the probability of a successful prediction by the prediction model; g is the growth coefficient, defaulting to 0.8; a larger g indicates greater confidence in the prediction model's output, suitable for use cases with strong periodicity; b is the skepticism coefficient, a larger b indicates less confidence in the prediction model. Based on capacity prediction t=w t1 + (1-w) t2, where t1 is the prediction time of the prediction model and t2 is the static prediction time.
[0034] It should be noted that regarding the dynamic adjustment of the weight ratio, since the prediction results output by the prediction model in this embodiment may have different values for the remaining storage pool capacity and success probability on a future day, the success probability on the day when the storage pool capacity alarm value is reached is taken. If the success probability is higher than 50%, it is adopted. The weight ratio can be obtained by fixing the coefficient, and the weight ratio is between 0.1 and 0.5.
[0035] Step S13: Based on the current capacity data collection results and the preset hash data distribution algorithm, simulate the impact of the preset hash data distribution algorithm and the changes in cluster topology on the capacity, so as to determine the simulation results of the current algorithm.
[0036] In this embodiment, after obtaining the current capacity data acquisition results, in addition to using the current capacity data acquisition results to predict the capacity, the current capacity data acquisition results are also used in conjunction with a preset hash data distribution algorithm to simulate the impact of algorithm changes and cluster topology changes on the capacity. That is: for any fault domain in the distributed storage system, the data volume corresponding to each placement group in the storage pool in the current fault domain is recorded based on the current capacity data acquisition results; the placement group distribution information of the storage pool in the current fault domain on each object storage device is recorded based on the current capacity data acquisition results; based on the data volume, the placement group distribution information, the preset hash data distribution algorithm, and the simulation algorithm, the data volume is recorded based on the current capacity data acquisition results; the data volume, the placement group distribution information, the preset hash data distribution algorithm, and the simulation algorithm are used to simulate the impact of algorithm changes and cluster topology changes on the capacity. A simulated faulty node or hard drive is used to simulate the impact of the preset hash data distribution algorithm and changes in cluster topology on capacity under fault conditions, in order to determine the fault simulation results. The fault simulation results include the difference in placement group distribution before and after the fault. Based on the difference in placement group distribution and the data volume, the amount of data to be migrated from each of the object storage devices is determined. Based on the amount of data to be migrated and the current capacity data collection results, the target water level information for each of the object storage devices within the current fault domain after the fault is determined. Based on the target water level information, the target placement group distribution information of the storage pools on each of the object storage devices within the current fault domain after the fault is determined. It is understood that the preset hash data distribution algorithm can be a CRUSH rule (Controlled Replication Under Scalable Hashing), which is a controllable and scalable hash data distribution algorithm used to determine the physical location of data objects in the storage cluster. A fault domain refers to the physical area in a distributed storage system where faults may occur simultaneously; multiple storage pools can exist within a fault domain.
[0037] It is important to understand that, in combination Figure 2As shown, the process of simulating the impact to achieve prediction is actually a tool that allows administrators to input planned changes, including but not limited to scaling down, fault simulation, and adjusting the rules of the preset hash data distribution algorithm. The aforementioned fault simulation type of change is mainly introduced. Other similar types of changes can also refer to the above logic. Based on the current preset hash data distribution algorithm and data distribution, the tool simulates the data migration and redistribution process after the change, and calculates the effective capacity and remaining available space of each storage pool in the distributed cluster after the change. The specific implementation steps are as follows: (1) Before simulating the fault, calculate and record the amount of data pg_c represented by PG in each storage pool under the fault domain according to the fault domain, and record the PG distribution of each storage pool on each OSD in each fault domain. (2) Output the node or hard disk to be faulted, and output the PG distribution of each storage pool on each OSD in each fault domain after the fault according to the CRUSH algorithm. (3) Calculate the difference d between the PG distribution of each storage pool on each OSD in each fault domain before and after the fault, and use pc_c d. The amount of data p_data that needs to be migrated after a failure on each OSD is obtained. (4) Combine the current water level of the OSDs in the storage pool with the p_data that needs to be migrated after a failure, calculate the final water level information of each OSD in the fault domain, and reflect the capacity of the storage pool by the water level information of the OSDs.
[0038] Thus, this embodiment proposes a tool to simulate CRUSH rule changes, allowing administrators to simulate the data migration and redistribution process after changes based on faults and scaling down, calculate the remaining available space of each storage pool in the distributed cluster after the change, and provide corresponding analysis reports.
[0039] Step S14: Determine the target capacity warning information corresponding to the current capacity data collection result based on the current capacity prediction result and the current algorithm simulation result, and trigger the information display operation according to the target capacity warning information.
[0040] In this embodiment, after completing capacity prediction and algorithm simulation, the obtained current capacity prediction results and current algorithm simulation results are used to analyze whether an early warning is needed, and if so, generate early warning-related information to be displayed. That is, based on the preset early warning mechanism, the current capacity prediction results, and the current algorithm simulation results, the target capacity early warning information corresponding to the current capacity data collection results is determined; the target capacity early warning information includes a capacity early warning report; and the target capacity early warning information, the current capacity prediction results, and the current algorithm simulation results are displayed on a human-computer interaction interface.
[0041] It's important to understand that this embodiment employs a multi-level warning mechanism for capacity alerts. It combines current capacity prediction results and current algorithm simulation results to analyze areas requiring warnings and determine their severity. Higher severity results in higher priority, and corresponding suggested actions are generated, culminating in a comprehensive warning report. Subsequently, the warning report, along with the current capacity prediction and algorithm simulation results, are displayed to the administrator through visualization, enabling timely action to address potential risks.
[0042] In summary, this embodiment aims to provide a distributed storage management scheme based on predictive capacity to address the problems of passive capacity management, lack of trend prediction, and impact assessment in existing technologies. This scheme achieves proactive capacity management and intelligent planning by introducing time series forecasting and CRUSH rule change simulation. Compared to existing related schemes, this scheme has the following beneficial effects: (1) Achieve a fundamental shift in operation and maintenance mode from passive response to proactive prevention. Through intelligent prediction algorithms based on time-series data, expandability plans can be planned in advance to avoid the risk of business interruption, while significantly reducing hardware procurement costs caused by over-configuration.
[0043] (2) Significantly enhances business continuity and system reliability. The CRUSH rule simulation module enables administrators to accurately assess the impact of planned changes and potential failures on the cluster, identify capacity risks in advance, and optimize data distribution strategies. Multi-level early warning mechanisms and visualization ensure that problems are detected and dealt with early, effectively preventing system failures caused by capacity exhaustion and providing more stable storage service guarantees for critical businesses.
[0044] (3) Significantly improves the level and efficiency of operation and maintenance automation. The system frees administrators from tedious manual inspections through real-time monitoring, intelligent analysis and automated alarms. It greatly reduces the complexity of operation and maintenance and management costs, and realizes refined and intelligent management of storage resources.
[0045] Therefore, in this embodiment of the invention, firstly, capacity performance statistics at the object storage daemon level and storage pool level of the distributed storage system are collected; then, based on the current capacity data collection results and the prediction model, the static exhaustion time and future usage of the capacity are predicted to obtain the current capacity prediction result; and based on the current capacity data collection results and the preset hash data distribution algorithm, the impact of changes in the algorithm and cluster topology on the capacity is simulated to obtain the current algorithm simulation result; finally, based on the current capacity prediction result and the current algorithm simulation result, the target capacity warning information corresponding to the current capacity data collection result is determined and displayed. The beneficial effect of this invention is that it can solve the problems existing in existing related solutions, thereby improving the automation level, efficiency, foresight, and intelligence of distributed storage system operation and maintenance, and further enhancing the reliability and business continuity of the distributed storage system.
[0046] As a preferred embodiment, considering that in large-scale clusters (such as thousands of OSDs and millions of PGs), fully simulating CRUSH rule changes requires traversing the mapping relationships of all PGs, resulting in a huge computational load. Real-time execution could consume the central processing unit / memory of the management node and even affect monitoring response speed. To avoid these adverse effects, this embodiment can also adopt the following simulation performance optimization and asynchronous processing measures: 1) The step of simulating the impact of changes in the cluster topology on capacity based on the current capacity data collection results and the preset hash data distribution algorithm can be placed in a background queue for execution, with the front end immediately returning a prompt such as "Simulation task submitted, expected to complete in 30 seconds." 2) For multiple simulations of the same topology and rules, cached results (key-value pairs based on topology hash + rule hash) can be used to avoid redundant calculations. 3) Incremental simulation updates can also be supported: after cluster changes (such as adding 10 OSDs), only the PG mappings involving these OSDs are recalculated, rather than the entire mapping.
[0047] See Figure 3 As shown, this embodiment of the invention also discloses a capacity management device for a distributed storage system, comprising: The capacity data acquisition module 11 is used to collect capacity performance statistics to obtain the current capacity data acquisition result; the capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system. The capacity prediction module 12 is used to predict the static exhaustion time and future usage of capacity based on the current capacity data acquisition results and prediction model, so as to determine the current capacity prediction result. The simulation analysis module 13 is used to simulate the impact of changes in the cluster topology on the capacity based on the current capacity data collection results and the preset hash data distribution algorithm, so as to determine the simulation results of the current algorithm. The early warning module 14 is used to determine the target capacity early warning information corresponding to the current capacity data collection result based on the current capacity prediction result and the current algorithm simulation result, and to trigger the information display operation according to the target capacity early warning information.
[0048] Therefore, in this embodiment of the invention, firstly, capacity performance statistics at the object storage daemon level and storage pool level of the distributed storage system are collected; then, based on the current capacity data collection results and the prediction model, the static exhaustion time and future usage of the capacity are predicted to obtain the current capacity prediction result; and based on the current capacity data collection results and the preset hash data distribution algorithm, the impact of changes in the algorithm and cluster topology on the capacity is simulated to obtain the current algorithm simulation result; finally, based on the current capacity prediction result and the current algorithm simulation result, the target capacity warning information corresponding to the current capacity data collection result is determined and displayed. The beneficial effect of this invention is that it can solve the problems existing in existing related solutions, thereby improving the automation level, efficiency, foresight, and intelligence of distributed storage system operation and maintenance, and further enhancing the reliability and business continuity of the distributed storage system.
[0049] For more detailed information on the working process of each of the above modules, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0050] Furthermore, embodiments of the present invention also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of the invention. Specifically, the electronic device may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the capacity management method of the distributed storage system disclosed in any of the foregoing embodiments. Furthermore, the electronic device in this embodiment may specifically be an electronic computer.
[0051] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this invention, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0052] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0053] The operating system 221 is used to manage and control the various hardware devices on the electronic device and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the capacity management method of the distributed storage system executed by the electronic device as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.
[0054] Furthermore, the present invention also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the capacity management method of the aforementioned distributed storage system. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0055] Furthermore, this application also discloses a computer program product, including a computer program / instructions; wherein, when the computer program / instructions are executed by a processor, they implement the aforementioned capacity management method for a distributed storage system. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0056] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0057] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0058] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0059] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0060] The technical solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A capacity management method for a distributed storage system, characterized in that, include: Collect capacity performance statistics to obtain the current capacity data collection results; The capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system. Based on the current capacity data collection results and prediction model, the static exhaustion time and future usage of capacity are predicted to determine the current capacity prediction results. Based on the current capacity data collection results and the preset hash data distribution algorithm, the impact of the preset hash data distribution algorithm and changes in cluster topology on capacity is simulated to determine the simulation results of the current algorithm. Based on the current capacity prediction results and the current algorithm simulation results, the target capacity early warning information corresponding to the current capacity data collection results is determined, and the information display operation is triggered according to the target capacity early warning information.
2. The capacity management method for a distributed storage system according to claim 1, characterized in that, The collection of capacity performance statistics to obtain the current capacity data collection results includes: The current capacity data collection results are obtained by collecting capacity performance statistics at the object storage daemon level and storage pool level in the distributed storage system through a data collection plugin. The data collection plugin is located in the manager of the distributed storage system. The current capacity data collection results include data read and write performance indicators corresponding to the object storage daemon, capacity usage information of each storage pool, and data distribution information of the storage pool on each object storage device. The data collection plugin stores the current capacity data collection results into a time-series database.
3. The capacity management method for a distributed storage system according to claim 2, characterized in that, The prediction of static capacity depletion time and future usage based on current capacity data collection results and prediction models includes: Based on a preset capacity threshold, the capacity usage information of each storage pool, and the data distribution information, a static exhaustion time prediction of the capacity is performed to determine the static prediction result; The future usage of capacity is predicted based on the prediction model, the capacity usage information, and the historical data write speed, to determine the capacity usage prediction result; wherein, the prediction model is a model built based on a long short-term memory network, a regularization layer, a fully connected layer, and an output layer; the historical data write speed is an indicator in the data read and write performance metrics; The target weight ratio is determined based on the capacity prediction results. Based on the target weight ratio, the static prediction result and the capacity usage prediction result are weighted and fused to determine the current capacity prediction result.
4. The capacity management method for a distributed storage system according to claim 3, characterized in that, The static exhaustion time prediction based on a preset capacity threshold, the capacity usage information of each storage pool, and the data distribution information includes: For any of the aforementioned storage pools, the data write speed of each placement group on the current storage pool is determined based on the corresponding capacity usage information and data distribution information; Based on the data write speed, the capacity usage information, and the data distribution information, predict the time it will take for the remaining capacity of the current storage pool to reach a first preset capacity threshold, so as to determine the first static prediction result; Based on the data write speed, the capacity usage information, and the data distribution information, predict the time it will take for the remaining capacity of the current storage pool to reach the second preset capacity threshold, so as to determine the second static prediction result; The static prediction result corresponding to the current storage pool is determined based on the first static prediction result and the second static prediction result.
5. The capacity management method for a distributed storage system according to claim 3, characterized in that, The prediction of future capacity usage based on the prediction model, the capacity usage information, and historical data write speed, to determine the capacity usage prediction result, includes: For any of the aforementioned storage pools, historical capacity information is obtained from the corresponding capacity usage information, and historical data write speed is obtained from the corresponding data read / write performance metrics. The historical capacity information and the historical data write speed are standardized to determine the processed historical capacity and the processed historical write speed. The processed historical capacity and the processed historical write speed are processed using a sliding window method to obtain time series data; The presence of missing values in the time series data is detected to determine the sequence detection result. If the sequence detection result indicates the presence of missing values, the time series data is processed based on the sequence detection result, forward imputation algorithm, and / or seasonal interpolation algorithm to determine the processed sequence data. The processed sequence data is input into the prediction model to predict future capacity usage and dynamic exhaustion time, so as to determine the current capacity usage prediction result of the storage pool.
6. The capacity management method for a distributed storage system according to claim 1, characterized in that, The process involves simulating the impact of changes in the cluster topology on capacity based on current capacity data collection results and a preset hash data distribution algorithm, in order to determine the simulation results of the current algorithm, including: For any fault domain in the distributed storage system, record the amount of data corresponding to each placement group in the storage pool within the current fault domain based on the current capacity data acquisition results; Based on the current capacity data collection results, record the distribution information of the storage pool in the current fault domain on each object storage device. Based on the data volume, the placement group distribution information, the preset hash data distribution algorithm, and the simulated faulty node or simulated faulty hard disk, the impact of the preset hash data distribution algorithm and the change in cluster topology on the capacity under fault conditions is simulated to determine the fault simulation results; the fault simulation results include the difference in placement group distribution before and after the fault. The amount of data to be migrated in each of the object storage devices is determined based on the distribution difference of the placement groups and the amount of data. Based on the amount of data to be migrated and the current capacity data collection results, determine the target water level information of each object storage device in the current fault domain after the fault. Based on the target water level information, the target placement group distribution information of the storage pool in the current fault domain after the fault is determined on each of the object storage devices.
7. The capacity management method for a distributed storage system according to any one of claims 1 to 6, characterized in that, The process of determining the target capacity early warning information corresponding to the current capacity data collection results based on the current capacity prediction results and the current algorithm simulation results, and triggering an information display operation based on the target capacity early warning information, includes: Based on the preset early warning mechanism, the current capacity prediction results, and the current algorithm simulation results, the target capacity early warning information corresponding to the current capacity data collection results is determined; the target capacity early warning information includes a capacity early warning report. The target capacity warning information, current capacity prediction results, and current algorithm simulation results are displayed through a human-computer interaction interface.
8. A capacity management device for a distributed storage system, characterized in that, include: The capacity data acquisition module is used to collect capacity performance statistics to obtain the current capacity data acquisition results; The capacity performance statistics are the capacity performance statistics at the object storage daemon level and the storage pool level in the distributed storage system. The capacity prediction module is used to predict the static exhaustion time and future usage of capacity based on the current capacity data acquisition results and prediction models, so as to determine the current capacity prediction results. The simulation analysis module is used to simulate the impact of changes in the cluster topology on the capacity based on the current capacity data collection results and the preset hash data distribution algorithm, so as to determine the simulation results of the current algorithm. The early warning module is used to determine the target capacity early warning information corresponding to the current capacity data collection results based on the current capacity prediction results and the current algorithm simulation results, and to trigger the information display operation according to the target capacity early warning information.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the capacity management method for the distributed storage system as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the capacity management method for the distributed storage system as described in any one of claims 1 to 7.