Storage performance prediction and elastic capacity expansion and contraction method for hyper-converged architecture
By constructing a storage performance prediction model based on historical scheduling logs and analyzing the performance anomaly propagation chain, combined with a supply-demand collaborative scaling strategy, the challenges of storage performance prediction and elastic scaling in hyperconverged architectures were solved, achieving efficient utilization of storage resources and business stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies in hyperconverged architectures struggle to fully utilize historical scheduling log information to build accurate storage performance prediction models. They also fail to accurately analyze load fluctuation characteristics and hardware health characteristics, making it difficult to quickly locate performance anomalies and implement effective elastic scaling strategies. This impacts the rational utilization of storage resources and business stability.
A key storage performance indicator prediction model is constructed, which takes the business load labels and historical indicator change trends in historical scheduling logs as input. Based on load fluctuation characteristics, hardware health characteristics, and topology correlation characteristics, the performance anomaly propagation chain is determined. Combined with the supply and demand coordinated expansion scale and cross-module collaborative execution mechanism, an elastic expansion strategy is selected and implemented.
It enables accurate prediction of future storage performance, allows for advance planning of resource allocation, avoids resource waste, ensures stable and efficient operation of the storage system, and meets business needs.
Smart Images

Figure CN121807230A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data storage, in particular to a storage performance prediction and elastic scaling method for a hyper-converged architecture. BACKGROUND
[0002] At present, in the digital era, the amount of data is growing explosively, which puts high demands on the performance and scalability of storage systems. Hyper-converged architecture, as an innovative IT infrastructure, integrates computing, storage, and network resources in depth, with the advantages of high efficiency, flexibility, and easy management, and is widely used in data centers, cloud computing, and other fields. It breaks the traditional architecture where storage, computing, and other resources are independent of each other, and realizes unified management and deployment of resources, providing more convenient and efficient IT solutions for enterprises. However, with the continuous development of business and the continuous increase of data volume, the storage under the hyper-converged architecture faces many challenges, therefore, the research on the storage performance prediction and elastic scaling method for the hyper-converged architecture is of great significance. This helps to ensure the continuity and stability of business. In the long run, with the further development of cloud computing, big data, artificial intelligence, and other technologies, the application of hyper-converged architecture will be more widespread, and the demand for such storage performance prediction and elastic scaling methods will continue to grow, with broad development prospects.
[0003] The existing technology has many shortcomings in the prediction of storage performance and elastic scaling of hyper-converged architecture. It is difficult to fully utilize the business load tags, historical index change trends, and other information in the historical scheduling log to build an accurate prediction model, resulting in the inability to know in advance the changes in storage performance in future periods. In addition, the analysis capability of the comprehensive information of the load fluctuation characteristics of each sub-cluster in the hyper-converged architecture, the hardware health characteristics of the storage nodes, and the topology correlation characteristics is insufficient, which cannot accurately determine the performance anomaly propagation chain, making it difficult to quickly locate the root cause and take effective measures when facing performance problems. Finally, in the formulation of elastic scaling strategy, it is unable to select the appropriate scaling strategy based on the principle of supply-demand coordination combined with the performance anomaly propagation chain, and lack of cross-module collaborative execution mechanism to ensure the effective implementation of the strategy. This leads to the inability to achieve precise and efficient elastic scaling when the storage resources need to be adjusted, affecting the normal operation of business and the rational use of resources.
[0004] Therefore, the present application proposes a storage performance prediction and elastic scaling method for a hyper-converged architecture. SUMMARY
[0005] The application provides a storage performance prediction and elastic scaling method for a hyper-converged architecture, which can accurately predict storage performance, plan resources in advance, avoid the impact of insufficient storage performance on business operation, and prevent waste caused by excessive resource configuration. The elastic scaling can flexibly adjust storage resources according to actual business needs, improve resource utilization, and reduce operating costs, thereby helping to ensure the continuity and stability of the business.
[0006] The application provides a storage performance prediction and elastic scaling method for a hyper-converged architecture, which includes: A key storage performance index prediction model is constructed, which takes business load labels and historical index change trends in historical scheduling logs as inputs, and takes key storage performance indexes of all storage nodes of the hyper-converged architecture in a preset future period as outputs. The key storage performance indexes of all storage nodes in the preset future period are predicted based on the key storage performance index prediction model. Based on the load fluctuation characteristics of all sub-clusters of the hyper-converged architecture, the hardware health characteristics, the topology correlation characteristics of all storage nodes, the key storage performance indexes in the preset future period, and the storage performance threshold range adapted to the business demand, a performance anomaly propagation chain is determined. An elastic scaling strategy is selected based on the supply-demand collaborative scaling scale and the performance anomaly propagation chain, and the elastic scaling strategy is executed based on a cross-module collaborative execution mechanism.
[0007] Preferably, the key storage performance index prediction model is constructed, which takes business load labels and historical index change trends in historical scheduling logs as inputs, and takes key storage performance indexes of all storage nodes of the hyper-converged architecture as outputs, and includes: The key storage performance indexes of the storage nodes of the hyper-converged architecture at multiple historical moments are collected, and the historical scheduling logs of the data center resource scheduling software in the historical period corresponding to each historical moment are obtained, wherein the historical scheduling logs cover the historical index change trends and business load labels of all key storage performance indexes in the corresponding historical period. Based on the key storage performance indexes of the storage nodes of the hyper-converged architecture at multiple historical moments and the historical scheduling logs of the data center resource scheduling software in the historical period corresponding to each historical moment, a random forest model is used as an initial model to construct a key storage performance index prediction model, which takes business load labels and historical index change trends in historical scheduling logs as inputs, and takes key storage performance indexes of all storage nodes of the hyper-converged architecture in a preset future period as outputs.
[0008] Preferably, the key storage performance indexes of all storage nodes in the preset future period are predicted based on the key storage performance index prediction model, and include: The service load tags and historical index change trends of all storage nodes of the hyper-converged architecture in the latest historical period are input into a key storage performance index prediction model to obtain the key storage performance indexes of all storage nodes in a preset future period.
[0009] Preferably, the partition method of all sub-clusters of the hyper-converged architecture comprises: The rack number to which all storage nodes belong, the network link belonging to the computing node, and the copy distribution rule of the distributed storage pool are extracted as hardware association features in the physical hardware topology of the hyper-converged architecture. All storage nodes are divided into multiple initial sub-clusters based on the hardware association features of all storage nodes. The historical scheduling log of the data center resource scheduling software in the latest historical period is called, and the load fluctuation features of each initial sub-cluster are extracted in the historical scheduling log, wherein the load fluctuation features include the IOPS fluctuation variance, the service load type proportion, and the storage fragmentation rate change trend of all storage nodes in the initial sub-cluster. Based on a preset clustering algorithm, the initial sub-clusters are divided and merged again based on the clustering target that the load fluctuation variance does not exceed the preset fluctuation variance threshold and the same type service proportion is not less than the preset proportion threshold, to obtain multiple sub-clusters, and based on a preset sub-cluster dynamic adjustment mechanism, all sub-clusters are dynamically adjusted.
[0010] Preferably, based on the load fluctuation features of all sub-clusters of the hyper-converged architecture and the hardware health features, topology association features, key storage performance indexes in a preset future period, and storage performance threshold range adapted to the service demand of all storage nodes, a performance anomaly propagation chain is determined, comprising: A topology association matrix is constructed based on the topology association features of all storage nodes. The storage performance threshold range adapted to the service demand is set. Based on the key storage performance indexes of all storage nodes in a preset future period, the performance index change slope, performance drop rate, and fluctuation sequence of each storage node in the future period are determined, and based on the storage performance threshold range, all storage performance abnormal nodes are screened out among all storage nodes. Based on the performance index change slope, performance drop rate, and fluctuation sequence of all storage nodes in the future period, the performance similarity between different storage nodes is calculated. Based on the load fluctuation features of all sub-clusters of the hyper-converged architecture, the hardware health features of all storage nodes, the topology association matrix, all storage performance abnormal nodes, and the performance similarity between different storage nodes, the performance anomaly propagation chain is determined.
[0011] Preferably, based on the load fluctuation characteristics of all sub-clusters of the hyper-converged architecture, the hardware health characteristics of all storage nodes, the topology correlation matrix, all storage performance abnormal nodes and the performance similarity between different storage nodes, the performance abnormal propagation chain is determined, including: Screening all health risk nodes among all storage nodes based on the hardware health characteristics of all storage nodes; Identifying a direct correlation node group based on the topology correlation matrix of all storage nodes, and constructing an initial node correlation graph based on the direct correlation node group; Taking all storage performance abnormal nodes and health risk nodes as key attention nodes, and determining all indirect correlation key node groups among all key attention nodes based on the initial node correlation graph; Based on the performance similarity between different storage nodes and the load fluctuation characteristics of all sub-clusters of the hyper-converged architecture, the comprehensive correlation strength of each indirect correlation key node group is calculated; Based on the comprehensive correlation strength of each indirect correlation key node group and the indirect correlation path between the two key attention nodes contained in the corresponding indirect correlation key node group, the reliability of each minimum unit path in the corresponding indirect correlation key node group is determined. The sum of the reliability of each minimum unit path in all indirect correlation key node groups in the initial node correlation graph is taken as the comprehensive reliability of each minimum unit path; All minimum unit paths in the initial node correlation graph with a comprehensive reliability less than a reliability threshold are removed to obtain an effective node correlation graph; Among all the hypothetical abnormal propagation paths in the effective node correlation graph, all the hypothetical abnormal propagation paths with a maximum coincidence degree not less than a coincidence degree threshold with all historical abnormal propagation paths are selected as performance abnormal propagation chains.
[0012] Preferably, based on the performance similarity between different storage nodes and the load fluctuation characteristics of all sub-clusters of the hyper-converged architecture, the comprehensive correlation strength of each indirect correlation key node group is calculated, including: Based on the load fluctuation characteristics of all sub-clusters of the hyper-converged architecture, the load coordination degree between different storage nodes is calculated; Based on the performance similarity between different storage nodes and the load coordination degree, the comprehensive correlation strength of each indirect correlation key node group is calculated.
[0013] Preferably, the expansion scale determination method of supply-demand coordination includes: Analyze the basic capacity demand of demand-side business growth, business peak elasticity demand, and data redundancy demand; Based on the historical capacity change data of the hyper-converged storage pool, the actual available capacity of the supply-side storage resource and the capacity growth loss are calculated; determine the expansion scale of supply and demand coordination based on the basic capacity demand of demand-side business growth, the business peak elasticity demand, the data redundancy demand, and the actual available capacity of supply-side storage resources, the capacity growth loss rate.
[0014] Preferably, the elastic expansion strategy is selected based on the expansion scale of supply and demand coordination and the performance anomaly propagation chain, including: When the business scenario of the hyper-converged architecture is a core business scenario, then based on the scale type to which the expansion scale of supply and demand coordination belongs and the influence range of the performance anomaly propagation chain, the corresponding elastic expansion strategy is selected from the expansion strategy library under the core business scenario; When the business scenario of the hyper-converged architecture is a non-core business scenario, then based on the scale type to which the expansion scale of supply and demand coordination belongs and the influence range of the performance anomaly propagation chain, the corresponding elastic expansion strategy is selected from the expansion strategy library under the non-core business scenario.
[0015] Preferably, the elastic expansion strategy is executed based on the cross-module collaborative execution mechanism, including: Based on the cross-module collaborative execution mechanism, the elastic expansion strategy is executed in collaboration with the computing module, the network module, and the monitoring module.
[0016] The beneficial effects of the present application relative to the prior art are: by constructing a prediction model with the business load label in the historical scheduling log and the historical index change trend as input and the key storage performance indicators of all storage nodes of the hyper-converged architecture in a preset future period as output, the future storage performance indicators can be effectively predicted to provide a basis for advance planning. Based on the model, the key storage performance indicators of all storage nodes in a preset future period can be predicted, so that the storage performance change trend can be known in advance by the operation and maintenance personnel. The performance anomaly propagation chain is determined based on the load fluctuation characteristics of all subsets of the hyper-converged architecture, the hardware health characteristics of the storage nodes, the topology correlation characteristics, the predicted key storage performance indicators, and the storage performance threshold range adapted to the business demand, which helps to accurately locate the possible performance problems and the propagation path. The elastic expansion strategy is selected based on the expansion scale of supply and demand coordination and the performance anomaly propagation chain, and is executed with the help of the cross-module collaborative execution mechanism, which can achieve more reasonable elastic expansion, effectively improve the storage system performance, ensure the stable and efficient operation of the hyper-converged architecture storage system, meet the business demand for storage performance, and at the same time avoid resource waste caused by excessive expansion.
[0017] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and obtained by the structure particularly pointed out in the application file.
[0018] The technical solutions of the present application are described in further detail below with reference to the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings are used to provide further understanding of the present application, and form a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation on the present application. In the drawings: Figure 1 A flowchart of the storage performance prediction and elastic scaling method for the hyper-converged architecture in the embodiments of the present application; Figure 2 A flowchart of the key storage performance index prediction model construction and index prediction in the embodiments of the present application; Figure 3 A flowchart of the elastic scaling strategy selection in the embodiments of the present application. DETAILED DESCRIPTION
[0020] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and do not constitute a limitation on the present application.
[0021] As shown in Figure 1 , the present application provides an embodiment of a storage performance prediction and elastic scaling method for a hyper-converged architecture, comprising: constructing a key storage performance index prediction model with the business load label in the historical scheduling log and the historical index change trend as input, and the key storage performance index of all storage nodes of the hyper-converged architecture in the preset future period as output; predicting the key storage performance index of all storage nodes in the preset future period based on the key storage performance index prediction model; determining the performance anomaly propagation chain based on the load fluctuation characteristics of all sub-clusters of the hyper-converged architecture and the hardware health characteristics, topology correlation characteristics, key storage performance index in the preset future period of all storage nodes, and the storage performance threshold range adapted to the business demand; selecting an elastic scaling strategy based on the supply-demand collaborative scaling size and the performance anomaly propagation chain, and executing the elastic scaling strategy based on the cross-module collaborative execution mechanism.
[0022] In this embodiment, the historical scheduling log is the content recorded by the data center resource scheduling software at each historical time corresponding to the historical period, covering all key storage performance index change sequences and business load labels in the corresponding historical period and other information.
[0023] In this embodiment, the business load label is used to identify different business load types running in the hyper-converged architecture, such as "database read-write", "file transfer", "backup archiving", etc. With these labels, different business loads can be distinguished.
[0024] In this embodiment, the historical indicator change trend refers to the change trend of key storage performance indicators of the storage node of the hyper-converged architecture in the past period of time, such as the change of indicators such as IOPS, read-write delay, and storage fragmentation rate over time.
[0025] In this embodiment, the hyper-converged architecture is an innovative IT infrastructure that deeply integrates resources such as computing, storage, and network, breaks the state of independence of each resource in the traditional architecture, and realizes unified management and deployment of resources. In the data center and other scenarios, the hyper-converged architecture effectively supports the operation of cloud computing, big data, and other businesses due to its high efficiency, flexibility, and ease of management.
[0026] In this embodiment, the storage node is a unit that undertakes data storage tasks in the hyper-converged architecture, and a large number of storage nodes jointly build the storage system of the hyper-converged architecture, similar to each small warehouse that stores a part of data in a large warehouse.
[0027] In this embodiment, the preset future period is a future time period set in advance, which is mainly used for predicting storage performance indicators in the technical solution of the present application. The specific duration can be set according to actual needs, such as 1 hour in the future, half a day in the future, etc.
[0028] In this embodiment, the key storage performance indicator is an important parameter for measuring the storage performance of the hyper-converged architecture, including IOPS (input / output operations per second), read-write delay, storage fragmentation rate, bandwidth utilization rate, cache hit rate, etc.
[0029] In this embodiment, the key storage performance indicator prediction model is a model that takes the business load label in the historical scheduling log and the historical indicator change trend as input, and takes the key storage performance indicators of all storage nodes of the hyper-converged architecture in the preset future period as output. Through learning of historical data, the model mines the internal relationship between business load and storage performance indicators, and then predicts future storage performance. For example, taking a random forest model as an initial model, after a series of training and optimization, the model can output the predicted values of the key storage performance indicators of each storage node in the preset future period.
[0030] In this embodiment, the sub-cluster is a plurality of clusters formed by dividing the storage nodes in the hyper-converged architecture according to hardware association characteristics such as the physical rack number to which the storage nodes belong, the network link belonging to the computing nodes, the replica distribution rule of the distributed storage pool, and load characteristics such as IOPS fluctuation variance, business load type proportion, and storage fragmentation rate change trend. The storage nodes in each sub-cluster after division have certain similarity in hardware and load.
[0031] In this embodiment, the load fluctuation characteristics reflect the changing characteristics of the storage node load in each sub-cluster of the hyperconverged architecture, specifically reflected by IOPS fluctuation variance, business load type proportion, and storage fragmentation rate change trend.
[0032] In this embodiment, the storage performance threshold range adapted to business needs is a performance indicator range set according to the storage performance requirements of different businesses. Different thresholds are set for different types of businesses. For example, for businesses with high real-time requirements, read / write latency is strictly limited to a certain time, and IOPS must be maintained above a certain value. When key storage performance indicators exceed this range, it may affect the normal operation of the business, thereby triggering corresponding processing measures, such as elastic scaling.
[0033] In this embodiment, the performance anomaly propagation chain is determined based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the hardware health characteristics of all storage nodes, topological correlation characteristics, key storage performance indicators in a preset future time period, and storage performance threshold ranges adapted to business needs. This chain represents the links that may cause storage performance anomalies to propagate between storage nodes. For example, by analyzing multiple factors such as the topological connections between nodes, the degree of load coordination, and the hardware health status, the specific path that a performance anomaly may take from one node to other nodes can be determined.
[0034] In this embodiment, the scale of capacity expansion in a supply-demand coordinated manner is determined by comprehensively considering the basic capacity requirements of demand-side business growth, peak elasticity requirements, and data redundancy requirements, as well as historical capacity change data of the supply-side hyperconverged storage pool, such as actual available capacity and capacity growth losses. The expansion scale is calculated using the specific formula "Expansion Scale = (Basic Capacity Requirements + Peak Elasticity Requirements + Redundancy Requirements) ÷ (1 - Capacity Growth Loss Rate) - Current Effective Capacity," thereby achieving reasonable and accurate expansion planning.
[0035] In this embodiment, the elastic scaling strategy is based on the business scenarios of the hyperconverged architecture, categorized into core and non-core business scenarios, the scale type of the scaling operation (e.g., small, medium, large), and the impact range of the performance anomaly propagation chain. Strategies to address storage performance changes are selected from the corresponding scaling strategy library. For example, in a core business scenario, if the performance anomaly is due to excessive fragmentation and the scaling scale is small, a "virtualized storage reorganization + data redistribution" strategy might be chosen. In a non-core business scenario with large-scale scaling, a "cold and hot data tiered migration + object storage expansion" strategy might be selected. This allows for targeted scaling decisions, meeting the storage performance adjustment needs under different circumstances.
[0036] like Figure 2As shown, to provide a reliable tool for storage performance prediction, a key storage performance indicator prediction model is proposed, which takes service load labels and historical indicator change trends in historical scheduling logs as input and key storage performance indicators of all storage nodes in the hyperconverged architecture as output. This model includes: Collect key storage performance metrics of storage nodes in hyperconverged architecture at multiple historical moments. At the same time, obtain historical scheduling logs of data center resource scheduling software for each historical period at each historical moment. The historical scheduling logs cover the historical trend of all key storage performance metrics and business load labels in the corresponding historical period. Based on the key storage performance indicators of storage nodes in the hyperconverged architecture at multiple historical moments and the historical scheduling logs of the data center resource scheduling software for each historical period, a key storage performance indicator prediction model is constructed using a random forest model as the initial model. This model takes the business load labels and historical indicator change trends in the historical scheduling logs as inputs and the key storage performance indicators of all storage nodes in the hyperconverged architecture in a preset future period as outputs.
[0037] In this embodiment, a historical moment refers to a specific point in the past.
[0038] In this embodiment, the data center resource scheduling software is a tool for managing the scheduling of various resources within a data center. It can schedule and arrange resources such as storage nodes in a hyperconverged architecture and record relevant information, such as when storage nodes handle what kind of business load and the resource usage at the corresponding time.
[0039] In this embodiment, the historical period corresponding to a historical moment is a specific time range prior to the historical moment, such as one day as a historical period.
[0040] In this embodiment, key storage performance metrics of storage nodes based on a hyperconverged infrastructure at multiple past specific points in time are collected, such as IOPS and read / write latency at different times, as well as historical scheduling logs generated by the data center resource scheduling software within a specific time period at each corresponding historical moment. These logs contain information such as service load labels and changes in various metrics. Using a random forest model as the starting model, and leveraging this rich data, a series of data processing, analysis, and training processes are conducted to mine the relationship between service load labels and historical metric trends in the historical scheduling logs and key storage performance metrics of all storage nodes in the hyperconverged infrastructure for a specific future time period. Ultimately, a key storage performance metric prediction model is constructed that can use the service load labels and historical metric trends in the historical scheduling logs as input and output the key storage performance metrics of all storage nodes in the hyperconverged infrastructure for a preset future time period. For example, the performance metrics and scheduling log data of storage nodes at each historical moment are first organized into a format acceptable to the model and input into the random forest model framework. After multiple iterations of training and adjustment of model parameters, the model learns to predict future storage node performance metrics from the input data, thereby completing the construction of the key storage performance metric prediction model.
[0041] To provide performance data for subsequent decision-making, a key storage performance indicator prediction model is proposed to predict the key storage performance indicators of all storage nodes in a preset future period, including: Input the business load labels and historical indicator change trends of all storage nodes in the hyperconverged architecture in the latest historical period into the key storage performance indicator prediction model to obtain the key storage performance indicators of all storage nodes in the preset future period.
[0042] To provide a reasonable basis for cluster partitioning for performance analysis and management, a method for partitioning all sub-clusters in a hyperconverged architecture is proposed, including: Extract the physical rack number of all storage nodes, the network link affiliation with compute nodes, and the replica distribution rules of the distributed storage pool as hardware association features from the physical hardware topology of the hyperconverged architecture. Based on the hardware association characteristics of all storage nodes, all storage nodes are divided into multiple initial sub-clusters; Retrieve the historical scheduling logs of the data center resource scheduling software in the latest historical period, and extract the load fluctuation characteristics of each initial sub-cluster from the historical scheduling logs: the load fluctuation characteristics include the IOPS fluctuation variance of all storage nodes in the initial sub-cluster, the proportion of business load types, and the trend of storage fragmentation rate. Based on a preset clustering algorithm, with the clustering objective being that the load fluctuation variance does not exceed a preset fluctuation variance threshold and the proportion of similar businesses is not less than a preset proportion threshold, the initial sub-clusters are divided and merged a second time to obtain multiple sub-clusters. At the same time, based on a preset sub-cluster dynamic adjustment mechanism, all sub-clusters are dynamically adjusted.
[0043] In this embodiment, the physical hardware topology of the hyperconverged architecture describes the connection relationships and layout between various physical hardware devices in the hyperconverged architecture. For example, how servers, storage devices, etc. are interconnected via network cables, and their location distribution in the data center, etc.
[0044] In this embodiment, the physical rack number to which the storage node belongs clearly indicates the specific physical rack where the storage node is located, which facilitates its location and management in the data center environment; Network link attribution to computing nodes refers to the attribution relationship of network connections between storage nodes and computing nodes, which determines the data transmission path between storage and computing stages; The replica distribution rules of a distributed storage pool specify how data replicas are stored on different storage nodes to ensure data reliability and availability. For example, some rules require that data replicas be evenly distributed across storage nodes on different racks.
[0045] In this embodiment, the IOPS fluctuation variance of the storage node reflects the degree of fluctuation of its input / output operations per second; the larger the variance, the more severe the fluctuation. The service load type ratio reflects the proportion of different types of service loads in the total amount of services processed by the storage node. The storage fragmentation rate trend shows the increase or decrease of the storage fragmentation rate over time.
[0046] In this embodiment, the preset clustering algorithm is the K-means clustering algorithm.
[0047] In this embodiment, based on a preset clustering algorithm, with the clustering objective of load fluctuation variance not exceeding a preset fluctuation variance threshold and the proportion of similar services not less than a preset proportion threshold, the initial sub-clusters are divided and merged a second time to obtain multiple sub-clusters. This is achieved by using a pre-set clustering algorithm, according to the standard that the load fluctuation variance should be controlled below a certain set value, and the proportion of similar services in the total number of services in the sub-clusters should reach a certain set proportion or higher, and the initially divided sub-clusters are split and merged again to finally form multiple new sub-clusters.
[0048] In this embodiment, the preset fluctuation variance threshold is a pre-set numerical standard used to measure whether the load fluctuation variance is within an acceptable range. For example, if the threshold is set to 5, and the load fluctuation variance of a certain sub-cluster is 8, then it exceeds the preset fluctuation variance threshold.
[0049] In this embodiment, load fluctuation variance is a statistical measure derived by calculating the fluctuation of performance indicators such as IOPS of storage nodes over a period of time, used to quantify the stability of the load. The calculation method is as follows: first, calculate the average IOPS over this period, then calculate the square of the difference between IOPS and the average value at each time point, sum these squared values and divide by the number of data points to obtain the load fluctuation variance.
[0050] In this embodiment, the preset percentage threshold is a pre-set ratio standard used to measure whether the proportion of similar services in the total number of services in a sub-cluster meets the requirements. For example, if it is set to 70%, and the proportion of similar services in a certain sub-cluster is 60%, then the preset percentage threshold has not been met.
[0051] In this embodiment, the proportion of similar services refers to the percentage of the same type of service load in the total service load processed by the sub-cluster. For example, in a certain sub-cluster, the volume of "database read / write" services accounts for 40% of the total service volume of the sub-cluster, and this 40% is the proportion of similar services.
[0052] In this embodiment, all sub-clusters are dynamically adjusted based on a preset sub-cluster dynamic adjustment mechanism. This mechanism establishes a sub-cluster dynamic adjustment mechanism that reassesses the rationality of sub-clusters every 7 days based on real-time monitoring data (hardware health status of storage nodes, load fluctuation amplitude) from the intelligent computing cluster monitoring system. If more than 30% of the storage nodes in a sub-cluster exhibit hardware health anomalies, such as a disk bad sector warning value ≥5, a cache module attenuation coefficient ≥0.3, or a load difference rate between sub-clusters (the IOPS ratio between the highest-loaded sub-cluster and the lowest-loaded sub-cluster) exceeding 2.5 times, then a sub-cluster re-partitioning is triggered. Simultaneously, a tag is added to each sub-cluster, containing "hardware topology identifier (e.g., rack number-link group), dominant business type (e.g., 'database read / write cluster'), and load level (high / medium / low)", providing a basis for subsequent load fluctuation characteristic analysis and expansion strategy matching.
[0053] To facilitate early detection and response to potential performance anomaly propagation, this paper proposes a performance anomaly propagation chain based on the load fluctuation characteristics of all sub-clusters in a hyperconverged architecture, the hardware health characteristics and topological correlation characteristics of all storage nodes, key storage performance indicators in a preset future time period, and storage performance threshold ranges adapted to business needs. This chain includes: A topological association matrix is constructed based on the topological association characteristics of all storage nodes; Set a storage performance threshold range that aligns with business requirements; Based on the key storage performance indicators of all storage nodes in a preset future period, the slope of performance indicator change, performance drop rate, and fluctuation sequence of each storage node in the future period are determined. Based on the storage performance threshold range, all storage nodes with abnormal storage performance are screened out from all storage nodes. Based on the slope of performance metric changes, performance drop rate, and fluctuation sequence of all storage nodes in future time periods, the performance similarity between different storage nodes is calculated. Based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the hardware health characteristics of all storage nodes, the topology correlation matrix, all storage performance abnormal nodes, and the performance similarity between different storage nodes, the performance abnormality propagation chain was determined.
[0054] In this embodiment, a topology association matrix is constructed based on the topological association characteristics of all storage nodes. This matrix is created based on the topological association characteristics of the physical or logical connections and mutual influences between storage nodes. The elements in the matrix represent the degree of association between two corresponding storage nodes. If two storage nodes are directly connected, the corresponding matrix element value may be set to 1; otherwise, it is set to 0, thus comprehensively presenting the topological relationships between storage nodes.
[0055] In this embodiment, setting a storage performance threshold range that adapts to business needs involves determining the value range of various key storage performance indicators based on the specific storage performance requirements of different businesses. For example, for businesses with high real-time requirements, the read / write latency is set to be less than 10 milliseconds and the IOPS to be greater than 1000, thus defining a reasonable range of storage performance for that business scenario.
[0056] In this embodiment, based on the key storage performance indicators of all storage nodes in a preset future time period, the slope of performance indicator change, performance drop rate, and fluctuation sequence of each storage node in the future time period are determined. Then, based on the storage performance threshold range, all storage nodes with abnormal storage performance are screened out from all storage nodes. Specifically, by analyzing the changes of key storage performance indicators (such as IOPS, read / write latency, etc.) over time within the preset future time period, the slope of performance indicator change is calculated (for example, if the IOPS of a storage node increases from 500 to 600 in the next hour, the slope is (600-500) / 1=100), the performance drop rate is calculated (by calculating the ratio of the decrease in the key storage performance indicator of a storage node within a certain time period to the initial indicator value), and the fluctuation sequence of indicator fluctuations is recorded. These calculation results or key storage performance indicators are then compared with a pre-set storage performance threshold range. If the indicator of a storage node exceeds the threshold range, it is determined to be a storage performance abnormal node.
[0057] In this embodiment, the performance similarity between different storage nodes is calculated based on the slope of performance metric changes, performance drop rate, and fluctuation sequence of all storage nodes over future periods. Specifically, a particular algorithm (such as calculating the square root of the sum of squares of the differences in each metric, where the difference in the fluctuation sequence is the average of the differences of all elements in the sequence) is used to measure the similarity of the performance of different storage nodes. For example, if the square root of the sum of squares of the differences in the performance metric changes, performance drop rate, and fluctuation sequence of nodes A and B is small, it indicates that their performance similarity is high.
[0058] To clearly understand the propagation path and patterns of performance anomalies in storage systems, and to provide key information for early prevention and effective response to performance issues, this paper proposes a system based on the load fluctuation characteristics of all sub-clusters in a hyperconverged architecture, the hardware health characteristics of all storage nodes, the topology correlation matrix, all storage nodes with performance anomalies, and the performance similarity between different storage nodes. This system identifies the performance anomaly propagation chain, including: Based on the hardware health characteristics of all storage nodes, all nodes with health risks were screened out from all storage nodes. Based on the topological association matrix of all storage nodes, directly related node groups are identified, and an initial node association graph is constructed based on the directly related node groups; All nodes with abnormal storage performance and nodes with health risks are treated as key nodes of concern, and based on the initial node association graph, all indirectly related key node groups are identified among all key nodes of concern. The comprehensive correlation strength of each indirectly related key node group is calculated based on the performance similarity between different storage nodes and the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture. Based on the comprehensive association strength of each indirect association key node group and the indirect association path between the two key attention nodes contained in the corresponding indirect association key node group, the reliability of each smallest unit path in the corresponding indirect association path under the corresponding indirect association key node group is determined. The sum of the reliability of each smallest unit path in the initial node association graph under all indirectly associated key node groups is taken as the comprehensive reliability of each smallest unit path. In the initial node association graph, all the smallest unit paths whose overall reliability is less than the reliability threshold are removed to obtain the effective node association graph; Among all hypothetical anomaly propagation paths in the effective node association graph, all hypothetical anomaly propagation paths whose maximum overlap with all historical anomaly propagation paths is not less than the overlap threshold are selected as performance anomaly propagation chains.
[0059] In this embodiment, all storage nodes at health risk are screened out based on their hardware health characteristics. This involves using various hardware status information of the storage nodes, such as operating temperature, hardware fault alarm records, and usage time, to determine which storage nodes may have health risks. For example, if the operating temperature of a storage node is consistently high and exceeds the normal range, it may be identified as a node at health risk.
[0060] In this embodiment, directly related node groups are identified based on the topological association matrix of all storage nodes, and an initial node association graph is constructed based on these directly related node groups. First, starting from the topological association matrix describing the topological relationships of the storage nodes, sets of nodes with direct connections to each other are identified; these sets constitute the directly related node groups. Then, based on these directly related node groups, the relationships between nodes are graphically displayed, such as connecting directly related nodes with lines, thereby constructing the initial node association graph.
[0061] In this embodiment, based on the initial node association graph, all indirectly related key node groups are identified among all key nodes of interest. In the pre-constructed initial node association graph, for those nodes pre-defined as key nodes of interest, graph analysis is used to find combinations of nodes indirectly connected to them through other nodes; these combinations constitute indirectly related key node groups. For example, if key nodes A and B are connected through an intermediate node C, then A and B might constitute an indirectly related key node group.
[0062] In this embodiment, the comprehensive association strength of the indirect association key node group is a value used to measure the degree of association between the nodes in the indirect association key node group.
[0063] In this embodiment, based on the overall association strength of each indirectly associated key node group and the indirect association path between the two key nodes of interest contained in the corresponding indirectly associated key node group, the reliability of each smallest unit path in the corresponding indirect association path under the corresponding indirectly associated key node group is determined. The ratio of the overall association strength of each indirectly associated key node group to the total number of smallest unit paths contained in the indirect association path between the two key nodes of interest contained in the corresponding indirectly associated key node group is taken as the reliability of the corresponding smallest unit path under the corresponding indirectly associated key node group.
[0064] In this embodiment, the reliability threshold is a pre-set numerical standard used to determine whether the reliability of the smallest unit path meets the requirements. For example, the reliability threshold is set to 0.8.
[0065] In this embodiment, it is assumed that the anomaly propagation path is a path that is artificially assumed to propagate from one node to other nodes when analyzing abnormal performance of storage nodes. For example, it is assumed that the path starts from node X, passes through nodes Y and Z, and finally propagates to node W as the anomaly propagation path.
[0066] In this embodiment, the maximum overlap between the assumed abnormal propagation path and all historical abnormal propagation paths refers to comparing the currently assumed abnormal propagation path with all past abnormal propagation paths, identifying the historical path with the most overlap with the assumed path, and calculating the proportion of this maximum overlap to the assumed path or historical path. For example, if the assumed abnormal propagation path has 5 nodes, and a certain historical abnormal propagation path also has 5 nodes, with 4 nodes overlapping, then the maximum overlap is 4 ÷ 5 = 80%.
[0067] In this embodiment, the overlap threshold is a pre-set proportional standard used to determine whether the degree of overlap between the hypothetical abnormal propagation path and the historical abnormal propagation path meets certain requirements. For example, if the overlap threshold is set to 70%, a maximum overlap of 80% would meet the requirements, while 50% would not.
[0068] To provide strong data support for subsequently identifying the propagation chain of performance anomalies, a comprehensive correlation strength for each indirectly related key node group is calculated based on the performance similarity between different storage nodes and the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture. This includes: The load synergy between different storage nodes is calculated based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture. The overall association strength of each indirectly related key node group is calculated based on the performance similarity and load synergy between different storage nodes.
[0069] In this embodiment, the load synergy between different storage nodes is calculated based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture. For IOPS fluctuation variance, it can be quantified by calculating the deviation between the actual IOPS value and the average value over a period of time; the greater the deviation, the higher the value. The change in the proportion of business load types can be quantified by calculating the difference in the proportion of the same type of business load over different time periods; the greater the difference, the higher the value. The trend of storage fragmentation rate can be quantified based on the increase or decrease of storage fragmentation rate over a period of time; the greater the increase or decrease, the higher the value. Then, corresponding weights are assigned to these quantified values. For example, the weight of IOPS fluctuation variance is set to 0.4, the weight of change in the proportion of business load types is set to 0.3, and the weight of the trend of storage fragmentation rate is set to 0.3. Finally, each quantified feature value is multiplied by its corresponding weight and then summed. The result is the load synergy between different storage nodes.
[0070] In this embodiment, the comprehensive association strength of each indirectly related key node group is calculated based on the performance similarity and load synergy between different storage nodes. This involves assigning specific weights to performance similarity and load synergy; for example, assuming a performance similarity weight of 0.6 and a load synergy weight of 0.4, the performance similarity and load synergy among the storage nodes involved in each indirectly related key node group are weighted and calculated accordingly. The result is the comprehensive association strength of that indirectly related key node group. For instance, if the indirectly related key node group consisting of nodes A and B has a performance similarity of 0.8 and a load synergy of 0.7, then the comprehensive association strength = 0.8 × 0.6 + 0.7 × 0.4 = 0.76.
[0071] To achieve reasonable and accurate capacity expansion planning, a method for determining the scale of capacity expansion based on supply and demand coordination is proposed, including: Analyze the basic capacity requirements, peak business elasticity requirements, and data redundancy requirements for demand-side business growth; Based on historical capacity change data of hyperconverged storage pools, calculate the actual available capacity and capacity growth loss of supply-side storage resources; Based on the basic capacity requirements of demand-side business growth, peak business elasticity requirements, data redundancy requirements, and the actual available capacity and capacity growth attrition rate of supply-side storage resources, the scale of supply-demand coordinated expansion is determined:
[0072] In this embodiment, the basic capacity requirements, peak business elasticity requirements, and data redundancy requirements of demand-side business growth are analyzed: Basic capacity requirements are the storage space necessary for the normal development of a business. For example, if a new database application is expected to grow its data volume to 100GB within the next year, this 100GB is the basic capacity requirement.
[0073] Peak demand elasticity takes into account the extra capacity required by the business during peak periods. For example, during promotional events, e-commerce platforms experience a surge in data access and storage. The additional capacity required at this time is the peak demand elasticity.
[0074] Data redundancy requirements are the space needed to store additional backup data in order to ensure data security and reliability. For example, if important data needs to be backed up twice, then the required capacity of the backup data is the data redundancy requirement.
[0075] In this embodiment, based on historical capacity change data of the hyperconverged storage pool, the actual available capacity and capacity growth loss rate of the supply-side storage resources are calculated: Actual available capacity refers to the capacity of the hyperconverged storage pool that is truly available for business use after deducting various usages. For example, if the initial total capacity of the hyperconverged storage pool is 500GB, the system uses 50GB, and 100GB has been allocated to other businesses, then the actual available capacity is 500-50-100=350GB.
[0076] Capacity growth loss rate is the percentage of capacity loss caused by various factors (such as hardware aging, storage format conversion, etc.) as storage capacity increases. Assuming the initial storage capacity is 100GB, and it grows to 120GB after a period of time, but the actual usable capacity only increases to 115GB, then the capacity growth loss = 120 - 115 = 5GB, and the capacity growth loss rate = 5 ÷ (120 - 100) × 100% = 25%.
[0077] like Figure 3 As shown, in order to achieve targeted scaling decisions, a flexible scaling strategy based on supply and demand coordination and the selection of scaling scale and performance anomaly propagation chain is proposed, including: When the business scenario of the hyperconverged architecture is the core business scenario, the corresponding elastic scaling strategy is selected from the scaling strategy library under the core business scenario based on the scale type of the scaling scale of supply and demand coordination and the impact range of the performance anomaly propagation chain. When the business scenario of the hyperconverged architecture is a non-core business scenario, the corresponding elastic scaling strategy is selected from the scaling strategy library for non-core business scenarios based on the scale type of the scaling scale of supply and demand coordination and the impact range of the performance anomaly propagation chain.
[0078] In this embodiment, core business scenarios refer to the operational environment and conditions that play a critical and core role in the entire business system. These businesses are typically essential to enterprise operations and service provision; problems in these areas can have serious negative impacts on the enterprise. For example, a bank's core transaction business, which concerns the flow of funds and customer account security, falls under the category of core business scenarios.
[0079] In this embodiment, the expansion strategy library for core business scenarios is a collection of expansion strategies specifically designed for core business scenarios. This library contains various strategies formulated based on the characteristics, needs, and performance metrics of core businesses, used to select appropriate expansion solutions when core businesses face changes in storage performance. For example, when the storage pressure on core businesses increases, the strategy library may include strategies such as immediately adding high-performance storage devices or optimizing the storage architecture to ensure the stable operation of core businesses.
[0080] For example, the scale types to which supply and demand coordination-based expansion belongs include: Small-scale expansion: Expansion size ≤ 10% of current effective capacity; Medium-sized expansion: 10% < expansion scale ≤ 30%; Large-scale expansion: Expansion scale > 30%.
[0081] The impact of the performance anomaly propagation chain includes: Local propagation: The number of affected nodes is ≤ 15% of the total number of nodes; Global propagation: Affects more than 15% of nodes.
[0082] If the propagation is localized and involves small-scale expansion, and the performance anomaly propagation chain originates from excessive fragmentation (not insufficient capacity), the "virtualized storage reorganization + data redistribution" strategy should be prioritized: By inputting the topology correlation matrix into a graph neural network (GNN), the network link weights and data dependencies between storage nodes are learned, and the optimal data redistribution path with "minimum cross-rack migration and minimum data synchronization latency" is planned. At the same time, the "incremental reorganization + cache mirroring" mechanism is initiated: a memory cache mirror (with a mirror capacity of not less than 8% of the volume capacity) is created for the storage volume corresponding to the core business. During the reorganization process, core business IO requests are written to the cache first. After the reorganization is completed, the cached data is written to the reorganized storage volume through asynchronous disk flushing technology to ensure zero interruption of core business IO. If the performance anomaly propagation is global and involves medium-sized expansion, and the propagation chain originates from insufficient capacity and there are idle storage nodes, the "node-level expansion + abnormal node isolation" strategy is adopted: First, idle nodes are quickly incorporated into the storage pool through a distributed consistency protocol (such as Raft), and metadata is synchronized (metadata synchronization time is controlled within 30 seconds). Then, based on the load fluctuation feature matrix, the source node in the performance anomaly propagation chain is identified and temporarily isolated (during the isolation period, the business is taken over by a backup node). At the same time, the storage pool load balancing algorithm is optimized, and data blocks are allocated according to the node's IO capacity (combining historical IOPS peak and real-time load). The initial data volume allocated to newly incorporated nodes does not exceed 30% of their total capacity to avoid a sudden increase in the load on new nodes. If it is a global expansion and a large-scale expansion, and there are no idle storage nodes, choose the strategy of "storage media upgrade + cold and hot data tiered migration + resource sharing": First, based on the storage access log, collect cold data (accessed less than 2 times in the last 72 hours), compress it using the LZ4 compression algorithm (compression level set to 4, compression ratio ≥35%) and migrate it to object storage to release low-speed HDD capacity; then upgrade the core business storage nodes with high-speed SSDs (the upgrade process uses incremental data synchronization, and the migration time is ≤15 minutes); at the same time, temporarily share idle cache resources of computing nodes (such as idle video memory of GPU nodes, the sharing ratio is ≤40% of idle video memory) to alleviate storage IO pressure until the expansion is completed.
[0083] In other cases, return to redetermine the scale type to which the supply and demand coordination expansion scale belongs and the scope of the impact of the performance anomaly propagation chain.
[0084] In this embodiment, non-core business scenarios, relative to core business scenarios, refer to business operating environments and conditions that, while having some effect on enterprise operations, are of lower importance. Even if such businesses experience temporary problems, the overall impact on the enterprise is relatively small. For example, the business of storing daily training materials for employees within the enterprise belongs to non-core business scenarios.
[0085] In this embodiment, the expansion strategy library for non-core business scenarios is a collection of expansion strategies specifically designed for non-core business scenarios. Because the impact of non-core businesses on an enterprise differs from that of core businesses, the strategies in this library will have different considerations regarding cost control and implementation complexity. For example, a relatively economical expansion method may be adopted, such as simple expansion using idle equipment, or expansion operations may be performed during periods of low business activity, in order to balance cost and business needs.
[0086] For example: If it is a localized expansion and a small-scale expansion, and the performance abnormality is due to insufficient capacity, choose the "lightweight node expansion" strategy: directly mount part of the capacity of the idle storage node (not exceeding 50% of the total node capacity) to the hyperconverged storage pool, without having to include all nodes, and shorten the expansion response time through a simplified version of the metadata synchronization protocol (such as a simplified Raft protocol, reducing the number of voting nodes); If it is a global propagation and large-scale expansion, and there are no idle storage nodes, choose the "cold and hot data tiered migration + object storage expansion" strategy: after compressing and migrating the cold data to object storage, enable the "downgrade expansion" mode for non-core business storage nodes, reduce the number of storage replicas for non-core businesses (from 3 replicas to 2 replicas), release redundant capacity, and at the same time expand the archive storage layer of object storage to meet the long-term storage needs of massive cold data.
[0087] To ensure the efficient and comprehensive implementation of the scaling strategy, a cross-module collaborative execution mechanism is proposed to execute the elastic scaling strategy, including: Based on a cross-module collaborative execution mechanism, it simultaneously collaborates with the computing module, network module, and monitoring module to execute elastic scaling strategies.
[0088] In this embodiment, a cross-module collaborative execution mechanism is implemented: through a unified scheduling interface, the operations of the computing module (resource allocation), network module (link configuration), and monitoring module (real-time monitoring) are synchronized to ensure that the expansion steps are executed in an orderly manner.
[0089] In this embodiment, based on a cross-module collaborative execution mechanism, the elastic scaling strategy is executed simultaneously in collaboration with the computing module, network module, and monitoring module. This means that a mechanism exists that allows each module to cooperate and jointly implement the elastic scaling strategy. In this process, the execution of the elastic scaling strategy is not performed independently by a single module, but rather the computing module, network module, and monitoring module collaborate with each other in a specific way and order. For example, when storage resources need to be expanded, the computing module provides additional computing resources to handle tasks such as data analysis brought about by the expansion; the network module adjusts the network configuration to ensure smooth communication between the newly expanded storage and other devices; and the monitoring module monitors performance indicators throughout the process in real time to ensure the expansion proceeds smoothly.
[0090] In this embodiment, the computing module is primarily responsible for providing computing power. In a flexible expansion scenario, as storage capacity increases, more data may need to be processed, and the computing module takes on tasks such as analyzing and performing calculations on this data. For example, it may classify the newly stored data and run retrieval algorithms.
[0091] In this embodiment, the network module's role is to ensure smooth network communication. When implementing an elastic expansion strategy, newly added storage devices need to establish network connections with other modules. The network module is responsible for configuring network parameters, such as IP address allocation and routing settings, to ensure that data can be accurately and quickly transmitted between the storage module, computing module, and other related modules.
[0092] In this embodiment, the monitoring module primarily monitors the entire elastic scaling process and the operational status of each system module in real time. It collects data such as performance metrics of the storage module (e.g., read / write speed, storage utilization), resource usage of the computing module (e.g., CPU utilization, memory usage), and communication status of the network module (e.g., bandwidth utilization, network latency). By analyzing this data, potential problems can be identified promptly, such as storage performance bottlenecks or network congestion, allowing for timely measures to ensure the effective execution of the elastic scaling strategy.
[0093] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for predicting storage performance and elastic scaling for hyperconverged architectures, characterized in that, include: Construct a key storage performance indicator prediction model that takes business load tags and historical indicator change trends in historical scheduling logs as inputs and key storage performance indicators of all storage nodes in the hyperconverged architecture in a preset future time period as outputs. Based on the key storage performance indicator prediction model, the key storage performance indicators of all storage nodes are predicted in a preset future time period. Based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the hardware health characteristics and topology association characteristics of all storage nodes, the key storage performance indicators in the preset future period, and the storage performance threshold range adapted to business needs, the performance anomaly propagation chain is determined. Based on the supply and demand coordination of the expansion scale and the performance anomaly propagation chain, an elastic expansion strategy is selected and executed based on the cross-module collaborative execution mechanism.
2. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 1, characterized in that, Construct a key storage performance indicator prediction model that takes service load tags and historical indicator change trends from historical scheduling logs as input and key storage performance indicators of all storage nodes in the hyperconverged architecture as output, including: Collect key storage performance metrics of storage nodes in hyperconverged architecture at multiple historical moments. At the same time, obtain historical scheduling logs of data center resource scheduling software for each historical period at each historical moment. The historical scheduling logs cover the historical trend of all key storage performance metrics and business load labels in the corresponding historical period. Based on the key storage performance indicators of storage nodes in the hyperconverged architecture at multiple historical moments and the historical scheduling logs of the data center resource scheduling software for each historical period, a key storage performance indicator prediction model is constructed using a random forest model as the initial model. This model takes the business load labels and historical indicator change trends in the historical scheduling logs as inputs and the key storage performance indicators of all storage nodes in the hyperconverged architecture in a preset future period as outputs.
3. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 1, characterized in that, Based on the key storage performance indicator prediction model, the key storage performance indicators of all storage nodes are predicted for a preset future period, including: Input the business load labels and historical indicator change trends of all storage nodes in the hyperconverged architecture in the latest historical period into the key storage performance indicator prediction model to obtain the key storage performance indicators of all storage nodes in the preset future period.
4. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 1, characterized in that, The methods for partitioning all sub-clusters in a hyperconverged architecture include: Extract the physical rack number of all storage nodes, the network link affiliation with compute nodes, and the replica distribution rules of the distributed storage pool as hardware association features from the physical hardware topology of the hyperconverged architecture. Based on the hardware association characteristics of all storage nodes, all storage nodes are divided into multiple initial sub-clusters; Retrieve the historical scheduling logs of the data center resource scheduling software in the latest historical period, and extract the load fluctuation characteristics of each initial sub-cluster from the historical scheduling logs: the load fluctuation characteristics include the IOPS fluctuation variance of all storage nodes in the initial sub-cluster, the proportion of business load types, and the trend of storage fragmentation rate. Based on a preset clustering algorithm, with the clustering objective being that the load fluctuation variance does not exceed a preset fluctuation variance threshold and the proportion of similar businesses is not less than a preset proportion threshold, the initial sub-clusters are divided and merged a second time to obtain multiple sub-clusters. At the same time, based on a preset sub-cluster dynamic adjustment mechanism, all sub-clusters are dynamically adjusted.
5. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 1, characterized in that, Based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the hardware health characteristics and topology correlation characteristics of all storage nodes, key storage performance indicators in a preset future period, and storage performance threshold ranges adapted to business needs, the performance anomaly propagation chain is determined, including: A topological association matrix is constructed based on the topological association characteristics of all storage nodes; Set a storage performance threshold range that aligns with business requirements; Based on the key storage performance indicators of all storage nodes in a preset future period, the slope of performance indicator change, performance drop rate, and fluctuation sequence of each storage node in the future period are determined. Based on the storage performance threshold range, all storage nodes with abnormal storage performance are screened out from all storage nodes. Based on the slope of performance metric changes, performance drop rate, and fluctuation sequence of all storage nodes in future time periods, the performance similarity between different storage nodes is calculated. Based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the hardware health characteristics of all storage nodes, the topology correlation matrix, all storage performance abnormal nodes, and the performance similarity between different storage nodes, the performance abnormality propagation chain was determined.
6. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 5, characterized in that, Based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the hardware health characteristics of all storage nodes, the topology correlation matrix, all storage performance anomaly nodes, and the performance similarity between different storage nodes, the performance anomaly propagation chain was determined, including: Based on the hardware health characteristics of all storage nodes, all nodes with health risks were screened out from all storage nodes. Based on the topological association matrix of all storage nodes, directly related node groups are identified, and an initial node association graph is constructed based on the directly related node groups; All nodes with abnormal storage performance and nodes with health risks are treated as key nodes of concern, and based on the initial node association graph, all indirectly related key node groups are identified among all key nodes of concern. The comprehensive correlation strength of each indirectly related key node group is calculated based on the performance similarity between different storage nodes and the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture. Based on the comprehensive association strength of each indirect association key node group and the indirect association path between the two key attention nodes contained in the corresponding indirect association key node group, the reliability of each smallest unit path in the corresponding indirect association path under the corresponding indirect association key node group is determined. The sum of the reliability of each smallest unit path in the initial node association graph under all indirectly associated key node groups is taken as the comprehensive reliability of each smallest unit path. In the initial node association graph, all the smallest unit paths whose overall reliability is less than the reliability threshold are removed to obtain the effective node association graph; Among all hypothetical anomaly propagation paths in the effective node association graph, all hypothetical anomaly propagation paths whose maximum overlap with all historical anomaly propagation paths is not less than the overlap threshold are selected as performance anomaly propagation chains.
7. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 6, characterized in that, Based on the performance similarity between different storage nodes and the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture, the comprehensive correlation strength of each indirectly related key node group is calculated, including: The load synergy between different storage nodes is calculated based on the load fluctuation characteristics of all sub-clusters in the hyperconverged architecture. The overall association strength of each indirectly related key node group is calculated based on the performance similarity and load synergy between different storage nodes.
8. The storage performance prediction and elastic scaling method for hyperconverged architectures according to claim 1, characterized in that, Methods for determining the scale of supply and demand coordination expansion include: Analyze the basic capacity requirements, peak business elasticity requirements, and data redundancy requirements for demand-side business growth; Based on historical capacity change data of hyperconverged storage pools, calculate the actual available capacity and capacity growth loss of supply-side storage resources; Based on the basic capacity requirements of demand-side business growth, peak business elasticity requirements, data redundancy requirements, and the actual available capacity and capacity growth loss rate of supply-side storage resources, the scale of capacity expansion in coordination with supply and demand is determined.
9. The method for predicting storage performance and elastic scaling for hyperconverged architectures according to claim 1, characterized in that, Based on supply and demand coordination, the scaling scale and performance anomaly propagation chain are used to select elastic scaling strategies, including: When the business scenario of the hyperconverged architecture is the core business scenario, the corresponding elastic scaling strategy is selected from the scaling strategy library under the core business scenario based on the scale type of the scaling scale of supply and demand coordination and the impact range of the performance anomaly propagation chain. When the business scenario of the hyperconverged architecture is a non-core business scenario, the corresponding elastic scaling strategy is selected from the scaling strategy library for non-core business scenarios based on the scale type of the scaling scale of supply and demand coordination and the impact range of the performance anomaly propagation chain.
10. The method for predicting storage performance and elastic scaling for hyperconverged architectures according to claim 1, characterized in that, The elastic scaling strategy is implemented based on a cross-module collaborative execution mechanism, including: Based on a cross-module collaborative execution mechanism, it simultaneously collaborates with the computing module, network module, and monitoring module to execute elastic scaling strategies.