Cluster management method and device, equipment, storage medium and computer program product

By monitoring and managing cluster performance parameters and dynamically controlling resource usage of cloud storage instances, the performance degradation caused by cloud hard disk instance speed limit is solved, service performance is improved, and service level agreement commitments are met.

CN119960695APending Publication Date: 2025-05-09CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510069885.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the cluster performance oversell scenario, the storage performance of cloud hard disk instances is degraded due to speed limit, resulting in service performance degradation, and the commitment to customer service level agreement (SLA) is not met, which is prone to cause customer complaints.

Method used

By monitoring cluster performance parameters, determine the preset performance parameters of the cloud storage instance set, and manage the cluster and cloud storage instance set based on these parameters, including closing the burst management function, migrating cloud storage instances, performing backpressure processing, etc., to dynamically control QOS flow and ensure storage performance.

Benefits of technology

It effectively solves the problem of storage performance degradation caused by speed limit of cloud hard disk instances, improves the service performance of cloud hard disk instances, and ensures the commitment to customer service level agreements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960695A_ABST
    Figure CN119960695A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster management method, and the method comprises the steps: monitoring a first cluster, and obtaining a first parameter value of a first performance parameter; determining a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster; wherein the first cloud storage instance set comprises cloud storage instances of an exclusive cloud storage instance type and cloud storage instances of a shared cloud storage instance type; and performing resource management on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value. The invention further discloses a cluster management device and equipment, a storage medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage application technology, and in particular to a cluster management method, apparatus, device, storage medium and computer program product. Background Art

[0002] With the rapid development of big data analysis, artificial intelligence and other fields, the demand for computing power to deliver single disk performance has increased sharply. In the context of overselling performance, the stability of cluster performance has been challenged. In the scenario of centralized cluster performance delivery, the cluster reaches the performance limit, which is referred to as performance explosion. The cluster business and IO request performance are degraded, and the service level agreement (SLA) commitment to customers cannot be met, which is easy to cause customer complaints. In order to maintain cluster stability, common means include cloud hard disk service quality (QoS) flow control, cloud hard disk instance migration, cluster expansion and other means. Cloud hard disk QOS flow control is usually used as an in-process processing means. Common cloud hard disk QOS flow control solutions include selecting high system input / output operations per second (IOPS) and high throughput cloud hard disks to limit the speed or amortize the cluster performance over-limit value to all volumes and limit the speed in proportion until the cluster performance level drops to a reasonable level. However, in the above cloud hard disk QOS flow control process, the cloud hard disk instance is limited in speed, resulting in reduced storage performance of the cloud hard disk instance and reduced service performance of the cloud hard disk instance.

[0003] Application Contents

[0004] In order to solve the above technical problems, the present application hopes to provide a cluster management method, apparatus, equipment, storage medium and computer program product, which solves the problem of speed limiting the cloud hard disk instance resulting in a decrease in the storage performance of the cloud hard disk instance. During the application process, the QOS flow of the cluster is dynamically controlled to ensure the storage performance of the cloud hard disk instance and improve the service performance of the cloud hard disk instance.

[0005] The technical solution of this application is implemented as follows:

[0006] The present application provides a cluster management method, the method comprising:

[0007] monitoring the first cluster to obtain a first parameter value of a first performance parameter;

[0008] Determine a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster; wherein the first cloud storage instance set includes cloud storage instances of exclusive cloud storage instance types and cloud storage instances of shared cloud storage instance types;

[0009] Based on the first parameter value and the second parameter value, resource management is performed on the first cluster and the first cloud storage instance set.

[0010] In the above solution, the resource management of the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value includes:

[0011] Determining a first performance over-limit value of the first cluster based on the first parameter value and a cluster performance threshold;

[0012] If the first performance limit value is greater than or equal to the second parameter value, disabling the burst management function of the first cluster;

[0013] Determining a cloud storage instance to be migrated from the first cloud storage instance set;

[0014] Determining a target cluster from the cluster system to which the first cluster belongs;

[0015] Migrate the cloud storage instance to be migrated to the target cluster.

[0016] In the above scheme, the method further comprises:

[0017] If the first performance over-limit value is less than the second parameter value, determine the cloud storage instances whose parameter values ​​of the second performance parameters are greater than or equal to the first performance threshold from the first cloud storage instance set, and obtain a second cloud storage instance set;

[0018] If the second cloud storage instance set does not include a performance-surge cloud storage instance, determine a sum of third parameter values ​​of preset performance parameters of all cloud storage instances of the shared cloud storage instance type in the first cloud storage instance set to obtain a first reference value;

[0019] Based on the first performance limit value and the first reference value, the cloud storage instances in the second cloud storage instance set are managed.

[0020] In the above scheme, the method further comprises:

[0021] If the second cloud storage instance set includes a performance-surge cloud storage instance, determining a parameter surge amount of the second performance parameter of each instance in the second cloud storage instance set;

[0022] If there is a first instance to be adjusted in the second cloud storage instance set in which the parameter sudden increment is less than the second reference value of the corresponding instance and the parameter sudden increment is greater than the first performance limit value, back pressure processing of the second performance parameter is performed on the first instance to be adjusted based on the parameter sudden increment of the first instance to be adjusted; wherein the second reference value is the difference between the current parameter value of the second performance parameter of the corresponding instance and the corresponding first performance threshold.

[0023] In the above scheme, the method further comprises:

[0024] If there is a second instance to be adjusted in the second cloud storage instance set whose parameter sudden increment is greater than or equal to the second reference value of the corresponding instance, the second performance parameter of the second instance to be adjusted is limited to the corresponding first performance threshold.

[0025] In the above scheme, the method further comprises:

[0026] monitoring a third parameter value of a first performance parameter of the first cluster;

[0027] determining a second performance over-limit value of the first cluster based on the third parameter value and the cluster performance threshold;

[0028] Determine a sum of third parameter values ​​of preset performance parameters of all cloud storage instances belonging to the shared cloud storage instance type in the first cloud storage instance set to obtain a first reference value;

[0029] The second cloud storage instance set is managed based on the second performance limit value and the first reference value.

[0030] In the above solution, the managing the second cloud storage instance set based on the second performance limit value and the first reference value includes:

[0031] If the second performance limit value is greater than or equal to the first reference value, disabling the burst management function of all cloud storage instances of the shared cloud storage instance type in the second cloud storage instance set;

[0032] monitoring a fourth parameter value of the first performance parameter of the first cluster;

[0033] If the fourth parameter value is greater than or equal to the cluster performance threshold, determine all instances of the exclusive cloud storage instance type in the second cloud storage instance set to obtain one or more third instances to be adjusted;

[0034] Managing one or more of the third instances to be adjusted.

[0035] In the above solution, the managing of one or more third instances to be adjusted includes:

[0036] Determining a preset configuration of each of the third instances to be adjusted;

[0037] Based on the preset configuration of one or more of the third instances to be adjusted, back pressure processing is performed on the second performance parameter of the corresponding third instances to be adjusted until the third parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each of the third instances to be adjusted is the corresponding first performance threshold.

[0038] In the above scheme, the method further comprises:

[0039] If the second performance limit value is less than the first reference value, the preset configuration of the fourth instance to be adjusted belonging to the shared cloud storage instance type in the second cloud storage instance set is used to perform back pressure processing on the corresponding fourth instance to be adjusted until the fifth parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each of the fourth instance to be adjusted is the corresponding first performance threshold.

[0040] In the above scheme, the first performance threshold of the cloud storage instance of the exclusive cloud storage instance type is provided based on the configured second performance parameter, the first performance threshold includes a first capacity and a second capacity related to the second performance parameter, and the second capacity has an associated relationship with the first capacity; the cloud storage instance of the shared cloud storage instance type is configured to provide a minimum performance value corresponding to the second performance parameter.

[0041] The present application provides a cluster management device, the device comprising at least: a monitoring unit, a determination unit and a management unit; wherein:

[0042] The monitoring unit is used to monitor the first cluster and obtain a first parameter value of a first performance parameter;

[0043] The determining unit is used to determine a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster; wherein the first cloud storage instance set includes cloud storage instances of exclusive cloud storage instance types and cloud storage instances of shared cloud storage instance types;

[0044] The management unit is used to manage the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value.

[0045] The present application provides a cluster management device, the device comprising at least: a communication interface, a memory, a processor and a communication bus; wherein:

[0046] The memory is used to store executable instructions;

[0047] The communication bus is used to realize the communication connection between the communication interface, the processor and the memory;

[0048] The processor is used to execute the cluster management program stored in the memory to implement the steps of any one of the cluster management methods described above.

[0049] The present application provides a storage medium, on which a cluster management program is stored. When the cluster management program is executed, it is used to implement the steps of any of the above-mentioned cluster management methods.

[0050] The present application provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the steps of the cluster management method as described in any one of the above items are implemented.

[0051] The embodiment of the present application provides a cluster management method, apparatus, device, storage medium and computer program product, which monitors the first cluster, obtains the first parameter value of the first performance parameter, determines the second parameter value of the preset performance parameter of the first cloud storage instance set corresponding to the first cluster, and performs resource management on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value. In this way, by comparing and analyzing the first parameter value of the first performance parameter of the first cluster with the second parameter value of the preset performance parameter, the first cluster and the first cloud storage instance set are managed, which solves the problem that the storage performance of the cloud hard disk instance is reduced due to speed limiting of the cloud hard disk instance. In the application process, the QOS flow of the cluster is dynamically controlled to ensure the storage performance of the cloud hard disk instance and improve the service performance of the cloud hard disk instance. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of a cluster management method provided in an embodiment of the present application;

[0053] Figure 2 A schematic diagram of a process flow for implementing an application embodiment of a cluster management method provided in an embodiment of the present application;

[0054] Figure 3 A schematic diagram of the structure of a cluster management device provided in an embodiment of the present application;

[0055] Figure 4 A schematic diagram of the structure of a cluster management device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0057] The embodiment of the present application provides a cluster management method, referring to Figure 1 As shown, the method is applied to a cluster management device, and the method comprises the following steps:

[0058] Step 101: monitor a first cluster to obtain a first parameter value of a first performance parameter.

[0059] In the embodiment of the present application, the cluster management device may be a device for managing the cluster, having functions such as computing, for example, a computer device, a server device, etc. The cluster management device is mainly arranged on the cluster operator side, and is used to perform corresponding cluster management for the first cluster provided by the user, including the read and write rate when providing read and write services. The first performance parameter is a service performance parameter used to represent the first cluster application process, for example, it may be a parameter used to quantify the data flow when the first cluster receives the customer data flow.

[0060] Exemplarily, the first performance parameter of the first cluster may be an input / output (IO) parameter of the first cluster. That is, when monitoring the first cluster, it is sufficient to monitor its IO changes.

[0061] Step 102: Determine a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster.

[0062] The first cloud storage instance set includes cloud storage instances of exclusive cloud storage instance types and cloud storage instances of shared cloud storage instance types.

[0063] In an embodiment of the present application, the first cloud storage instance set corresponding to the first cluster includes one or more cloud storage instances corresponding to the first cluster. The cloud storage instance is a block storage product provided by the operator to the user. During the application process, the public cloud operator mainly uses capacity reservation as the billing model for block storage, that is, the customer selects a specific block storage product specification, specifies the instance capacity, and charges according to the length of use. The preset performance parameter is an allowable fluctuation parameter that is allocated to each cloud storage instance in the first cloud storage instance set and is allowed to exceed the purchased capacity. A cloud storage instance of the exclusive cloud storage instance type is a cloud storage instance that sets the SLA commitment based on the performance parameter value of the cloud storage instance ordered by the customer, and a cloud storage instance of the shared cloud storage instance type is a cloud storage instance for which the operator promises the most detailed performance parameter value to the customer. It should be noted that a cloud storage instance can be referred to as an instance.

[0064] Exemplarily, the preset performance parameter of the first cloud storage example set corresponding to the first cluster may be a preset allowed burst IO value.

[0065] Step 103: Perform resource management on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value.

[0066] In an embodiment of the present application, the first parameter value and the second parameter value are analyzed and processed to obtain the analysis result. Then, it is determined whether to shut down the relevant functions of the first cluster or to perform back pressure processing on the cloud storage instances in the first cloud storage instance set based on the analysis result, thereby reducing the possibility of a decrease in the storage performance of the cluster when traffic bursts occur.

[0067] Exemplarily, the real-time IO of the first cluster and the preset burst IO value are analyzed to determine the specific burst condition of the first cluster, and the resources of the first cluster and the first cloud storage instance set are managed according to the specific burst condition.

[0068] Based on the foregoing embodiment, in other embodiments of the present application, resource management is performed on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value, including:

[0069] Determining a first performance over-limit value of the first cluster based on the first parameter value and the cluster performance threshold;

[0070] If the first performance over-limit value is greater than or equal to the preset performance parameter, shutting down the burst management function of the first cluster;

[0071] Determine a cloud storage instance to be migrated from the first cloud storage instance set;

[0072] Determine a target cluster from the cluster system to which the first cluster belongs;

[0073] Migrate the cloud storage instance to be migrated to the target cluster.

[0074] In the embodiment of the present application, the difference between the first parameter value and the cluster performance threshold is calculated to obtain the first performance over-limit value of the first cluster. Generally, when the first cluster is over-limited, the first performance over-limit value is greater than 0. The preset performance parameter can be a certain performance parameter that is allocated to the first cluster after the customer purchases the first cluster and allows over-limit, so as to ensure that when the customer's data traffic bursts, the first cluster can still ensure the processing efficiency of the data traffic and ensure the user's experience effect.

[0075] When the first performance limit value is greater than or equal to the preset performance parameter, it indicates that the current data traffic of the first cluster has exceeded the maximum allowed value. At this time, the burst management function of the first cluster is turned off, and the target cluster is determined from the cluster system to which the first cluster belongs. The cloud storage instance that the first cluster cannot currently process is determined from the first cloud storage instance set, and the cloud storage instance to be migrated is obtained. Then, the cloud storage instance to be migrated is migrated to the target cluster. In this way, the corresponding data flow of the customer can be quickly processed, which ensures the corresponding business processing capability and improves the corresponding user experience.

[0076] Based on the foregoing embodiment, in other embodiments of the present application, the method further includes:

[0077] If the first performance over-limit value is less than the second parameter value, determine the cloud storage instances whose parameter values ​​of the second performance parameters are greater than or equal to the first performance threshold from the first cloud storage instance set to obtain a second cloud storage instance set;

[0078] If the second cloud storage instance set does not include a performance surge cloud storage instance, determine a sum of third parameter values ​​of preset performance parameters of all cloud storage instances of a shared cloud storage instance type in the first cloud storage instance set to obtain a first reference value;

[0079] Based on the first performance limit value and the first reference value, the cloud storage instances in the second cloud storage instance set are managed.

[0080] In the embodiment of the present application, the second performance parameter is used to represent the performance of the cloud storage instance, for example, it can be represented by IOPS. The first performance threshold is the burst performance parameter allowed by the corresponding cloud storage instance. The first performance thresholds of different cloud storage instances can be the same or different, which is specifically determined by the performance of the corresponding cloud storage instance. The preset performance parameter is the allowed burst value of the second performance parameter of each cloud storage instance, for example, a burst IO parameter.

[0081] An instance analysis is performed on the cloud storage instances in the determined second cloud storage instance set to determine whether the second cloud storage instance set includes a performance-surge cloud storage instance. When the second cloud storage instance set does not include a performance-surge cloud storage instance, a third parameter value of a preset performance parameter of all instances of the shared cloud storage instance type in the first cloud storage instance set is determined, and a sum of the third parameter values ​​of all the determined shared cloud storage instances is calculated to obtain a first reference value. Finally, the first performance over-limit value and the first reference value of the first cluster are analyzed to manage the cloud storage instances in the second cloud storage instance set.

[0082] Exemplarily, the third parameter value may be a burst IO volume of a shared cloud storage instance.

[0083] Based on the foregoing embodiment, in other embodiments of the present application, the method further includes the following steps:

[0084] If the second cloud storage instance set includes a performance burst cloud storage instance, determining a parameter burst amount of a second performance parameter of each instance in the second cloud storage instance set;

[0085] If there is a first instance to be adjusted in the second cloud storage instance set whose parameter sudden increment is less than the second reference value of the corresponding instance and whose parameter sudden increment is greater than the first performance limit value, the second performance parameter of the first instance to be adjusted is back pressured based on the parameter sudden increment of the first instance to be adjusted; wherein the second reference value is the difference between the current parameter value of the second performance parameter of the corresponding instance and the corresponding first performance threshold.

[0086] In the embodiment of the present application, when the first performance over-limit value of the first cluster is less than the second parameter value, but the first cluster still exceeds the limit, the cloud storage instances in the first cloud storage instance set can be processed at this time, and the corresponding processing process is: determine the cloud storage instances whose parameter values ​​of the corresponding second performance parameters are greater than or equal to the corresponding first performance threshold from the first cloud storage instance set, and after obtaining the second cloud storage instance set, count the parameter sudden increment corresponding to the second performance parameter of each instance in the second cloud storage instance set, that is, the difference between the actual value of the second performance parameter of each cloud storage instance in the second cloud storage instance set and the actual value of the second performance parameter of the corresponding instance collected in the adjacent previous collection cycle. Determine the cloud storage instances in the second cloud storage instance set whose parameter sudden increment is less than the corresponding second reference value and whose parameter sudden increment is also greater than the first performance over-limit value of the first cluster, and obtain one or more first instances to be adjusted. Perform back pressure processing on the corresponding first instance to be adjusted according to the parameter sudden increment corresponding to each first instance to be adjusted, and reduce the value of the second performance parameter of the corresponding first instance to be adjusted, so as to ensure the stability of the second performance parameter of the corresponding first instance to be adjusted and reduce the possibility of performance explosion.

[0087] Based on the foregoing embodiment, in the embodiment of the present application, the method further includes:

[0088] If there is a second instance to be adjusted in the second cloud storage instance set whose parameter sudden increment is greater than or equal to the second reference value of the corresponding instance, the second performance parameter of the second instance to be adjusted is limited to the corresponding first performance threshold.

[0089] In an embodiment of the present application, when there is a second instance to be adjusted in the second cloud storage instance set whose parameter sudden increment is greater than or equal to the second reference value of the corresponding instance, the second performance parameter of the second instance to be adjusted is restricted to the upper limit value of the performance parameter value corresponding to the second instance to be adjusted, which is the corresponding first performance threshold.

[0090] In this way, the corresponding cloud storage instance can be guaranteed to provide services according to the maximum performance allowed, ensuring storage efficiency.

[0091] Based on the foregoing embodiment, in other embodiments of the present application, the method further includes:

[0092] monitoring a third parameter value of the first performance parameter of the first cluster;

[0093] Determining a second performance over-limit value of the first cluster based on a third parameter value and a cluster performance threshold;

[0094] Determine a sum of third parameter values ​​of preset performance parameters of all cloud storage instances of the shared cloud storage instance type in the first cloud storage instance set to obtain a first reference value;

[0095] Based on the second performance limit value and the first reference value, the second cloud storage instance set is managed.

[0096] In an embodiment of the present application, after back pressure processing of the second performance parameter is performed on one or more first instances to be adjusted, or after the value of the second performance parameter of each instance in the second cloud storage instance is set to the corresponding first performance threshold, the first performance parameter of the first cluster is continuously monitored to determine whether the value of the first performance parameter of the first cluster has recovered to a normal level, that is, whether the second parameter value is less than the cluster performance threshold. If the second parameter value is less than the cluster performance threshold, it indicates that the first performance parameter of the first cluster has recovered to a normal level. If the second parameter value of the first cluster has not recovered to a normal level, the difference between the second parameter value and the cluster performance threshold at this time is determined to obtain a second performance over-limit value, and all cloud storage instances of the type of shared cloud storage instance are determined from the second cloud storage instance set, and the sum of the third parameter values ​​of the preset performance parameters of all cloud storage instances belonging to the shared cloud storage instance, such as the burst IO parameter, is calculated to obtain a first reference value. The first reference value and the first performance over-limit value of the first cluster at this time are analyzed, and the second cloud storage instance set is managed according to the analysis result.

[0097] Based on the foregoing embodiment, in other embodiments of the present application, managing the second cloud storage instance set based on the second performance limit value and the first reference value includes:

[0098] If the second performance limit-exceeding value is greater than or equal to the first reference value, shutting down the burst management function of all cloud storage instances of the shared cloud storage instance type in the second cloud storage instance set;

[0099] monitoring a fourth parameter value of the first performance parameter of the first cluster;

[0100] If the fourth parameter value is greater than or equal to the cluster performance threshold, all instances of the exclusive cloud storage instance type in the second cloud storage instance set are determined to obtain one or more third instances to be adjusted;

[0101] One or more third instances to be adjusted are managed.

[0102] In an embodiment of the present application, when it is determined that the second performance limit value is greater than or equal to the first reference value, the burst management function of all cloud storage instances belonging to the shared cloud storage instance type in the second cloud storage instance set is turned off, and after turning off the burst management function of all cloud storage instances belonging to the shared cloud storage instance type in the second cloud storage instance set, the change of the first performance parameter of the first cluster is monitored to obtain the third parameter value of the first performance parameter of the first cluster at this time. If the third parameter value is greater than the cluster performance threshold, it indicates that the first performance parameter of the first cluster has not yet returned to normal. At this time, the instances of the exclusive cloud storage instance type in the second cloud storage instance set are determined to obtain one or more second instances to be adjusted, and then the second performance parameters of one or more second instances to be adjusted are adjusted.

[0103] Based on the foregoing embodiment, in other embodiments of the present application, managing one or more second instances to be adjusted includes:

[0104] Determining a preset configuration for each second instance to be adjusted;

[0105] Based on the preset configuration of one or more third instances to be adjusted, back pressure processing is performed on the second performance parameter of the corresponding third instances to be adjusted until the fifth parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each third instance to be adjusted is the corresponding first performance threshold.

[0106] In the embodiment of the present application, the preset configuration of each third instance to be adjusted is determined according to the performance of the third instance to be adjusted, such as capacity parameters, when allocating the corresponding third instance to be adjusted to the customer. The difference between the third parameter value and the cluster performance threshold is calculated to obtain the third performance over-limit value of the first cluster at this time. The third performance over-limit value is allocated according to the preset configuration of one or more second instances to be adjusted to obtain the adjusted value of the second performance parameter of each third instance to be adjusted. Back pressure processing is performed on the second performance parameter of the corresponding third instance to be adjusted according to each adjusted value. During each back pressure process, the value of the second performance parameter of each third instance to be adjusted is adjusted downward by the corresponding adjusted value. After each adjustment, the third instances to be adjusted whose values ​​have reached the corresponding first performance threshold are excluded, and it is determined whether the parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold at this time. If it is less than the cluster performance threshold, the first cluster is monitored and analyzed normally. Otherwise, the third instances to be adjusted whose parameter values ​​of the first performance parameters are still greater than the cluster performance threshold continue to be back pressured according to the corresponding adjusted value until the parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the parameter values ​​of the second performance parameters of all the third instances to be adjusted are the corresponding first performance threshold.

[0107] Based on the foregoing embodiment, in other embodiments of the present application, the method further includes:

[0108] If the second performance limit value is less than the first reference value, the preset configuration of the fourth instance to be adjusted belonging to the shared cloud storage instance type in the second cloud storage instance set is used to perform back pressure processing on the corresponding fourth instance to be adjusted until the fifth parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each fourth instance to be adjusted is the corresponding first performance threshold.

[0109] In an embodiment of the present application, when the third parameter value is less than the cluster performance threshold, the preset configuration of the fourth instance to be adjusted that belongs to the shared cloud storage smooth type in the second cloud storage instance set is the corresponding first performance threshold, that is, the corresponding fourth instance to be adjusted is back-pressured according to the first performance threshold of the fourth instance to be adjusted until the first cluster returns to a normal level, or the second performance parameter of the fourth instance to be adjusted is the corresponding first performance threshold.

[0110] Based on the foregoing embodiments, in other embodiments of the present application, a first performance threshold of a cloud storage instance of an exclusive cloud storage instance type is determined based on a configured second performance parameter, and the capacity of the cloud storage instance of the exclusive cloud storage instance type includes at least a first capacity and a second capacity related to the second performance parameter, and the second capacity has an associated relationship with the first capacity; the cloud storage instance of the shared cloud storage instance type is configured to provide a minimum performance value corresponding to the second performance parameter.

[0111] In an embodiment of the present application, the first capacity is the actual capacity applied for by the cloud storage instance type of the exclusive cloud storage instance type, that is, the actual size of the cloud storage instance of the cloud storage instance type purchased by the user, and the second capacity is the fluctuating capacity provided by the instance developer for the cloud storage instance of the cloud storage instance type in order to ensure the reliability of service provision.

[0112] Based on the above embodiments, the embodiments of the present application provide a cluster management method, which is mainly used to realize intelligent hierarchical flow control of block storage. First, the corresponding cluster includes the following designs: two specifications, IOPS exclusive type and IOPS shared type, are defined based on each cloud hard disk type. Correspondingly, the performance indicators, capacity technical indicators, billing metering formulas, etc. of the cloud hard disk instance of the new cloud hard disk specification system based on cluster performance can refer to the following description, which at least includes:

[0113] Cloud disk type: The cloud disk type is related to the underlying architecture of the block storage cluster, data redundancy algorithm, hardware type, and model. It usually determines the performance limit of the cloud disk instance, such as single IO latency and the IOPS limit that a single instance can support.

[0114] The cloud hard disk specifications include at least two specifications: IOPS exclusive and IOPS shared. Among them: IOPS exclusive cloud hard disk instances mainly make SLA commitments based on the ordered IOPS values, charge according to the ordered IOPS volume, and are equipped with a certain amount of free capacity, where, for example, the ordering step can be 100 IOPS. The free capacity equipped with IOPS exclusive cloud hard disk instances is related to the IOPS ordering volume and has an upper limit on the free capacity. IOPS shared cloud hard disk instances promise a minimum IOPS performance value and charge according to the ordered capacity, where the ordering step can be in gigabit (GB).

[0115] Minimum capacity of a single disk: The minimum capacity of a single cloud disk instance supported by the cluster. For exclusive IOPS, this value can be set as a fixed value, which can be determined based on actual conditions. For shared IOPS, this value can be set as the minimum capacity to be ordered for a single disk, which can be determined based on actual conditions.

[0116] Maximum capacity of a single disk: The maximum capacity of a single cloud disk instance supported by the cluster. The exclusive IOPS type can be set to a fixed value, which can be determined based on actual conditions. The shared IOPS type can be set to the maximum orderable capacity of a single disk, which can be determined based on actual conditions.

[0117] The minimum IOPS per disk for an IOPS exclusive cloud disk instance refers to the minimum IOPS per disk instance that the cluster can support when the IOPS exclusive cloud disk instance is of a certain size, such as 4 kilobytes (KB). It should be noted that in order to effectively distinguish between IOPS exclusive cloud disk instances and IOPS shared cloud disk instances, under normal circumstances, the minimum IOPS subscription for a single disk of an IOPS exclusive cloud disk instance should be at least greater than the IOPS SLA value for an IO shared cloud disk. The specific amount can be determined based on actual conditions.

[0118] The maximum IOPS per disk for IOPS-only cloud disk instances refers to the maximum IOPS per cloud disk instance that the cluster can support when the block size is a certain size, such as 4KB. The specific IOPS can be determined based on actual conditions.

[0119] The following are some definitions of terms used in the embodiments of this application, including:

[0120] Burst IO limit: Based on the IOPS promised by the SLA, the cluster allows the cloud disk instance to perform IOPS bursts within a certain time and within a certain traffic range. This indicator is the burst IO limit value of a single cloud disk instance supported by the cluster. The burst IO mechanism is related to the current operation status and performance margin of the cluster. The burst IO limit is a fixed value and can be determined based on actual conditions.

[0121] Maximum throughput of a single disk: The maximum storage bandwidth of a single cloud disk instance supported by the cluster. The maximum throughput of a single disk is a fixed value and can be determined based on actual conditions. The unit can be expressed in megabits per second (MB / s), for example.

[0122] Single-path random write average latency: The average latency of a single-path random IO of a cloud disk instance when the storage block size is a certain size, such as 4KB. This indicator is usually determined by the system architecture, network capabilities, and underlying storage hardware performance of the cluster. It can be determined based on actual conditions.

[0123] Single-disk IOPS SLA commitment value: When the storage block size is a certain size, such as 4KB block size, the IOPS value committed by the customer's cloud disk instance SLA is usually between the minimum IOPS of a cluster single disk and the maximum IOPS of a cluster single disk. The IOPS SLA commitment value of a single-disk exclusive IOPS is the customer's IOPS ordering quantity. In some application scenarios, it is generally recommended to order in steps of 100 IOPS; the IOPS SLA commitment value of a single-disk shared IOPS is a fixed value, determined based on factors such as cluster performance, capacity, and market demand. In order to mitigate the risk of cluster performance explosion, guide customers to order IOPS exclusive cloud disk instances. It is recommended that the IOPS SLA commitment value of a single-disk shared IOPS be set lower than the minimum IOPS of a single disk of the IOPS exclusive specification.

[0124] Capacity calculation formula: The capacity of the customer's cloud disk instance, usually in GB. IOPS sharing is related to the customer's instance subscription capacity. The capacity of the IOPS exclusive cloud disk is divided into free capacity and additional subscription capacity. The free capacity is related to the IOPS volume subscribed by the customer instance. Correspondingly, the capacity of the IOPS exclusive cloud disk can be calculated and determined using the following formula: IOPS exclusive cloud disk instance capacity = min{minimum capacity of a single disk + capacity factor*(IOPS subscription volume - minimum IOPS of a single disk), free capacity upper limit}+[extra subscription capacity].

[0125] Burst IO size: Based on the IOPS SLA commitment value, the upper limit of IOPS that can be reached by the cloud disk instance within a certain time and a certain traffic size range. IOPS exclusive type can be set gradually within the burst IO upper limit range according to the IOPS subscription amount; IOPS shared cloud disk instance can be set gradually within the burst IO upper limit range according to the capacity subscription amount. For example, it can be recorded in the following table:

[0126] IOPS subscription / capacity subscription Burst IO size Order Quantity < Value 1 Burst value 1 Order Quantity∈[value1,value2) Burst value 2 Order Quantity∈[value2,value3) Burst value 3 ... ...

[0127] Maximum throughput of a single disk: the maximum storage bandwidth of a cloud disk instance.

[0128] (1) The maximum throughput of a single disk of an IOPS-only cloud disk instance is related to the IOPS subscription quantity and can be determined using the following calculation formula: Maximum throughput of a single disk of an IOPS-only cloud disk instance = min{basic value+throughput coefficient*IOPS subscription quantity, maximum throughput of a single disk}MBps.

[0129] (2) The maximum throughput of a single disk of an IOPS shared cloud disk instance is related to the IOPS capacity subscription, and can be determined by the following calculation formula: The maximum throughput of a single disk of an IOPS shared cloud disk instance = min{basic value+throughput coefficient*capacity, maximum throughput of a single disk}MBps.

[0130] Cluster sold-out mechanism: (1) The cluster’s saleable capacity is 0. (2) The cluster’s saleable IOPS is 0. It should be noted that the burst IO value allowed by the cloud disk instance is not counted as sold IOPS.

[0131] Correspondingly, an intelligent hierarchical flow control solution adapted to the cloud disk specification system defined in this proposal can be summarized as follows: In order to improve the market competitiveness of the product, the cloud disk instance is equipped with a burst IO function, which allows the cloud disk instance to run within a certain range and exceed the IOPS SLA commitment value. The upper limit of the burst IO can be set to the IOPS SLA commitment value or higher under the same type of conditions. In order to avoid the cluster performance explosion problem in the scenario of centralized redemption of cluster cloud disk instance performance, maintain cluster stability, and ensure the SLA indicators promised by the cloud disk instance to customers, a corresponding implementation process can be as follows Figure 2 As shown, the specific implementation steps include the following:

[0132] Step a11, start.

[0133] Step a12: Obtain cluster performance parameters.

[0134] The acquisition of cluster performance parameters includes cluster performance analysis or cloud disk instance performance analysis corresponding to the cluster.

[0135] Step a13: determine whether the cluster exceeds the limit based on the cluster performance parameters. If not, execute step a14; if so, execute step a29.

[0136] When monitoring the cluster status, you can use cloud logs, cloud monitoring and other technologies to collect cluster performance and cloud disk instance performance information, and use big data analysis, intelligent algorithms and other technologies to achieve situational awareness analysis of cluster operation and determine whether the cluster is overloaded. Among them, factors such as the information collection and reporting time interval, big data batch processing time, and algorithm models will affect the accuracy of intelligent flow control.

[0137] Step a14: determine whether the intelligent flow control function should be terminated or continued according to the intelligent flow control operating conditions. If terminated, execute step a28; otherwise, execute step a12.

[0138] Step a15: determine whether the cluster performance over-limit value is greater than or equal to the total cluster burst IO. If the cluster performance over-limit value is greater than or equal to the total cluster burst IO, execute step a16; otherwise, execute step a17.

[0139] The total cluster burst IO is the cumulative sum of the allowed burst IO of all cloud disk instances corresponding to the cluster.

[0140] Step a16: Disable the cluster IO burst function, call the cloud disk instance migration function, migrate the cloud disk instance to the remaining low-water-level clusters, balance the cluster load, and then execute step a12.

[0141] Step a17: From all cloud disks corresponding to the cluster, select cloud disk instances whose real-time IOPS exceeds the corresponding SLA commitment value to form a flow control list.

[0142] Step a18: Determine whether the flow control list includes the performance-surge cloud hard disk instance. If it does, execute step a19; otherwise, execute step a24.

[0143] Step a19, determine whether the performance surge value of the performance surge cloud hard disk instance is greater than or equal to the difference between the real-time IOPS of the performance surge cloud hard disk instance and the corresponding SLA commitment value. If it is greater than or equal to the difference, execute step a20; otherwise, execute step a22.

[0144] The performance burst value of the performance burst cloud disk instance is the change between the current real-time IOPS of the performance burst cloud disk instance and the IOPS collected in the previous collection period.

[0145] Step a20: determine whether the performance burst value of the instance is greater than or equal to the cluster performance limit value. If so, execute step a21; otherwise, execute step a24.

[0146] Step a21: Perform IOPS back pressure on the performance-surge cloud hard disk instance with the performance-surge value of the performance-surge cloud hard disk instance.

[0147] Step a22: limit the IOPS of the performance burst instance to the SLA committed value.

[0148] Among them, when the performance burst value of the performance burst cloud hard disk instance is greater than or equal to the difference between the real-time IOPS of the performance burst cloud hard disk instance and the corresponding SLA commitment value, and when the performance burst value of the performance burst cloud hard disk instance is greater than or equal to the cluster performance limit value, the IOPS of the performance burst cloud hard disk instance is limited to the SLA commitment value.

[0149] Step a23: determine whether the cluster performance level has returned to normal. If so, execute step a12; if not, execute step a24.

[0150] Step a24: Determine whether the cluster performance limit value is greater than or equal to the total burst IO of all IOPS shared cloud disk instances. If the cluster performance limit value is greater than or equal to the total burst IO of all IOPS shared cloud disk instances, execute step a25; otherwise, execute step a28.

[0151] Step a25: Disable the burst IO function of all IOPS shared cloud disk instances.

[0152] Step a26: Determine whether the cluster performance level has returned to normal. If not, execute step a27; if so, execute step a12.

[0153] Step a27: Perform IOPS back pressure on all IOPS exclusive cloud disk instances in the flow control list in proportion to the IOPS subscription amount, up to the SLA commitment value.

[0154] In this way, hierarchical flow control is performed for IOPS-exclusive cloud hard disks and IOPS-shared cloud hard disks in the same cluster, ensuring the flow control efficiency of the cluster, effectively reducing the risk of cluster performance explosion, and improving customer satisfaction.

[0155] Step a28: Perform IOPS back pressure on all IOPS shared cloud disk instances in the flow control list in proportion to the capacity subscription amount, up to the SLA commitment value.

[0156] Step a29, end.

[0157] Through the above intelligent flow control solution, hierarchical flow control can be performed for IOPS-exclusive cloud hard disks and IOPS-shared cloud hard disks in the same cluster to improve customer satisfaction.

[0158] The cluster management method provided in the embodiment of the present application monitors the first cluster, obtains the first parameter value of the first performance parameter, determines the second parameter value of the first cloud storage instance set corresponding to the first cluster, and performs resource management on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value. In this way, by comparing and analyzing the first parameter value of the first performance parameter of the first cluster with the second parameter value, the first cluster and the first cloud storage instance set are managed, which solves the problem that the storage performance of the cloud hard disk instance is reduced due to speed limiting of the cloud hard disk instance. In the application process, the QOS flow of the cluster is dynamically controlled to ensure the storage performance of the cloud hard disk instance and improve the service performance of the cloud hard disk instance.

[0159] Based on the above embodiments, the embodiments of the present application provide a cluster management device, which can be applied to Figure 1 In the cluster management method provided in the corresponding embodiment, refer to Figure 3 As shown, the cluster management device 2 may include: a monitoring unit 21, a determining unit 22 and a management unit 23; wherein:

[0160] A monitoring unit 21, configured to monitor the first cluster and obtain a first parameter value of a first performance parameter;

[0161] The determining unit 22 is used to determine a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster; wherein the first cloud storage instance set includes cloud storage instances of exclusive cloud storage instance types and cloud storage instances of shared cloud storage instance types;

[0162] The management unit 23 is used to manage the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value.

[0163] In other embodiments of the present application, the management unit includes: a first determination module, a closing module and a migration module; wherein:

[0164] A first determination module, configured to determine a first performance over-limit value of a first cluster based on a first parameter value and a cluster performance threshold;

[0165] A shut-down module, configured to shut down a burst management function of the first cluster if the first performance over-limit value is greater than or equal to a preset performance parameter;

[0166] The first determination module is further used to determine the cloud storage instance to be migrated from the first cloud storage instance set;

[0167] The first determination module is further used to determine a target cluster from the cluster system to which the first cluster belongs;

[0168] The migration module is used to migrate the cloud storage instance to be migrated to the target cluster.

[0169] In other embodiments of the present application, the management unit further includes: a second determination module and a management module; wherein:

[0170] A second determination module is used to determine, from the first cloud storage instance set, cloud storage instances whose parameter values ​​of the second performance parameters are greater than or equal to the first performance threshold value if the first performance limit value is less than the second parameter value, to obtain a second cloud storage instance set;

[0171] The second determination module is further configured to determine the sum of the third parameter values ​​of the preset performance parameters of all cloud storage instances of the shared cloud storage instance type in the first cloud storage instance set to obtain the first reference value if the second cloud storage instance set does not include the performance surge cloud storage instance;

[0172] The management module is used to manage the cloud storage instances in the second cloud storage instance set based on the first performance limit value and the first reference value.

[0173] In other embodiments of the present application, the management unit further includes: a management module; wherein:

[0174] The second determination module is further configured to determine a parameter burst increment of a second performance parameter of each instance in the second cloud storage instance set if the second cloud storage instance set includes a performance burst cloud storage instance;

[0175] An adjustment module is used to perform back pressure processing on the second performance parameter of the first instance to be adjusted based on the parameter sudden increment of the first instance to be adjusted, if there is a first instance to be adjusted in the second cloud storage instance set, whose parameter sudden increment is less than the second reference value of the corresponding instance and whose parameter sudden increment is greater than the first performance overlimit value; wherein the second reference value is the difference between the current parameter value of the second performance parameter of the corresponding instance and the corresponding first performance threshold.

[0176] In other embodiments of the present application, the adjustment module is also used to limit the second performance parameter of the second instance to be adjusted to the corresponding first performance threshold if there is a second instance to be adjusted in the second cloud storage instance set whose parameter sudden increment is greater than or equal to the second reference value of the corresponding instance.

[0177] In other embodiments of the present application,

[0178] The monitoring unit is further used to monitor a third parameter value of the first performance parameter of the first cluster;

[0179] The determining unit is further used to determine a second performance over-limit value of the first cluster based on the third parameter value and the cluster performance threshold;

[0180] The determination unit is further used to determine the sum of the third parameter values ​​of the preset performance parameters of all cloud storage instances belonging to the shared cloud storage instance type in the first cloud storage instance set to obtain a first reference value;

[0181] The management unit is further used to manage the second cloud storage instance set based on the second performance limit value and the first reference value.

[0182] In other embodiments of the present application, when the management unit executes the step of managing the second cloud storage instance set based on the second performance limit value and the first reference value, it can be specifically implemented by the following steps:

[0183] If the second performance limit-exceeding value is greater than or equal to the first reference value, shutting down the burst management function of all cloud storage instances of the shared cloud storage instance type in the second cloud storage instance set;

[0184] monitoring a fourth parameter value of the first performance parameter of the first cluster;

[0185] If the fourth parameter value is greater than or equal to the cluster performance threshold, all instances of the exclusive cloud storage instance type in the second cloud storage instance set are determined to obtain one or more third instances to be adjusted;

[0186] One or more third instances to be adjusted are managed.

[0187] In other embodiments of the present application, when the management unit executes the step of managing one or more third instances to be adjusted, it can be implemented by the following steps:

[0188] Determining a preset configuration of each third instance to be adjusted;

[0189] Based on the preset configuration of one or more third instances to be adjusted, back pressure processing is performed on the second performance parameter of the corresponding third instances to be adjusted until the third parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each third instance to be adjusted is the corresponding first performance threshold.

[0190] In other embodiments of the present application, the management unit is further configured to perform the following steps:

[0191] If the second performance limit value is less than the first reference value, the preset configuration of the fourth instance to be adjusted belonging to the shared cloud storage instance type in the second cloud storage instance set is used to perform back pressure processing on the corresponding fourth instance to be adjusted until the fifth parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each fourth instance to be adjusted is the corresponding first performance threshold.

[0192] In other embodiments of the present application, the first performance threshold of a cloud storage instance of an exclusive cloud storage instance type is determined based on a configured second performance parameter, the capacity of the cloud storage instance of the exclusive cloud storage instance type includes at least a first capacity and a second capacity related to the second performance parameter, and the second capacity has an associated relationship with the first capacity; the cloud storage instance of a shared cloud storage instance type is configured to provide a minimum performance value corresponding to the second performance parameter.

[0193] It should be noted that the process of information interaction between the units and modules in this embodiment can refer to the description in other embodiments and will not be repeated here.

[0194] The cluster management device provided in the embodiment of the present application monitors the first cluster, obtains the first parameter value of the first performance parameter, determines the second parameter value of the first cloud storage instance set corresponding to the first cluster, and performs resource management on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value. In this way, by comparing and analyzing the first parameter value of the first performance parameter of the first cluster with the second parameter value, the first cluster and the first cloud storage instance set are managed, which solves the problem that the storage performance of the cloud hard disk instance is reduced due to speed limiting of the cloud hard disk instance. In the application process, the QOS flow of the cluster is dynamically controlled to ensure the storage performance of the cloud hard disk instance and improve the service performance of the cloud hard disk instance.

[0195] Based on the above embodiments, the embodiments of the present application provide a cluster management device, which can be applied to Figure 1 In the cluster management method provided in the corresponding embodiment, refer to Figure 4 As shown, the cluster management device 3 may include: a communication interface 31, a memory 32, a processor 33 and a communication bus 34; wherein:

[0196] A memory 32, for storing executable information;

[0197] A communication bus 34, used to realize communication connection between the communication interface 31, the processor 33 and the memory 32;

[0198] The processor 33 is used to execute the cluster management program stored in the memory 32 to implement the following Figure 1 The implementation process of the cluster management method provided in the corresponding embodiment will not be repeated here.

[0199] Based on the foregoing embodiments, the embodiments of the present application provide a computer-readable storage medium, referred to as a storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement reference Figure 1 The implementation process of the cluster management method provided in the corresponding embodiment will not be repeated here.

[0200] Based on the above embodiments, the embodiments of the present application further provide a computer program product, including a computer program, which can be executed by the processor 33 of the cluster management device 3 to complete any of the above method steps.

[0201] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0202] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0203] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0204] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0205] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.

Claims

1. A cluster management method, characterized in that: The method comprises: monitoring the first cluster to obtain a first parameter value of a first performance parameter; Determine a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster; wherein the first cloud storage instance set includes cloud storage instances of exclusive cloud storage instance types and cloud storage instances of shared cloud storage instance types; Based on the first parameter value and the second parameter value, resource management is performed on the first cluster and the first cloud storage instance set.

2. The method according to claim 1, characterized in that The performing resource management on the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value includes: Determining a first performance over-limit value of the first cluster based on the first parameter value and a cluster performance threshold; If the first performance limit value is greater than or equal to the second parameter value, disabling the burst management function of the first cluster; Determining a cloud storage instance to be migrated from the first cloud storage instance set; Determining a target cluster from the cluster system to which the first cluster belongs; Migrate the cloud storage instance to be migrated to the target cluster.

3. The method according to claim 2, characterized in that The method further comprises: If the first performance over-limit value is less than the second parameter value, determine the cloud storage instances whose parameter values ​​of the second performance parameters are greater than or equal to the first performance threshold from the first cloud storage instance set, and obtain a second cloud storage instance set; If the second cloud storage instance set does not include a performance-surge cloud storage instance, determine a sum of third parameter values ​​of preset performance parameters of all cloud storage instances of the shared cloud storage instance type in the first cloud storage instance set to obtain a first reference value; Based on the first performance limit value and the first reference value, the cloud storage instances in the second cloud storage instance set are managed.

4. The method according to claim 3, characterized in that: The method further comprises: If the second cloud storage instance set includes a performance-surge cloud storage instance, determining a parameter surge amount of the second performance parameter of each instance in the second cloud storage instance set; If there is a first instance to be adjusted in the second cloud storage instance set in which the parameter sudden increment is less than the second reference value of the corresponding instance and the parameter sudden increment is greater than the first performance limit value, back pressure processing of the second performance parameter is performed on the first instance to be adjusted based on the parameter sudden increment of the first instance to be adjusted; wherein the second reference value is the difference between the current parameter value of the second performance parameter of the corresponding instance and the corresponding first performance threshold.

5. The method according to claim 4, characterized in that The method further comprises: If there is a second instance to be adjusted in the second cloud storage instance set whose parameter sudden increment is greater than or equal to the second reference value of the corresponding instance, the second performance parameter of the second instance to be adjusted is limited to the corresponding first performance threshold.

6. The method according to claim 5, characterized in that The method further comprises: monitoring a third parameter value of a first performance parameter of the first cluster; determining a second performance over-limit value of the first cluster based on the third parameter value and the cluster performance threshold; Determine the sum of all cloud storage instances in the first cloud storage instance set that belong to the shared cloud storage instance type to obtain a first reference value; The second cloud storage instance set is managed based on the second performance limit value and the first reference value.

7. The method according to claim 4, characterized in that The managing the second cloud storage instance set based on the second performance limit value and the first reference value includes: If the second performance limit value is greater than or equal to the first reference value, disabling the burst management function of all cloud storage instances of the shared cloud storage instance type in the second cloud storage instance set; monitoring a fourth parameter value of the first performance parameter of the first cluster; If the fourth parameter value is greater than or equal to the cluster performance threshold, determine all instances of the exclusive cloud storage instance type in the second cloud storage instance set to obtain one or more third instances to be adjusted; Managing one or more of the third instances to be adjusted.

8. The method according to claim 7, characterized in that The managing one or more third instances to be adjusted includes: Determining a preset configuration of each of the third instances to be adjusted; Based on the preset configuration of one or more of the third instances to be adjusted, back pressure processing is performed on the second performance parameter of the corresponding third instances to be adjusted until the third parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each of the third instances to be adjusted is the corresponding first performance threshold.

9. The method according to claim 7, characterized in that: The method further comprises: If the second performance limit value is less than the first reference value, the preset configuration of the fourth instance to be adjusted belonging to the shared cloud storage instance type in the second cloud storage instance set is used to perform back pressure processing on the corresponding fourth instance to be adjusted until the fifth parameter value of the first performance parameter of the first cluster is less than the cluster performance threshold, or the value of the second performance parameter of each of the fourth instance to be adjusted is the corresponding first performance threshold.

10. The method according to claim 1, characterized in that The first performance threshold of the cloud storage instance of the exclusive cloud storage instance type is determined based on the configured second performance parameter. The capacity of the cloud storage instance of the exclusive cloud storage instance type includes at least a first capacity and a second capacity related to the second performance parameter, and the second capacity has an associated relationship with the first capacity. The cloud storage instance of the shared cloud storage instance type is configured to provide a minimum performance value corresponding to the second performance parameter.

11. A cluster management device, characterized in that: The device at least comprises: a monitoring unit, a determination unit and a management unit; wherein: The monitoring unit is used to monitor the first cluster and obtain a first parameter value of a first performance parameter; The determining unit is used to determine a second parameter value of a preset performance parameter of a first cloud storage instance set corresponding to the first cluster; wherein the first cloud storage instance set includes cloud storage instances of exclusive cloud storage instance types and cloud storage instances of shared cloud storage instance types; The management unit is used to manage the first cluster and the first cloud storage instance set based on the first parameter value and the second parameter value.

12. A cluster management device, characterized in that: The device at least includes: a communication interface, a memory, a processor and a communication bus; wherein: The memory is used to store executable instructions; The communication bus is used to realize the communication connection between the communication interface, the processor and the memory; The processor is used to execute the cluster management program stored in the memory to implement the steps of the cluster management method according to any one of claims 1 to 10.

13. A storage medium, characterized in that: The storage medium stores a cluster management program, which is used to implement the steps of the cluster management method according to any one of claims 1 to 10 when executed.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the cluster management method according to any one of claims 1 to 10.