Cloud disk scheduling method, cloud server, storage medium and program product
Patent Information
- Application Number
- CN202510237968.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]本申请提供一种云盘调度方法、云服务器、存储介质及程序产品,用以解决集群发生故障的概率高,影响集群稳定性的问题
[0012] The cloud disk scheduling method, cloud server, storage medium, and program product provided in this application divide cloud disks within a cluster into cloud disk groups according to their performance modes, ensuring that cloud disks within the same cloud disk group have the same or similar performance modes. When traffic risk is detected in any cluster (referred to as the first cluster), a cloud disk scheduling plan that satisfies the cloud disk group constraints is generated based on the traffic information of each cluster. According to the cloud disk scheduling plan, the target cloud disk is migrated to the target cloud disk group within the corresponding target cluster, achieving elastic scaling of the cloud disk and improving cluster stability. By introducing cloud disk group constraints, the target cloud disks to be scheduled are distributed as widely as possible across multiple cloud disk groups within the first cluster. This avoids scheduling multiple cloud disks with the same or similar performance modes within a single cloud disk group, reducing the probability of traffic risk arising from the target cloud disk being embedded in the target cluster, thereby reducing the probability of cluster failure and improving cluster stability.
Smart Images

Figure CN122661284A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and more particularly to a cloud disk scheduling method, a cloud server, a storage medium, and a program product. Background Technology
[0002] With the development of internet technology and the arrival of the big data era, the amount of data generated by enterprises and individuals is growing exponentially. This massive amount of data needs to be effectively stored and managed to support various business needs. Storing data in the cloud on block devices (i.e., cloud disks) has become a trend. Effective resource scheduling can improve the utilization rate of existing infrastructure, thereby reducing the total cost of ownership of cloud storage.
[0003] However, different users use cloud drives in different ways, and their usage patterns vary greatly. Multiple cloud drives within a cluster may experience synchronous high loads or resource contention (i.e., cloud drive resonance), leading to a surge in cluster traffic or even overload. This increases the probability of cluster failure and affects cluster stability. Summary of the Invention
[0004] This application provides a cloud disk scheduling method, a cloud server, a storage medium, and a program product to solve the problem of high probability of cluster failure and its impact on cluster stability.
[0005] Firstly, this application provides a cloud disk scheduling method, including:
[0006] Monitor the traffic information of each cluster, wherein any cluster is divided into at least one cloud disk group, and the cloud disks in the same cloud disk group have the same or similar performance modes.
[0007] If the traffic information of each cluster indicates that the first cluster has a traffic risk, a cloud disk scheduling plan for the first cluster is generated based on the traffic information of each cluster. The cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group where the target cloud disk is located, and the second cluster and the second cloud disk group to which the target cloud disk is migrated. The cloud disk scheduling plan satisfies the cloud disk group constraint, which includes the target cloud disk being distributed in multiple cloud disk groups of the first cluster.
[0008] According to the cloud disk scheduling plan, the target cloud disk is migrated to the second cloud disk group within the corresponding second cluster.
[0009] In a second aspect, this application provides a cloud server, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the cloud server to perform the method provided in any of the foregoing aspects.
[0010] Thirdly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method provided in any of the foregoing aspects.
[0011] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods provided in any of the foregoing aspects.
[0012] The cloud disk scheduling method, cloud server, storage medium, and program product provided in this application divide cloud disks within a cluster into cloud disk groups according to their performance modes, ensuring that cloud disks within the same cloud disk group have the same or similar performance modes. When traffic risk is detected in any cluster (referred to as the first cluster), a cloud disk scheduling plan that satisfies the cloud disk group constraints is generated based on the traffic information of each cluster. According to the cloud disk scheduling plan, the target cloud disk is migrated to the target cloud disk group within the corresponding target cluster, achieving elastic scaling of the cloud disk and improving cluster stability. By introducing cloud disk group constraints, the target cloud disks to be scheduled are distributed as widely as possible across multiple cloud disk groups within the first cluster. This avoids scheduling multiple cloud disks with the same or similar performance modes within a single cloud disk group, reducing the probability of traffic risk arising from the target cloud disk being embedded in the target cluster, thereby reducing the probability of cluster failure and improving cluster stability. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0014] Figure 1 A flowchart illustrating a cloud disk scheduling method provided in an exemplary embodiment of this application;
[0015] Figure 2 A flowchart illustrating a method for grouping cloud disks as provided in an exemplary embodiment of this application;
[0016] Figure 3 A general architecture diagram of cloud disk scheduling provided for an exemplary embodiment of this application;
[0017] Figure 4 A general flowchart of cloud disk scheduling provided for an exemplary embodiment of this application;
[0018] Figure 5 A detailed flowchart of cloud disk scheduling provided for another exemplary embodiment of this application;
[0019] Figure 6 A detailed flowchart of cloud disk scheduling provided for another exemplary embodiment of this application;
[0020] Figure 7 This is a schematic diagram of the structure of a cloud server provided in an embodiment of this application.
[0021] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0023] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0024] First, let me explain the terms used in this application:
[0025] Cloud disk: another name for block devices in the public cloud.
[0026] Cloud disk group: Group cloud disks with the same or similar performance patterns (IO patterns) together.
[0027] Resonance: Same usage mode.
[0028] Anomaly scheduling: In public cloud scenarios of cloud computing, it is a type of adaptive resource scheduling performed to alleviate the situation of unreasonable resource allocation.
[0029] Throughput: Bytes per second, measured in MBps, used to measure the performance of cloud disks / clusters.
[0030] Mean Squared Error (MSE) is a measure of the difference between the estimated and the estimated. MSE is the average of the sum of squared errors and is used to evaluate the accuracy of a predictive model. A smaller MSE indicates higher accuracy of the predictive model.
[0031] Performance Pattern (IO Pattern): Also known as I / O pattern, it refers to the data access pattern on a storage system (such as a cloud disk). It describes the regularity or characteristics of data during read and write processes. IO pattern has a significant impact on cloud disk performance; different IO patterns will lead to different performance indicators such as read / write speed and throughput.
[0032] To address the issue of cluster failures and instability caused by synchronized high loads or resource contention among multiple cloud disks within a cluster (i.e., cloud disk resonance), which can lead to surges in cluster traffic or even overload, this application provides a cloud disk scheduling method. This method divides cloud disks within a cluster into cloud disk groups based on performance modes, ensuring that cloud disks within the same group have similar performance modes. When traffic risk is detected in any cluster (referred to as the first cluster), a cloud disk scheduling plan that satisfies the constraints of the cloud disk groups is generated based on the traffic information of each cluster. According to the scheduling plan, the target cloud disk is migrated to the corresponding target cloud disk group within the target cluster, achieving elastic scaling of the cloud disks and improving cluster stability. By introducing cloud disk group constraints, the target cloud disks to be scheduled are distributed as widely as possible across multiple cloud disk groups within the first cluster. This avoids scheduling multiple cloud disks with the same or similar performance modes within a single cloud disk group, reducing the probability of traffic risk arising from the target cloud disk being embedded in the target cluster, thereby reducing the probability of cluster failures and improving cluster stability.
[0033] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0034] Figure 1 This is a flowchart illustrating a cloud disk scheduling method provided in an exemplary embodiment of this application. The executing entity in this embodiment is a cloud server, such as a cloud disk scheduling server, used for cloud disk scheduling. Figure 1 As shown, the specific steps of this method are as follows:
[0035] Step S101: Monitor the traffic information of each cluster, wherein any cluster is divided into at least one cloud disk group, and the cloud disks in the same cloud disk group have the same or similar performance modes.
[0036] Among them, performance pattern (IO pattern), also known as I / O pattern, refers to the access pattern of data on a storage system (such as a cloud disk). It describes the regularity or characteristics of data during the read and write process. IO pattern has a significant impact on the performance of cloud disks, and different IO patterns will lead to different performance indicators such as read and write speed and throughput.
[0037] In this embodiment, the cloud disks in each cluster are divided into multiple cloud disk groups, so that the cloud disks in the same cloud disk group have the same or similar performance modes.
[0038] The cloud server can monitor the traffic information of each cluster. The traffic information of any cluster includes, but is not limited to, the total traffic of the cluster, the traffic of each cloud disk group within the cluster, and the traffic of each cloud disk within the cluster.
[0039] Step S102: If the traffic information of each cluster indicates that the first cluster has a traffic risk, then a cloud disk scheduling plan for the first cluster is generated based on the traffic information of each cluster. The cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group where the target cloud disk is located, and the second cluster and the second cloud disk group to which the target cloud disk is migrated. The cloud disk scheduling plan satisfies the cloud disk group constraints, which include distributing the target cloud disk as much as possible among multiple cloud disk groups in the first cluster.
[0040] The cloud server can determine whether there is traffic risk in each cluster based on the traffic information of each cluster. When any cluster (referred to as the first cluster) is determined to have traffic risk, cloud disk scheduling for the first cluster is triggered.
[0041] For example, based on the total traffic of any cluster, it is determined whether the cluster meets the cluster traffic limit conditions. The first cluster that does not meet the cluster traffic limit conditions is identified as a cluster with traffic risk. The cluster traffic limit conditions include: the total traffic reaches the cluster traffic water level, and / or, the difference between the total traffic and the average total traffic of clusters within the same availability zone meets a first difference condition. An availability zone typically includes multiple clusters. Clusters within the same availability zone can perform cloud disk scheduling with each other (except where constraints prohibit mutual scheduling), while cloud disk scheduling is not allowed between different availability zones. The availability zone settings can be configured according to actual application needs and are not specifically limited here.
[0042] The cluster traffic limits can differ for different clusters. The cluster traffic level for any given cluster can be determined based on the cluster's rated traffic.
[0043] For example, the cluster flow rate level can be the first percentage of the cluster's rated flow rate, such as 80% or 70%. The first percentage can be set and adjusted according to actual application needs, and no specific limitation is made here.
[0044] For example, the cluster flow level can be the cluster's rated flow minus the first flow difference. The first flow difference can be set and adjusted according to actual application needs, and no specific limitation is made here.
[0045] The first difference condition in the cluster traffic limit conditions for any cluster can be: the difference is less than 0 and the absolute value of the difference is greater than a first preset difference. The first preset difference can be set by the rated traffic of the cloud disk group; no specific limitation is made here.
[0046] When a first cluster with potential traffic risks is detected, and cloud disk scheduling is performed for that cluster, the cloud server generates a cloud disk scheduling plan for the first cluster based on the traffic information of each cluster and the set cloud disk group constraints. This cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group where the target cloud disk resides, and the second cluster and second cloud disk group to which the target cloud disk will be migrated. The generated cloud disk scheduling plan satisfies the cloud disk group constraints.
[0047] The cloud disk group constraint aims to distribute the target cloud disks as widely as possible across multiple cloud disk groups within the first cluster. By introducing this constraint, the target cloud disks to be scheduled are distributed across multiple cloud disk groups within the first cluster. This avoids scheduling multiple cloud disks with the same or similar performance modes within a single cloud disk group, thereby reducing the probability of cloud disk resonance within the cluster, lowering the probability of cluster failure, and improving cluster stability.
[0048] Optionally, when generating a cloud disk scheduling plan for any cluster, other constraints besides the aforementioned cloud disk group constraints can be incorporated, including but not limited to blacklist constraints, cluster traffic constraints, cluster capacity constraints, and cluster characteristic constraints. These constraints, together with the aforementioned cloud disk group constraints, constitute the constraints used during cloud disk scheduling, i.e., cloud disk scheduling constraints.
[0049] When generating the cloud disk scheduling plan for the first cluster based on the traffic information of each cluster, the cloud server generates the cloud disk scheduling plan for the first cluster based on the traffic information of each cluster and the set cloud disk scheduling constraints. Among them, the cloud disk scheduling constraints include at least cloud disk group constraints.
[0050] Optionally, in addition to the aforementioned cloud disk group constraints, the cloud disk group constraints may further include the following constraint: after migrating the target cloud disk to the second cloud disk group, the second cloud disk group meets the cloud disk group traffic limit conditions. This constraint ensures that after migrating the target cloud disk to the second cloud disk group, the second cloud disk group meets the cloud disk group traffic limit conditions, thus avoiding traffic risks to the second cloud disk group due to cloud disk scheduling. By introducing cloud disk group constraints, the scope of traffic failures is narrowed from the cluster granularity to the cloud disk group granularity, reducing the impact of failures.
[0051] The cloud disk group traffic limit conditions may include: the traffic reaches the preset group traffic level, and / or the difference between the traffic and the average traffic of the cloud disk group in the cluster meets the second difference condition.
[0052] Different cloud disk groups can have different traffic limits. The preset traffic threshold for any cloud disk group can be automatically updated based on the proportion of that cloud disk group's traffic to the total traffic of the cluster.
[0053] For example, the preset group traffic threshold can be a second percentage of the rated traffic of the cluster to which the cloud disk group belongs. This second percentage can be determined based on the proportion of the cloud disk group's traffic to the total traffic of the cluster. For instance, if a cloud disk group's traffic accounts for 80% of the total traffic of its cluster, then the preset group traffic threshold for that cloud disk group can be 80% of the rated traffic of the cluster. The second difference condition in the cloud disk group traffic limit conditions for any cloud disk group can be: the difference is less than 0 and the absolute value of the difference is greater than the second preset difference. The second preset difference can be set based on the rated traffic of the cloud disk group, and is not specifically limited here.
[0054] Optionally, in addition to the cloud disk group constraints mentioned above, the cloud disk scheduling constraints may also include at least one of the following: blacklist constraints, cluster traffic constraints, and cluster capacity constraints.
[0055] The blacklist constraints include: the second cluster is not on the blacklist of the first cluster that cannot schedule with each other, and the second cloud disk group is not on the blacklist of the first cloud disk group that cannot schedule with each other.
[0056] For example, clusters 1 and 2 are non-user-specific storage domains currently shared by users A and B. Cluster 3 is user C's dedicated storage domain. In this example scenario, clusters 1 and 2 can be added to a blacklist of clusters 3 that cannot be scheduled between them, ensuring that user C's dedicated cloud disk in cluster 3 will not be migrated to clusters 1 or 2.
[0057] By setting blacklists that prevent inter-cluster scheduling and blacklists that prevent inter-disk scheduling between different cloud disk groups, it is ensured that cloud disks in any cluster cannot be migrated to clusters on the blacklist of non-inter-cluster scheduling, and cloud disks in any cloud disk group cannot be migrated to cloud disk groups on the blacklist of non-inter-cluster scheduling. The blacklists for each cluster and each cloud disk group can be set and adjusted according to actual application needs, and no specific limitations are set here.
[0058] The cluster traffic constraint includes: after migrating the target cloud disk to the second cloud disk group of the second cluster, the second cluster meets the cluster traffic constraint conditions. These cluster traffic constraint conditions include: the total traffic reaches the cluster traffic water level, and / or, the difference between the total traffic and the average total traffic of the clusters within it satisfies the first difference condition.
[0059] By setting cluster traffic constraints, it can be ensured that after the target cloud disk is migrated to the second cluster, the second cluster meets the cluster traffic limit conditions, thus avoiding traffic risks to the second cluster due to cloud disk scheduling.
[0060] The cluster capacity constraint includes: after migrating the target cloud disk to the second cloud disk group of the second cluster, the second cluster must meet the cluster capacity limit conditions. Specifically, the cluster capacity limit conditions include: the cluster capacity must reach the cluster capacity water level.
[0061] The capacity water level lines for different clusters can be different. The capacity water level line for any cluster can be determined based on its rated capacity. For example, the capacity water level line for a cluster can be a third percentage of its rated capacity, such as 80% or 70%. This third percentage can be set and adjusted according to actual application needs, and is not specifically limited here. Alternatively, the capacity water level line for a cluster can be the rated capacity minus the capacity difference. This capacity difference can be set and adjusted according to actual application needs, and is not specifically limited here.
[0062] Optionally, in addition to the aforementioned cloud disk scheduling constraints, the cloud disk scheduling constraints also include: cluster characteristic constraints. Cluster characteristic constraints include one or more custom constraints. Cluster characteristic constraints can be set according to the usage characteristics / requirements of the cluster in the actual application scenario, and are not specifically limited here.
[0063] For example, in some scenario examples, some clusters can support cloud disks of specific IO types. Cluster characteristic constraints can include the following custom constraints: specifying that a cluster only allows the migration of cloud disks of specific IO types. Which clusters are included in the specified clusters can be set and adjusted according to actual application needs, and are not specifically limited here. IO types include, but are not limited to: random read, random write, sequential read, and sequential write.
[0064] For example, in some scenario examples, users or relevant operations and maintenance personnel can set a whitelist for specific clusters, which may include one or more clusters. During cloud disk scheduling, cloud disks in a specific cluster are only allowed to be migrated to clusters that are on the whitelist corresponding to that specific cluster. Cluster characteristic constraints may include the following custom constraint: cloud disks in a specific cluster are only allowed to be migrated to clusters that are on the whitelist corresponding to that specific cluster. The specific clusters included, and the whitelist for that specific cluster, can be set and adjusted according to actual application needs, and are not specifically limited here.
[0065] In this embodiment, a cloud disk group dimension is introduced during cloud disk scheduling. Cloud disks within the same group have the same or similar performance patterns (IOPattern), while cloud disks in different groups have different and dissimilar performance patterns (IOPattern). On one hand, cloud disk group constraints are incorporated into the cloud disk scheduling strategy, determining which cloud disks to schedule for computation is not solely based on cloud disk traffic, allowing for a more comprehensive consideration of the rationality and security of the cloud disk scheduling plan. On the other hand, it enables better monitoring of cluster stability, preventing cluster traffic risks caused by the resonance of cloud disks with the same performance pattern (IOPattern), thus improving cluster stability.
[0066] Step S103: According to the cloud disk scheduling plan, migrate the target cloud disk to the second cloud disk group in the corresponding second cluster.
[0067] After generating the cloud disk scheduling plan for the first cluster, the cloud disk scheduling plan is executed to migrate the target cloud disk to the second cloud disk group within the corresponding second cluster, thereby eliminating the traffic risk present in the first cluster, reducing the probability of cluster failure, and improving the stability of the cluster.
[0068] This embodiment divides cloud disks within a cluster into cloud disk groups based on performance modes, ensuring that cloud disks within the same group have the same or similar performance modes. When traffic risk is detected in any cluster (referred to as the first cluster), a cloud disk scheduling plan that satisfies the cloud disk group constraints is generated based on the traffic information of each cluster. According to the cloud disk scheduling plan, the target cloud disk is migrated to the corresponding target cloud disk group within the target cluster, achieving elastic scaling of the cloud disks and improving cluster stability. By introducing cloud disk group constraints, the target cloud disks to be scheduled are distributed as widely as possible across multiple cloud disk groups within the first cluster. This avoids scheduling multiple cloud disks with the same or similar performance modes within a single cloud disk group, reducing the probability of traffic risk arising from the target cloud disk being embedded in the target cluster, thereby reducing the probability of cluster failure and improving cluster stability.
[0069] In one optional embodiment, before scheduling cloud disks, the cloud server can divide the cloud disks in the same cluster into multiple cloud disk groups based on the attribute information and historical throughput information of the cloud disks in each cluster, so that the cloud disks in the same cloud disk group have the same or similar performance modes.
[0070] Figure 2 A flowchart illustrating a cloud disk grouping method provided for an exemplary embodiment of this application. Figure 2 As shown, the specific steps of this method are as follows:
[0071] Step S201: Obtain the attribute information and historical throughput information of the cloud disks in each cluster.
[0072] The attribute information of a cloud drive refers to its inherent characteristics, including but not limited to: the user to whom the cloud drive belongs, the size of the cloud drive, and the type of cloud drive. The user to whom the cloud drive belongs is used to distinguish cloud drives belonging to different users, as different users have different usage patterns. The type of cloud drive can be determined based on the classification of cloud drives in the actual application scenario; for example, the type of cloud drive may include, but is not limited to, system disks and data disks.
[0073] Historical throughput information of cloud disks refers to the throughput information of cloud disks within a historical period, which may include, but is not limited to: the average read throughput of cloud disks within the historical period, the peak read throughput of cloud disks within the historical period, the average write throughput of cloud disks within the historical period, the peak write throughput of cloud disks within the historical period, the average mixed IO throughput (i.e., read + write throughput) of cloud disks within the historical period, and the peak mixed IO throughput of cloud disks within the historical period.
[0074] The historical time period can be set according to actual application needs. For example, the historical time period can be the past day, the past several hours, etc. There is no specific limitation here.
[0075] In this step, the cloud server can obtain the attribute information and historical throughput information of the cloud disks in each cluster.
[0076] Step S202: Encode the attribute information and historical throughput information of the cloud disks in each cluster into vectors to obtain the feature vectors of each cloud disk in each cluster.
[0077] In this step, for any cloud disk in any cluster, the cloud server encodes the cloud disk's attribute information and historical throughput information into a vector, which serves as the feature vector of that cloud disk. This allows the acquisition of feature vectors for each cloud disk in each cluster.
[0078] When encoding the attribute information and historical throughput information of the cloud disk into a vector, the cloud server can use one-hot encoding or other similar encoding methods (such as encoding methods based on deep learning models). This embodiment does not make specific limitations here.
[0079] Step S203: Based on the feature vectors of each cloud disk in any cluster, divide the cloud disks in the cluster into multiple cloud disk groups. Each cloud disk group contains at least one cloud disk, and the cloud disks in the same cloud disk group have the same or similar performance modes.
[0080] In this step, for any given cluster, the cloud server clusters the feature vectors of each cloud disk within that cluster, resulting in multiple clusters. Each cluster contains the feature vector of at least one cloud disk. Feature vectors within the same cluster have high similarity, while feature vectors between different clusters have low similarity.
[0081] Each cluster corresponds to a cloud disk group. Cloud disks are assigned to the corresponding cloud disk groups based on the cluster in which their feature vectors belong. Since feature vectors within the same cluster have high similarity, while feature vectors in different clusters have low similarity, cloud disks within the same cloud disk group have the same or similar performance patterns, while cloud disks in different cloud disk groups have different and dissimilar performance patterns.
[0082] Optionally, when clustering the feature vectors of each cloud disk within the cluster, the cloud server can use the K-means clustering algorithm to cluster the feature vectors of each cloud disk within the cluster into K clusters. Each cluster corresponds to a cluster center (or cluster center), and each feature vector belongs to the cluster closest to the cluster center.
[0083] The value of K can be set and adjusted according to actual application needs and empirical values. Alternatively, experiments can be conducted by simulating different values of K multiple times, and the value with the smallest MSE of the cluster center can be selected as the value of K. No specific limitation is made here.
[0084] In addition, when clustering the feature vectors of each cloud disk in the cluster, the cloud server can use other clustering algorithms, such as hierarchical clustering algorithms, etc., without making specific restrictions here.
[0085] In another optional embodiment, the cloud server can further divide the cloud disks within the same cluster into multiple cloud disk groups based on the historical throughput information of the cloud disks within each cluster, so that the cloud disks within the same cloud disk group have the same or similar performance patterns. Specifically, the historical throughput information of the cloud disks within each cluster is encoded into vectors to obtain the feature vectors of each cloud disk in each cluster; based on the feature vectors of each cloud disk in any cluster, the cloud disks in that cluster are divided into multiple cloud disk groups, each cloud disk group containing at least one cloud disk, and the cloud disks within the same cloud disk group have the same or similar performance patterns. The specific implementation principle is similar to the aforementioned steps S202-S203, and will not be repeated here.
[0086] In one optional embodiment, after dividing the cloud disks within the cluster into multiple cloud disk groups, the cloud server can visualize the scope of each cloud disk group (i.e., which cloud disks are included) and the number of cloud disks within each group. During cloud disk scheduling, when changes occur within a cloud disk group, the visualized content is updated promptly, allowing relevant technical personnel to intuitively see the process and results of cloud disk scheduling, thereby optimizing cloud disk scheduling.
[0087] In one optional embodiment, the cloud server can periodically update the cloud disk groups within each cluster. Specifically, at regular intervals (e.g., several days), the cloud server can re-execute the aforementioned cloud disk grouping process based on the attribute information of the cloud disks within each cluster and the historical throughput information of the cloud disks in the most recent historical period (e.g., the past day or several hours). Subsequent cloud disk scheduling, based on the latest cloud disk groups within each cluster, can improve the accuracy of cloud disk scheduling, thereby better reducing the probability of cluster failures and improving cluster stability.
[0088] The solution in this embodiment takes into account that cloud disks with the same or similar performance patterns (IO Pattern) are prone to resonance and have similar attributes, which are more likely to cause traffic risks within the cluster. By clustering the feature vectors of cloud disks, the cloud disks are divided into cloud disk groups, and cloud disks are balanced and scheduled based on the cloud disk group. This can reduce the probability of cloud disk resonance within the cluster to a certain extent, better consider the imbalance within the cluster (i.e., when cloud disks have the same or similar performance patterns, they are very likely to cause resonance, which leads to cluster traffic risks), and select the most suitable cloud disk for scheduling in the current state to ensure the stability of the cluster.
[0089] For example, Figure 3 This is a general architecture diagram of cloud disk scheduling provided for an exemplary embodiment of this application. (See diagram below.) Figure 3 As shown, the overall architecture of cloud disk scheduling is as follows:
[0090] First, divide the cloud disks in each cluster into cloud disk groups so that the cloud disks in the same cloud disk group have the same or similar performance modes.
[0091] For example, with Figure 3 Taking cluster 1 as an example, the cloud disks within cluster 1 (such as...) Figure 3 The CS1-CS9 shown in the figure are divided into 3 cloud disk groups as shown in the figure (e.g., Figure 3 (As shown in G1, G2, and G3). Cloud disks within each cloud disk group have the same or similar performance modes, while cloud disks in different cloud disk groups have different and dissimilar performance modes. For example, cloud disks CS1 and CS2 in cloud disk group G1 have the same or similar performance modes, while cloud disk CS1 in cloud disk group G1 and cloud disk CS5 in cloud disk group G2 have different and dissimilar performance modes.
[0092] Then, the cloud server monitors the traffic information of each cluster in real time and determines whether there is any traffic risk in each cluster. When a traffic risk is detected in the first cluster (e.g., cluster 1), based on the set cloud disk scheduling constraints (e.g., ... Figure 3 Using the blacklist constraints, cluster capacity constraints, cluster traffic constraints, and cloud disk group constraints shown, a cloud disk scheduling plan is generated. The cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group where the target cloud disk resides, and the second cluster and second cloud disk group to which the target cloud disk will be migrated.
[0093] Furthermore, execute the cloud disk scheduling plan to migrate one or more target cloud disks from the first cluster to the second cluster (such as cluster 2).
[0094] For example, such as Figure 3 As shown, cloud disk CS1 belonging to cloud disk group G1 in cluster 1 and cloud disk CS6 belonging to cloud disk group G2 are migrated to cloud disk groups G4 and G5 in cluster 2, respectively. Cloud disk group G4 in cluster 2 also includes cloud disks G10 and G11, and cloud disk group G5 in cluster 2 also includes cloud disks G12 and G13.
[0095] For example, Figure 4 This is a general flowchart of cloud disk scheduling provided for an exemplary embodiment of this application. Figure 4 As shown, the overall process of cloud disk scheduling is as follows:
[0096] The cloud server periodically groups the various clusters. During this phase, the cloud server periodically divides the cloud disks within the cluster into multiple cloud disk groups.
[0097] Set cloud disk scheduling constraints, including but not limited to: cloud disk group constraints, blacklist constraints, cluster capacity constraints, cluster traffic constraints, and cluster characteristic constraints.
[0098] The cloud server monitors risks in real time, and when any cluster / cloud disk is found to have traffic risks, it generates a cloud disk scheduling plan based on cloud disk scheduling constraints.
[0099] The scheduling task is generated based on the cloud disk scheduling plan, and the cloud disk scheduling can be realized by executing the scheduling task.
[0100] In this embodiment, cloud disks within a cluster are clustered on a per-cluster basis, and classified into different cloud disk groups based on their attributes and historical throughput performance characteristics. When traffic risks are detected, cloud disk scheduling strategies are decided and selected on a per-cloud disk-group basis. Cloud disk scheduling is carried out under the premise of meeting cloud disk scheduling constraints, thereby maintaining traffic balance and stability within the cluster and within the cloud disk groups, and reducing the probability of failure.
[0101] Figure 5 A detailed flowchart of cloud disk scheduling is provided for another exemplary embodiment of this application. (See attached diagram.) Figure 5 As shown, the detailed process of cloud disk scheduling is as follows:
[0102] Step S500: Obtain the attribute information and historical throughput information of the cloud disks in each cluster.
[0103] The implementation principle of this step is the same as that of step S201 above, and the relevant content of the above embodiment is also provided. It will not be repeated here.
[0104] Step S501: Based on the attribute information and historical throughput information of the cloud disks in each cluster, divide the cloud disks in the same cluster into multiple cloud disk groups, so that the cloud disks in the same cloud disk group have the same or similar performance modes.
[0105] The implementation principle of this step is described in steps S202-S203 above, and the relevant content of the above embodiments will not be repeated here.
[0106] Step S502: Monitor the traffic information of each cluster.
[0107] In this step, the cloud server can monitor the traffic information of each cluster, including but not limited to the total traffic of the cluster, the traffic of each cloud disk group within the cluster, and the traffic of each cloud disk within the cluster.
[0108] Step S503: Based on the total traffic of each cluster, determine that the first cluster that does not meet the cluster traffic limit conditions has a traffic risk.
[0109] In this step, the cloud server determines whether each cluster meets the cluster traffic limit conditions based on the total traffic of each cluster, and thus determines whether there is a traffic risk in the cluster.
[0110] For any cluster, if the cluster meets the cluster traffic limit conditions, then the cluster is considered to have traffic risk. If the cluster does not meet the cluster traffic limit conditions, then the cluster is considered to have no traffic risk.
[0111] In this embodiment, the cluster flow limiting conditions include: the total flow reaches the cluster flow level, and / or the difference between the total flow and the average total flow of the cluster in the available area satisfies the first difference condition.
[0112] The cluster traffic limits can differ for different clusters. The cluster traffic level for any given cluster can be determined based on the cluster's rated traffic.
[0113] For example, the cluster flow rate level can be the first percentage of the cluster's rated flow rate, such as 80% or 70%. The first percentage can be set and adjusted according to actual application needs, and no specific limitation is made here.
[0114] For example, the cluster flow level can be the cluster's rated flow minus the first flow difference. The first flow difference can be set and adjusted according to actual application needs, and no specific limitation is made here.
[0115] When the total traffic of the cluster reaches the cluster traffic level, it indicates that the cluster's own traffic has exceeded the limit of the cluster traffic level, and the cluster is considered to have a traffic risk.
[0116] The first difference condition in the cluster traffic limit conditions for any cluster can be: the difference is less than 0 and the absolute value of the difference is greater than the first preset difference. The first preset difference can be set by the rated traffic of the cloud disk group, and is not specifically limited here.
[0117] When the difference between the total traffic of a cluster and the average total traffic of clusters in the available area meets the first difference condition, it indicates that the traffic of each cluster is unbalanced and the traffic of the cluster is significantly higher than the average traffic of clusters in the available area, and the cluster is considered to have traffic risk.
[0118] Step S504: Based on the traffic information of each cluster, generate a cloud disk scheduling plan for the first cluster. The cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group where the target cloud disk is located, and the second cluster and the second cloud disk group to which the target cloud disk is migrated. The cloud disk scheduling plan satisfies the cloud disk scheduling constraints.
[0119] When a traffic risk is detected in the first cluster, the cloud server generates a cloud disk scheduling plan for the first cluster based on the traffic information of each cluster.
[0120] When a first cluster with potential traffic risks is detected, and cloud disk scheduling is performed on the first cluster, the cloud server generates a cloud disk scheduling plan for the first cluster based on the traffic information of each cluster and the set cloud disk scheduling constraints. This cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group to which the target cloud disk belongs, and the second cluster and second cloud disk group to which the target cloud disk will be migrated. The cloud disk scheduling constraints include at least cloud disk group constraints, and may also include at least one of the following: blacklist constraints, cluster traffic constraints, cluster capacity constraints, and cluster characteristic constraints.
[0121] Specifically, the cloud server performs the following cloud disk scheduling processing (a1-a4) for the first cluster iteration:
[0122] a1. Sort the cloud disk groups in the first cluster according to the simulated traffic, and select the cloud disk group with the highest simulated traffic as the first cloud disk group to be scheduled.
[0123] Simulated traffic refers to the estimated traffic after simulating scheduling for a given target cloud disk. In the first iteration, the traffic of each cloud disk group monitored is used as the initial simulated traffic.
[0124] a2. Based on the traffic of each cloud disk in the first cloud disk group, select one cloud disk from the first cloud disk group as the target cloud disk.
[0125] Optionally, the cloud server can select the cloud disk with the highest traffic in the first cloud disk group as the target cloud disk based on the traffic of each cloud disk in the first cloud disk group.
[0126] Optionally, in some instance scenarios, the cloud server can select one or more cloud disks with higher traffic as the target cloud disk based on the traffic of each cloud disk in the first cloud disk group.
[0127] It should be noted that, in order to meet the cloud disk group constraints, the target cloud disks are distributed as widely as possible in multiple cloud disk groups of the first cluster. In each iteration, as few cloud disks as possible are selected as target cloud disks to avoid scheduling multiple cloud disks with the same or similar performance modes within a single cloud disk group. This reduces the probability of cloud disk resonance within the cluster, lowers the probability of cluster failure, and improves the stability of the cluster.
[0128] a3. Based on the traffic information of each cluster and the set cloud disk scheduling constraints, determine the scheduling plan for the target cloud disk.
[0129] The scheduling plan for the target cloud disk includes the migration of the target cloud disk to the second cluster and the second cloud disk group.
[0130] The scheduling plan for the target cloud disk satisfies cloud disk scheduling constraints. Cloud disk scheduling constraints include at least cloud disk group constraints, and may also include at least one of the following: blacklist constraints, cluster traffic constraints, cluster capacity constraints, and cluster characteristic constraints.
[0131] In this step, the cloud server can select the target cluster and target cloud disk group to which the target cloud disk is migrated based on the preset scheduling policy, as well as blacklist constraints and cluster characteristic constraints, to ensure that the selection of the target cluster meets the blacklist constraints and cluster characteristic constraints.
[0132] Furthermore, the cloud server performs a simulation of migrating the target cloud disk to the target cluster and target cloud disk group based on the currently selected target cluster and target cloud disk group, and calculates the simulated traffic of the target cluster and the simulated traffic of the target cloud disk group after migrating the target cloud disk to the target cluster and target cloud disk group.
[0133] Based on the simulated traffic of the target cluster and the target cloud disk group, verify whether the selection of the target cluster and the target cloud disk group meets the cluster traffic constraints, cluster capacity constraints, and cloud disk group constraints.
[0134] If the selection of the target cluster and target cloud disk group does not meet any of the cloud disk scheduling constraints, then the target cluster and target cloud disk group will be reselected for the target cloud disk according to the preset scheduling strategy.
[0135] Until the selection of the target cluster and target cloud disk group satisfies all cloud disk scheduling constraints, the currently selected target cluster and target cloud disk group will be determined as the second cluster and second cloud disk group to which the target cloud disk will be migrated.
[0136] It should be noted that the selection of the destination cluster and destination cloud disk group does not need to satisfy constraints not included in the cloud disk scheduling constraints. If the cloud disk scheduling constraints do not include any of the following constraints: blacklist constraint, cluster traffic constraint, cluster capacity constraint, or cluster characteristic constraint, then the selection of the destination cluster and destination cloud disk group does not need to satisfy that constraint.
[0137] For example, taking the case where cloud disk scheduling constraints do not include cluster characteristic constraints, when selecting the target cluster and target cloud disk group, you can choose according to the preset scheduling strategy and blacklist constraints.
[0138] a4. Add the target cloud disk's scheduling plan to the cloud disk scheduling plan of the first cluster, and update the simulated traffic of the first cluster, the first cloud disk group, the second cluster, and the second cloud disk group.
[0139] After generating the scheduling plan for the target cloud disk, the scheduling plan for the target cloud disk is added to the cloud disk scheduling plan of the first cluster. The scheduling plans for each target cloud disk in the cloud disk scheduling plan of the first cluster can be arranged in the order they were generated.
[0140] Furthermore, the cloud server updates the simulated traffic of the first cluster, the first cloud disk group, the second cluster, and the second cloud disk group.
[0141] Furthermore, based on the updated simulated traffic of the first cluster (referring to the simulated value of the total traffic of the first cluster), it is determined whether the first cluster has traffic risks. If it is determined that the first cluster still has traffic risks, then a1-a4 are iteratively executed based on the updated simulated traffic of each cluster and cloud disk group. This continues until the first cluster no longer has traffic risks, at which point the cloud disk scheduling plan for the first cluster is obtained.
[0142] Step S505: According to the cloud disk scheduling plan, migrate the target cloud disk to the second cloud disk group within the corresponding second cluster.
[0143] After generating the cloud disk scheduling plan for the first cluster, the cloud disk scheduling plan is executed to migrate the target cloud disk to the second cloud disk group within the corresponding second cluster, thereby eliminating the traffic risk present in the first cluster, reducing the probability of cluster failure, and improving the stability of the cluster.
[0144] In this embodiment, cloud disk scheduling is performed at the cloud disk group level. This better considers the imbalance within the cluster (i.e., when cloud disk performance modes are the same or similar, resonance is easily triggered, leading to cluster traffic risks). The most suitable cloud disk for scheduling under the current state is selected, thereby reducing the probability of cloud disk resonance within the cluster and ensuring cluster stability. By introducing cloud disk group constraints, the scope of traffic failures is narrowed from the cluster level to the cloud disk group level, reducing the impact of failures. This embodiment provides a cloud disk scheduling strategy based on multiple constraints, effectively achieving balanced scheduling of distributed cloud disks, realizing elastic scaling of cloud disks within the cluster, improving the efficiency of cloud disk scheduling, reducing the potential risks of cloud disk resonance, and ensuring cluster stability.
[0145] Figure 6 A detailed flowchart of cloud disk scheduling is provided for another exemplary embodiment of this application. (See attached diagram.) Figure 6 As shown, the detailed process of cloud disk scheduling is as follows:
[0146] Step S600: Obtain the attribute information and historical throughput information of the cloud disks in each cluster.
[0147] The implementation principle of this step is the same as that of step S201 above, and the relevant content of the above embodiment is also provided. It will not be repeated here.
[0148] Step S601: Based on the attribute information and historical throughput information of the cloud disks in each cluster, divide the cloud disks in the same cluster into multiple cloud disk groups, so that the cloud disks in the same cloud disk group have the same or similar performance modes.
[0149] The implementation principle of this step is described in steps S202-S203 above, and the relevant content of the above embodiments will not be repeated here.
[0150] Step S602: Monitor the traffic information of each cluster.
[0151] In this step, the cloud server can monitor the traffic information of each cluster, including but not limited to the total traffic of the cluster, the traffic of each cloud disk group within the cluster, and the traffic of each cloud disk within the cluster.
[0152] Step S603: Based on the traffic of cloud disk groups within each cluster, cloud disk groups that do not meet the traffic limit conditions are designated as the third cloud disk group with traffic risk, and the cluster where the third cloud disk group is located is determined to be the first cluster with traffic risk.
[0153] In this step, the cloud server determines whether each cloud disk group meets the traffic limit conditions based on the traffic of the cloud disk groups within each cluster, thereby determining whether each cloud disk group has traffic risks. Cloud disk groups that do not meet the cloud disk group traffic limit conditions are designated as the third cloud disk group with traffic risks, and the cluster containing the third cloud disk group is determined as the first cluster with traffic risks.
[0154] The cloud disk group traffic limit conditions include: the traffic reaches the preset group traffic level, and / or the difference between the traffic and the average traffic of the cloud disk group in the cluster meets the third difference condition.
[0155] Different cloud disk groups can have different traffic limits. The preset traffic threshold for any cloud disk group can be automatically updated based on the proportion of that cloud disk group's traffic to the total traffic of the cluster.
[0156] For example, the preset group traffic threshold can be a fourth percentage of the rated traffic of the cluster to which the cloud disk group belongs. The fourth percentage can be determined based on the proportion of the cloud disk group's traffic to the total traffic of the cluster. For example, if the traffic of a cloud disk group accounts for 80% of the total traffic of the cluster, then the preset group traffic threshold for that cloud disk group can be 80% of the rated traffic of the cluster.
[0157] When the traffic of a cloud disk group reaches the preset traffic level, it means that the traffic of the cloud disk group itself has exceeded the limit of the preset traffic level, and the cloud disk group is considered to have a traffic risk.
[0158] The third difference condition in the traffic limit conditions for any cloud disk group can be: the difference is less than 0 and the absolute value of the difference is greater than the third preset difference. The third preset difference can be set by the rated traffic of the cloud disk group, and is not specifically limited here.
[0159] When the difference between the traffic of a cloud disk group and the average traffic of the cloud disk groups in the same cluster meets the third difference condition, it indicates that the traffic of the cloud disk groups in the cluster is unbalanced and the traffic of the cloud disk group is significantly higher than the average traffic of the cloud disk groups in the same cluster. Therefore, the cloud disk group is considered to have traffic risk.
[0160] Step S604: Based on the traffic information of each cluster, generate a cloud disk scheduling plan for the first cluster. The cloud disk scheduling plan includes the target cloud disk to be migrated, the third cloud disk group where the target cloud disk is located, and the second cluster and the second cloud disk group to which the target cloud disk is migrated. The cloud disk scheduling plan satisfies cloud disk scheduling constraints.
[0161] When a traffic risk is detected in the third cloud disk group within the first cluster, the cloud server generates a cloud disk scheduling plan for the first cluster based on the traffic information of each cluster. This cloud disk scheduling plan includes the target cloud disk to be migrated, the third cloud disk group where the target cloud disk is located, and the second cluster and second cloud disk group to which the target cloud disk will be migrated. The cloud disk scheduling plan satisfies cloud disk scheduling constraints.
[0162] It should be noted that when any cloud disk group (the third cloud disk group) in the first cluster has traffic risks, the generated cloud disk scheduling plan is only used to schedule cloud disks within the third cloud disk group. To distinguish it from the scheduling plan of the first cluster in the aforementioned embodiments, the cloud disk scheduling plan generated in this step is referred to as the cloud disk scheduling plan of the third cloud disk group.
[0163] In this step, the cloud server generates a cloud disk scheduling plan for the third cloud disk group based on the traffic information of each cluster and the set cloud disk scheduling constraints. This cloud disk scheduling plan includes the target cloud disk to be migrated, the third cloud disk group where the target cloud disk resides, and the second cluster and second cloud disk group to which the target cloud disk will be migrated. The cloud disk scheduling constraints include at least cloud disk group constraints, and may also include at least one of the following: blacklist constraints, cluster traffic constraints, cluster capacity constraints, and cluster characteristic constraints.
[0164] Specifically, the third cloud disk group iteration is performed as follows: b1-b4 scheduling process:
[0165] b1. Based on the traffic of each cloud disk in the third cloud disk group, select the cloud disk with the highest traffic in the third cloud disk group as the target cloud disk.
[0166] b2. Based on the traffic information of each cluster and the set cloud disk scheduling constraints, determine the scheduling plan for the target cloud disk.
[0167] The scheduling plan for the target cloud disk must satisfy the cloud disk scheduling constraints. The cloud disk scheduling constraints must include at least cloud disk group constraints, and may also include at least one of the following: blacklist constraints, cluster traffic constraints, cluster capacity constraints, and cluster characteristic constraints.
[0168] The implementation principle of this step is the same as that of step a3 mentioned above. For details, please refer to the relevant content of the aforementioned embodiment. The steps are not repeated here.
[0169] b3. Add the target cloud disk's scheduling plan to the cloud disk scheduling plan of the first cluster, and update the simulated traffic of the first cluster, the third cloud disk group, the second cluster, and the second cloud disk group.
[0170] After generating the scheduling plan for the target cloud disk, the scheduling plan for the target cloud disk is added to the cloud disk scheduling plan of the first cluster (i.e., the cloud disk scheduling plan of the third cloud disk group). The scheduling plans for each target cloud disk in the cloud disk scheduling plan of the third cloud disk group can be arranged in the order they were generated.
[0171] Furthermore, the cloud server updates the simulated traffic of the first cluster, the third cloud disk group, the second cluster, and the second cloud disk group.
[0172] Furthermore, based on the updated simulated traffic of the third cloud disk group, it is determined whether the third cloud disk group still has traffic risks. If it is determined that the third cloud disk group still has traffic risks, then b1-b4 are executed iteratively based on the updated simulated traffic of each cluster and cloud disk group. This continues until the third cloud disk group no longer has traffic risks, at which point the cloud disk scheduling plan for the third cloud disk group is obtained.
[0173] Step S605: According to the cloud disk scheduling plan, migrate the target cloud disk to the second cloud disk group within the corresponding second cluster.
[0174] After generating the cloud disk scheduling plan for the third cloud disk group, the cloud disk scheduling plan is executed to migrate the target cloud disk to the second cloud disk group within the corresponding second cluster, thereby eliminating the traffic risk of the third cloud disk group, reducing the probability of cluster failure, and improving the stability of the cluster.
[0175] Figure 7 This is a schematic diagram of the structure of a cloud server provided in an embodiment of this application. Figure 7 As shown, the cloud server includes a memory 701 and a processor 702. The memory 701 stores computer-executed instructions and can be configured to store various other data to support operations on the cloud server. The processor 702 is communicatively connected to the memory 701 and executes the computer-executed instructions stored in the memory 701 to implement the technical solutions provided in any of the above method embodiments. Their specific functions and the technical effects they achieve are similar and will not be repeated here.
[0176] Optional, such as Figure 7 As shown, the cloud server also includes other components such as a firewall 703, a load balancer 704, a communication component 705, and a power supply component 706. Figure 7 The diagram only shows some components and does not mean that a cloud server includes only these components. Figure 7 The components shown.
[0177] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the method of any of the foregoing embodiments. The specific functions and technical effects to be achieved are not described here.
[0178] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments. The computer program is stored in a readable storage medium, and at least one processor of the cloud server can read the computer program from the readable storage medium. The execution of the computer program by the at least one processor causes the cloud server to perform the technical solution provided in any of the above method embodiments. The specific functions and the technical effects that can be achieved are not described here.
[0179] This application provides a chip, including a processing module and a communication interface. The processing module is capable of executing the technical solution of the cloud server in the aforementioned method embodiments. Optionally, the chip further includes a storage module (e.g., a memory), which stores instructions. The processing module executes the instructions stored in the storage module, and the execution of the instructions stored in the storage module causes the processing module to execute the technical solution provided in any of the aforementioned method embodiments.
[0180] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0181] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), a graphics processing unit (GPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules in at least one processor.
[0182] The memory may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0183] The aforementioned storage device can be object storage service (OSS).
[0184] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0185] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as mobile hotspots (WiFi), second-generation (2G), third-generation (3G), fourth-generation (4G) / Long Term Evolution (LTE), fifth-generation (5G), or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be based on Radio Frequency Identification (RFID), infrared, Ultra Wide Band (UWB), Bluetooth, and other technologies.
[0186] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0187] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0188] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0189] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0190] The order of the embodiments described above is merely for illustrative purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The sequence numbers are merely used to distinguish different operations, and the sequence numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. "Multiple" means two or more, unless otherwise explicitly specified.
[0191] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0192] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0193] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A cloud disk scheduling method, characterized in that, include: Monitor the traffic information of each cluster, wherein any cluster is divided into at least one cloud disk group, and the cloud disks in the same cloud disk group have the same or similar performance modes. If the traffic information of each cluster indicates that the first cluster has a traffic risk, a cloud disk scheduling plan for the first cluster is generated based on the traffic information of each cluster. The cloud disk scheduling plan includes the target cloud disk to be migrated, the first cloud disk group where the target cloud disk is located, and the second cluster and the second cloud disk group to which the target cloud disk is migrated. The cloud disk scheduling plan satisfies the cloud disk group constraint, which includes the target cloud disk being distributed in multiple cloud disk groups of the first cluster. According to the cloud disk scheduling plan, the target cloud disk is migrated to the second cloud disk group within the corresponding second cluster.
2. The method according to claim 1, characterized in that, Traffic information for any cluster, including the total traffic of the cluster and the traffic of cloud disk groups within the cluster. The determination of traffic risk in the first cluster based on the traffic information of each cluster includes: Based on the total traffic of each cluster, the first cluster that does not meet the cluster traffic limit conditions is determined to have traffic risk; The cluster traffic limiting conditions include: the total traffic reaches the cluster traffic level, and / or the difference between the total traffic and the average total traffic of the cluster in the available area satisfies the first difference condition.
3. The method according to claim 2, characterized in that, The step of generating a cloud disk scheduling plan for the first cluster based on the traffic information of each cluster includes: The first cluster is iteratively processed using the following cloud disk scheduling method until there is no longer any traffic risk in the first cluster: The cloud disk groups in the first cluster are sorted according to the simulated traffic, and the cloud disk group with the highest simulated traffic is selected as the first cloud disk group to be scheduled. The simulated traffic refers to the estimated traffic after simulated scheduling of the determined target cloud disk. Based on the traffic of each cloud disk in the first cloud disk group, select one cloud disk from the first cloud disk group as the target cloud disk; Based on the traffic information of each cluster and the set cloud disk scheduling constraints, the scheduling plan of the target cloud disk is determined, wherein the scheduling plan of the target cloud disk satisfies the cloud disk scheduling constraints, and the cloud disk scheduling constraints include the cloud disk group constraints; Add the scheduling plan of the target cloud disk to the cloud disk scheduling plan of the first cluster, and update the simulated traffic of the first cluster, the first cloud disk group, the second cluster, and the second cloud disk group.
4. The method according to claim 1, characterized in that, Traffic information for any cluster, including the total traffic of the cluster and the traffic of cloud disk groups within the cluster. The determination of traffic risk in the first cluster based on the traffic information of each cluster includes: Based on the traffic of each cloud disk group within the cluster, cloud disk groups that do not meet the cloud disk group traffic limit conditions are designated as the third cloud disk group with traffic risk. The cluster to which the third cloud disk group is located is determined to be the first cluster with traffic risk.
5. The method according to claim 4, characterized in that, The cloud disk group traffic limit conditions include: The flow rate has reached the preset group flow rate level. And / or, The difference between the traffic and the average traffic of the cloud disk group in the cluster satisfies the second difference condition.
6. The method according to claim 4, characterized in that, The step of generating a cloud disk scheduling plan for the first cluster based on the traffic information of each cluster includes: The third cloud disk group is iteratively scheduled as follows until the third cloud disk group meets the cloud disk group traffic limit conditions: Based on the traffic of each cloud disk in the third cloud disk group, the cloud disk with the highest traffic in the third cloud disk group is selected as the target cloud disk. Based on the traffic information of each cluster and the set cloud disk scheduling constraints, the scheduling plan of the target cloud disk is determined, wherein the scheduling plan of the target cloud disk satisfies the cloud disk scheduling constraints, and the cloud disk scheduling constraints include the cloud disk group constraints; Add the scheduling plan of the target cloud disk to the cloud disk scheduling plan of the first cluster, and update the simulated traffic of the first cluster, the third cloud disk group, the second cluster, and the second cloud disk group.
7. The method according to claim 1, characterized in that, The cloud disk group constraints also include: After migrating the target cloud disk to the second cloud disk group, the second cloud disk group meets the cloud disk group traffic limit conditions.
8. The method according to claim 3 or 6, characterized in that, The cloud disk scheduling constraints also include at least one of the following: Blacklist constraints, cluster traffic constraints, and cluster capacity constraints; The blacklist constraint includes: the second cluster is not in the blacklist of the first clusters that cannot schedule each other, and the second cloud disk group is not in the blacklist of the first cloud disk group that cannot schedule each other. The cluster traffic constraint includes: after migrating the target cloud disk to the second cloud disk group of the second cluster, the second cluster meets the cluster traffic constraint conditions; The cluster capacity constraint includes: after migrating the target cloud disk to the second cloud disk group of the second cluster, the second cluster meets the cluster capacity limit condition.
9. The method according to claim 8, characterized in that, The cloud disk scheduling constraints also include: cluster characteristic constraints. The cluster characteristic constraints include at least one of the following custom constraints: Specify that a cluster allows the migration of cloud disks with specific IO types; Cloud disks in a specific cluster can only be migrated to clusters that are on the whitelist corresponding to that specific cluster.
10. The method according to any one of claims 1-7, characterized in that, Also includes: Based on the attribute information and historical throughput information of the cloud disks in each cluster, the cloud disks in the same cluster are divided into multiple cloud disk groups, so that the cloud disks in the same cloud disk group have the same or similar performance modes.
11. The method according to claim 10, characterized in that, The step of dividing cloud disks within the same cluster into multiple cloud disk groups based on the attribute information and historical throughput information of the cloud disks in each cluster, so that cloud disks within the same cloud disk group have the same or similar performance modes, includes: The attribute information and historical throughput information of the cloud disks in each cluster are encoded into vectors to obtain the feature vectors of each cloud disk in each cluster. Based on the feature vectors of each cloud disk in any cluster, the cloud disks in the cluster are divided into multiple cloud disk groups. Each cloud disk group contains at least one cloud disk, and the cloud disks in the same cloud disk group have the same or similar performance modes.
12. A cloud server, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the server to perform the method according to any one of claims 1-11.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-11.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-11.