Resource scheduling method, distributed storage system, electronic device and storage medium
Patent Information
- Application Number
- PCT/CN2025/141298
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-12-09
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025141298_01102026_PF_FP_ABST
Abstract
Description
Resource scheduling methods, distributed storage systems, electronic devices and storage media Technical Field
[0001] This disclosure relates to the fields of cloud storage and system scheduling control, and more specifically, to a resource scheduling method, a distributed storage system, an electronic device, and a storage medium. Background Technology
[0002] In a distributed cloud storage environment, when users create cloud disks based on cluster snapshots, the conventional cluster selection strategy tends to distribute disks across multiple clusters. This often leads to unnecessary cross-cluster traffic consumption and the risk of capacity redundancy in the target cluster. Specifically, when cluster snapshot data is copied to another cluster to create a new cloud disk, especially in scenarios with large amounts of snapshot data, not only will the space snapshot of the target cluster increase significantly, but cross-cluster data transfer will also consume a large amount of bandwidth, significantly impacting network performance. Furthermore, if multiple disks are created simultaneously based on the same snapshot, it may cause cluster resource overload, affecting the real-time throughput and write volume of the user's cloud disk, ultimately threatening the overall stability and performance of the cloud storage system. In other words, there is a risk of creating cloud disks across storage clusters based on snapshot data, resulting in ineffective cross-storage cluster traffic and capacity redundancy.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This disclosure provides a resource scheduling method, a distributed storage system, an electronic device, and a storage medium to at least solve the technical problem of invalid cross-storage cluster traffic and capacity guarantees caused by creating cloud disks across storage clusters based on snapshot data.
[0005] According to one aspect of the present disclosure, a resource scheduling method is provided, applied to a scheduling system connected to multiple storage clusters. The method includes: responding to a cloud disk creation request for snapshot data, determining the amount of snapshot data and a source storage cluster storing the snapshot data; if the amount of data is greater than or equal to a preset amount of data and the source storage cluster meets preset resource scheduling conditions, sending a cloud disk creation request to the source storage cluster, so that the source storage cluster creates a cloud disk based on the cloud disk creation request; if the amount of data is less than the preset amount of data or the source storage cluster does not meet the preset resource scheduling conditions, determining a target storage cluster from among the multiple storage clusters, excluding the source storage cluster, and sending a cloud disk creation request to the target storage cluster, so that the target storage cluster creates a cloud disk based on the cloud disk creation request.
[0006] According to another aspect of the embodiments of this disclosure, a distributed storage system is also provided, including: multiple storage clusters; and a scheduling system connected to the multiple storage clusters and configured to execute the methods in the various embodiments of this disclosure.
[0007] According to another aspect of the present disclosure, an electronic device is also provided, including: a memory storing an executable program; and a processor connected to the memory via a bus for running the program, wherein the program executes the methods in various embodiments of the present disclosure during runtime.
[0008] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the methods of the various embodiments of the present disclosure.
[0009] According to another aspect of the embodiments of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this disclosure.
[0010] According to another aspect of the embodiments of this disclosure, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods of various embodiments of this disclosure.
[0011] According to another aspect of the embodiments of this disclosure, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this disclosure.
[0012] In this embodiment of the disclosure, in response to a cloud disk creation request for snapshot data, the amount of snapshot data and the source storage cluster storing the snapshot data are determined. If the amount of data is greater than or equal to a preset amount and the source storage cluster meets preset resource scheduling conditions, a cloud disk creation request is sent to the source storage cluster, enabling the source storage cluster to create a cloud disk based on the request. If the amount of data is less than the preset amount or the source storage cluster does not meet the resource scheduling conditions, a target storage cluster is determined from multiple storage clusters, excluding the source storage cluster, and a cloud disk creation request is sent to the target storage cluster, enabling the target storage cluster to create a cloud disk based on the request. It's noteworthy that when a cloud disk creation request based on snapshot data is received, the scheduling system can dynamically create a cloud disk based on the snapshot data. If the snapshot data volume is greater than or equal to a preset data volume and the source storage cluster meets the preset resource scheduling conditions, a new disk can be created in the source storage cluster to which the snapshot data belongs. This achieves affinity scheduling, avoiding data transfer across storage clusters. Since the proportion of snapshot data greater than or equal to the preset data volume is relatively small, creating a cloud disk in the source storage cluster has minimal impact on the actual default scheduling. It also mitigates the traffic and space issues caused by creating disks across clusters for large snapshots. Furthermore, before creating a new disk for snapshot data greater than or equal to the preset data volume in the source storage cluster, the source storage cluster must meet the preset resource scheduling conditions to ensure that the source storage cluster does not excessively consume resources due to requests to create multiple disks based on snapshots, effectively avoiding traffic congestion. This system addresses the issue of data clustering, preventing sudden impacts on cluster performance and capacity caused by snapshot data creation, thus ensuring service continuity and stability. Furthermore, when the snapshot data volume is less than the preset volume, or when the source storage cluster's quota is exhausted and no new disk can be created within the source storage cluster, the system can initiate a process to determine alternative clusters. This enables rapid and intelligent cluster selection and data transfer strategies when the source storage cluster cannot meet the request or the snapshot data volume is small. This ensures fast response and processing of cloud disk creation requests, avoiding prolonged waiting times and resource allocation failures. Through this process, the system effectively reduces cross-cluster data transfer of snapshot data, avoids rigid redemption of traffic and storage capacity between clusters, optimizes resource utilization, and ultimately solves the technical problem of ineffective cross-storage cluster traffic and capacity redemption caused by creating cloud disks based on snapshot data across storage clusters.
[0013] It is worth noting that the above general description and the following detailed description are merely for illustrative and explanatory purposes and do not constitute a limitation thereof. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:
[0015] Figure 1 is a schematic diagram of the hardware structure of a computer terminal according to an embodiment of the present disclosure;
[0016] Figure 2 is a schematic diagram of a computer terminal as a computing node in a computing environment according to an embodiment of the present disclosure;
[0017] Figure 3 is a schematic diagram of a structural block diagram of using a computer terminal as a service mesh according to an embodiment of the present disclosure;
[0018] Figure 4 is a flowchart of a resource scheduling method according to an embodiment of the present disclosure;
[0019] Figure 5 is a schematic diagram of an optional cloud disk creation process according to an embodiment of the present disclosure;
[0020] Figure 6 is a schematic diagram of an optimized cross-cluster snapshot data creation cloud disk according to an embodiment of the present disclosure;
[0021] Figure 7 is a schematic diagram of creating a cloud disk from snapshot data according to a soft rejection retry fallback mechanism according to an embodiment of the present disclosure;
[0022] Figure 8 is an interactive diagram of creating a cloud disk from snapshot data according to a dynamic parameter mechanism of an embodiment of the present disclosure;
[0023] Figure 9 is a schematic diagram of the relationship between the capacity and performance of an optional cloud disk according to an embodiment of the present disclosure;
[0024] Figure 10 is a schematic diagram of the relationship between single gigabyte performance and capacity of an optional cloud disk according to an embodiment of the present disclosure;
[0025] Figure 11 is a structural block diagram of a distributed storage system according to an embodiment of the present disclosure;
[0026] Figure 12 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] First, some nouns or terms that appear in the description of the embodiments of this disclosure shall be interpreted as follows:
[0030] Ocean: A block storage region-level management component.
[0031] RiverMaster; a block storage availability zone-level management component.
[0032] BlockMaster: A block storage cluster-level management and control component.
[0033] Storage components are deployed at the block storage cluster level.
[0034] According to embodiments of this disclosure, a resource scheduling method is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0035] The method embodiments provided in this disclosure can be executed in a mobile terminal, computer terminal, or similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a resource scheduling method. As shown in Figure 1, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a cursor control device, a keyboard, a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that the structure shown in Figure 1 is merely illustrative and does not limit the structure of the above-described electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a different configuration than shown in Figure 1.
[0036] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in embodiments of this disclosure, the data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0037] The memory 104 can be used to store software programs and modules of application software, such as program instructions and / or data storage devices corresponding to the methods in the embodiments of this disclosure. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the methods in the above embodiments. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0038] The transmission device 106 is used to receive or send data via a network and can establish wired and / or wireless network connections. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. The transmission device 106 may also include a network interface. In one example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0039] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0040] The hardware structure block diagram shown in Figure 1 can serve as an exemplary block diagram not only for the computer terminal 10 (or mobile device) described above, but also as an exemplary block diagram for the server described above. In an optional embodiment, Figure 2 shows a block diagram of an example of using the computer terminal 10 (or mobile device) shown in Figure 1 as a computing node in the computing environment 201.
[0041] Figure 2 is a schematic diagram of a computer terminal as a computing node in a computing environment according to an embodiment of the present disclosure. As shown in Figure 2, the computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 210-1, 210-2, ... in the figure). Each computing node contains local processing and memory resources, and the terminal user 202 can remotely run applications or store data in the computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in the computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0042] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).
[0043] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0044] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in Figure 2, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. The proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.
[0045] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201. Executing one or more functions of one service requires calling one or more functions of another service. As shown in Figure 2, service "A" 220-1 receives a user request from end user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to execute one or more functions.
[0046] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0047] In another alternative embodiment, FIG3 illustrates a block diagram of an example using the computer terminal 10 (or mobile device) shown in FIG1 above as a service mesh. FIG3 shows a structural block diagram of a service mesh 300, which is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them across different clusters / machines.
[0048] As shown in Figure 3, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 300. In one implementation, application service instance A runs as a container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs as a container / process 310 on machine / workload container group 316 (Pod).
[0049] As shown in Figure 3, application service instance A and grid proxy (sidecar) 303 coexist in machine / workload container group 314, and application service instance B and grid proxy 305 coexist in machine / workload container group 316. Grid proxy 303 and grid proxy 305 form the data plane layer of service mesh 300. Grid proxy 303 and grid proxy 305 run as containers / processes 304 and 306 respectively, and can receive requests 312 for product query services. Grid proxy 303 and application service instance A can communicate bidirectionally, and grid proxy 305 and application service instance B can also communicate bidirectionally. Furthermore, grid proxy 303 and grid proxy 305 can also communicate bidirectionally with each other.
[0050] In one implementation, traffic from application service instance A is routed to the appropriate destination via mesh proxy 303, and network traffic from application service instance B is routed to the appropriate destination via mesh proxy 305. It should be noted that the network traffic mentioned here includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), high-performance, general-purpose open-source frameworks (Google Remote Procedure Call, gRPC), and open-source in-memory data structure storage systems (Redis).
[0051] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the proxy (Envoy) in service mesh 300. The service mesh proxy configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh proxies 303 and 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0052] As shown in Figure 3, the service mesh 300 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, managed by a managed control plane component 301 within machine / workload container groups (machine / Pods) 302. As shown in Figure 3, the managed control plane component 301 communicates bidirectionally with mesh agents 303 and 305. The managed control plane component 301 is configured to perform control and management functions. For example, it receives telemetry data from mesh agents 303 and 305 and can further aggregate this telemetry data. The managed control plane component 301 can also provide a user-facing Application Programming Interface (API) for these services, facilitating easier manipulation of network behavior and providing configuration data to mesh agents 303 and 305.
[0053] In the above operating environment, this disclosure provides a resource scheduling method as shown in Figure 4, which is applied to a scheduling system connected to multiple storage clusters. Figure 4 is a flowchart of a resource scheduling method according to an embodiment of this disclosure. As shown in Figure 4, it may specifically include the following steps:
[0054] Step S402: Respond to the cloud disk creation request for the snapshot data, and determine the amount of snapshot data and the source storage cluster storing the snapshot data.
[0055] The aforementioned storage cluster can refer to multiple storage servers working collaboratively to provide storage services and achieve high availability and performance for data. In a cloud storage system, data can be distributed across different storage clusters, each with its own storage and computing resources. These clusters can also be connected via a network for data replication and transfer.
[0056] The aforementioned snapshot data can be a capture or copy of data state at a specific point in time. Snapshot data can be used for scenarios such as data recovery, copying, and creating new disks. In cloud storage, snapshot data can be created for cloud disks, forming an independent data copy. The creation and use of snapshot data pose challenges to the management of storage resources, especially large snapshots, i.e., snapshots with large data volumes. The creation and use of large snapshots consume significant storage resources, not only occupying a large amount of storage space but also consuming a large amount of network bandwidth during data reading, copying, and transmission, as well as during cross-cluster transmission.
[0057] The aforementioned cloud disk creation request can be initiated by the system or a user, requesting the cloud storage service to create a new cloud disk instance. A cloud disk is a persistent storage resource in a cloud environment, where systems or users can deploy operating systems, applications, or store data. A cloud disk creation request based on snapshot data, where the system or user needs to copy snapshot data to create a cloud disk instance identical to the snapshot content, can be used for scenarios such as data migration, environment replication, or fault recovery.
[0058] The aforementioned data volume refers to the total size of the data contained in the snapshot, which can be measured in bytes, such as GB, TB, PB, etc. In cloud storage scenarios, the size of the data can affect the consumption of storage resources and the efficiency of data transmission.
[0059] In an optional embodiment, the resource scheduling method proposed in this disclosure can be implemented as a scheduling system in cloud storage, specifically as a storage availability zone level management component (RiverMaster). Here, an availability zone can refer to one or more independent data centers of a cloud computing service provider within a geographical region. In the cloud storage field, a storage availability zone level management component can refer to a component capable of managing and controlling storage resources within different availability zones. By storing data in different availability zones, data availability and disaster recovery capabilities can be improved, ensuring that data remains available even if an availability zone fails. When the system receives a cloud disk creation request based on snapshot data, it can first obtain the amount of snapshot data specified in the request. The amount of snapshot data can be determined based on the actual space occupied by the snapshot data in the source storage cluster, which can include the metadata and actual amount of snapshot data. The system can also identify the source storage cluster storing the snapshot data, which can be done by querying the metadata information of the snapshot data. The metadata includes the creation location and storage cluster number of the snapshot data. In this process, by determining the amount of snapshot data and the source storage cluster through the disk creation request, the cloud platform can more accurately understand the data status of the snapshot data, facilitating subsequent data transfer and storage resource allocation based on the amount of snapshot data and the source storage cluster where the snapshot data resides.
[0060] Step S404: If the data volume is greater than or equal to the preset data volume and the source storage cluster meets the preset resource scheduling conditions, a cloud disk creation request is sent to the source storage cluster so that the source storage cluster can create a cloud disk based on the cloud disk creation request.
[0061] The aforementioned preset data volume can be a pre-set judgment value used to measure the data volume of snapshot data. An appropriate preset data volume can be set so that the proportion of snapshot data greater than or equal to the preset data volume, i.e., large snapshots, is small. In this way, the creation of cloud disks for subsequent source storage clusters for snapshot data with a data volume greater than or equal to the preset data volume can have a smaller impact on the actual default scheduling. At the same time, it can resolve the traffic and space issues caused by creating disks across clusters for large snapshots. Preferably, the preset data volume can be set to 5T. According to historical data statistics, the proportion of snapshot data greater than or equal to 5T is about 0.3%. The preset data volume can be determined according to actual needs and is not limited here.
[0062] The aforementioned preset resource scheduling conditions can refer to the resource scheduling conditions set in advance for the source storage cluster. For example, it can be a quota mechanism. The quota mechanism can be used to limit the amount of disk space used, including hard disk space quotas and file number quotas. It can effectively manage system resources, prevent resource abuse and resource bottlenecks, and better control the allocation and utilization of system resources.
[0063] In one optional embodiment, when a cloud disk creation request based on snapshot data is received, the scheduling system first checks whether the amount of snapshot data involved in the cloud disk creation request is greater than or equal to a preset data amount, such as whether it is greater than or equal to 5TB. For snapshot data with a data amount greater than or equal to the preset data amount, the system can further determine whether the source storage cluster meets the preset resource scheduling conditions. If the data amount of the snapshot data is greater than or equal to the preset data amount and the source storage cluster meets the preset resource scheduling conditions, a new disk can be created in the source storage cluster to which the snapshot data belongs. This implements an affinity strategy to avoid data transfer across storage clusters. Affinity strategy scheduling can maintain a certain association between the creation of new disks and the source storage cluster where the snapshot data is stored, thereby reducing cross-cluster data transfer. Based on affinity scheduling, the preset resource scheduling conditions further ensure that the source storage cluster will not excessively consume resources due to requests to create multiple disks based on snapshots. Each large snapshot can be allocated a certain number of quotas. In the source storage cluster, the creation of a new cloud disk is only allowed when there are available quotas. This mechanism effectively avoids traffic congestion. Even with a large number of disk creation requests in a short period, it can control and balance resource allocation through preset resource scheduling conditions, preventing the bandwidth and storage space of the source cluster from being rapidly consumed. In the above process, the combination of affinity strategies and preset resource scheduling conditions significantly reduces the need for cross-cluster data transfer, reduces unnecessary network traffic consumption, and prevents the storage capacity of the source and destination clusters from being consumed by large snapshot disk creation activities in a short period. By reducing cross-cluster data replication, the response time for cloud disk creation is shortened. Preset resource scheduling conditions ensure reasonable resource allocation while avoiding sudden impacts on cluster performance and capacity from snapshot data disk creation, ensuring service continuity and stability. This reduces cross-cluster data transfer, avoids rigid redemption of traffic and storage capacity between clusters, thereby optimizing resource utilization and improving service efficiency and user experience.
[0064] Step S406: If the data volume is less than the preset data volume or the source storage cluster does not meet the preset resource scheduling conditions, determine the target storage cluster from multiple storage clusters other than the source storage cluster, and send the cloud disk creation request to the target storage cluster so that the target storage cluster can create a cloud disk based on the cloud disk creation request.
[0065] In an optional embodiment, when the data volume is less than a preset data volume, or the source storage cluster does not meet the preset resource scheduling conditions—that is, in cloud storage applications, the snapshot data volume is less than the preset data volume, for example, the snapshot data volume is less than 5TB—or the quota of the source storage cluster is exhausted and a new disk cannot be created in the source storage cluster, the system can initiate a process to determine alternative clusters. This process may include resource evaluation, intelligent scheduling, and data transmission optimization of multiple storage clusters to ensure that the cloud disk creation activity can be completed smoothly with minimal resource consumption and optimized performance. Specifically, the current resource status of multiple storage clusters can be evaluated, including storage capacity and network bandwidth. The resource evaluation can be based on the real-time resource level and historical data of the clusters to identify clusters with sufficient resources and suitable as the target cluster for new cloud disk creation. During the evaluation process, the system can prioritize clusters with more balanced resource utilization. Specifically, it selects the target storage cluster from multiple storage clusters, excluding the source cluster. Once the target storage cluster is determined, the system sends the cloud disk creation request to it. This allows the target storage cluster to create the cloud disk based on the request, essentially copying the snapshot data from the source cluster to the target cluster. During data transmission, the system employs effective traffic control and optimization strategies, such as compression, deduplication, and incremental transmission, to reduce network resource consumption and improve transmission efficiency. In situations where the source storage cluster cannot meet the request or the snapshot data volume is small, the rapid and intelligent cluster selection and data transmission strategy ensures quick response and processing of cloud disk creation requests, avoiding prolonged waiting times and resource allocation failures, thus significantly improving the user experience. Even in resource-constrained environments, users can obtain stable service quality. The alternative cluster selection mechanism helps achieve load balancing among clusters, preventing any single cluster from becoming overloaded due to data creation activities. By balancing the load across multiple clusters, the system can handle high-concurrency disk creation requests more efficiently, enhancing the overall processing capacity and response speed of the system.
[0066] In this embodiment of the disclosure, in response to a cloud disk creation request for snapshot data, the amount of snapshot data and the source storage cluster storing the snapshot data are determined. If the amount of data is greater than or equal to a preset amount and the source storage cluster meets preset resource scheduling conditions, a cloud disk creation request is sent to the source storage cluster, enabling the source storage cluster to create a cloud disk based on the request. If the amount of data is less than the preset amount or the source storage cluster does not meet the resource scheduling conditions, a target storage cluster is determined from multiple storage clusters, excluding the source storage cluster, and a cloud disk creation request is sent to the target storage cluster, enabling the target storage cluster to create a cloud disk based on the request. It's noteworthy that when a cloud disk creation request based on snapshot data is received, the scheduling system can dynamically create a cloud disk based on the snapshot data. If the snapshot data volume is greater than or equal to a preset data volume and the source storage cluster meets the preset resource scheduling conditions, a new disk can be created in the source storage cluster to which the snapshot data belongs. This achieves affinity scheduling, avoiding data transfer across storage clusters. Since the proportion of snapshot data greater than or equal to the preset data volume is relatively small, creating a cloud disk in the source storage cluster has minimal impact on the actual default scheduling. It also mitigates the traffic and space issues caused by creating disks across clusters for large snapshots. Furthermore, before creating a new disk for snapshot data greater than or equal to the preset data volume in the source storage cluster, the source storage cluster must meet the preset resource scheduling conditions to ensure that the source storage cluster does not excessively consume resources due to requests to create multiple disks based on snapshots, effectively avoiding traffic congestion. This system addresses the issue of data clustering, preventing sudden impacts on cluster performance and capacity caused by snapshot data creation, thus ensuring service continuity and stability. Furthermore, when the snapshot data volume is less than the preset volume, or when the source storage cluster's quota is exhausted and no new disk can be created within the source storage cluster, the system can initiate a process to determine alternative clusters. This enables rapid and intelligent cluster selection and data transfer strategies when the source storage cluster cannot meet the request or the snapshot data volume is small. This ensures fast response and processing of cloud disk creation requests, avoiding prolonged waiting times and resource allocation failures. Through this process, the system effectively reduces cross-cluster data transfer of snapshot data, avoids rigid redemption of traffic and storage capacity between clusters, optimizes resource utilization, and ultimately solves the technical problem of ineffective cross-storage cluster traffic and capacity redemption caused by creating cloud disks based on snapshot data across storage clusters.
[0067] In the above embodiments of this disclosure, the source storage cluster is determined to meet the preset resource scheduling conditions by: obtaining the current resource scheduling quota of the snapshot data; if the current resource scheduling quota is greater than or equal to the preset quota, the source storage cluster is determined to meet the preset resource scheduling conditions; if the current resource scheduling quota is less than the preset quota, the source storage cluster is determined not to meet the preset resource scheduling conditions.
[0068] The aforementioned current resource scheduling quota can be the resource scheduling quota allocated to snapshot data. Specifically, the current resource scheduling quota can be the resource scheduling quota directly allocated to snapshot data, or it can be the resource scheduling quota obtained by reducing the initial resource scheduling quota allocated to snapshot data. The current resource scheduling quota can be allocated to snapshot data according to the total resources of the source storage cluster. When the total resources of the source storage cluster are large, a larger current resource scheduling quota can be allocated to snapshot data; when the total resources of the source storage cluster are small, a smaller current resource scheduling quota can be allocated to snapshot data. This achieves dynamic allocation of an appropriate current resource scheduling quota to snapshot data. The current resource scheduling quota can be determined according to actual needs and is not limited here.
[0069] The aforementioned preset quota can refer to the quota pre-allocated to the source storage cluster. It can be used to measure the remaining capacity of the source storage cluster to create new disks for snapshot data. The preset quota can be determined according to actual needs and is not limited here.
[0070] In one optional embodiment, the current resource scheduling quota of the snapshot data can be obtained. By determining whether the source storage cluster meets preset resource scheduling conditions, it can be ensured that the source storage cluster will only accept cloud disk creation requests based on snapshot data when resources allow, thereby avoiding excessive resource consumption and the risk of guaranteed returns. Specifically, when the system receives a cloud disk creation request based on snapshot data, it first checks whether the current resource scheduling quota of the snapshot data in the source storage cluster is greater than or equal to a preset quota. The preset quota can be set according to system design and resource management strategies, aiming to balance resource consumption and user service needs. If the current quota meets or exceeds the preset quota, the system determines that the source storage cluster meets the preset resource scheduling conditions and can accept the cloud disk creation request; if the current quota is lower than the preset quota, the system determines that the source storage cluster does not meet the resource scheduling conditions, and the request can be passed to other storage clusters. In the above process, by introducing preset resource scheduling conditions, the system can control the use of snapshot data in a fine-grained manner, avoiding a sudden increase in resource consumption due to snapshot-based disk creation activities at a certain point in time, which could cause the capacity or bandwidth of the source storage cluster to reach its limit. This avoids the problem of resource rigidity and ensures a stable supply of system resources. The dynamic allocation and management mechanism of quotas enables the creation of new cloud disks to be more evenly distributed among multiple storage clusters in the entire availability zone, avoiding excessive concentration or idleness of resources, improving the overall utilization efficiency of resources, and reducing the frequency of cross-cluster data transmission, thus saving network resources.
[0071] In the above embodiments of this disclosure, the current resource scheduling quota of the snapshot data is less than or equal to the initial resource scheduling quota of the snapshot data, and the initial resource scheduling quota of the snapshot data is determined based on the total resource amount of the source storage cluster.
[0072] In one optional embodiment, an initial resource scheduling quota can be allocated to the snapshot data. The initial resource scheduling quota can be set based on historical data and statistical analysis. For example, if the source storage cluster size is less than 48, one quota can be allocated to each large snapshot, meaning the initial resource scheduling quota for that snapshot data is 1. If the source storage cluster size is greater than or equal to 48, two quotas can be allocated to each large snapshot, meaning the initial resource scheduling quota for that snapshot data is 2. The method for determining the initial resource scheduling quota for the snapshot data can be determined according to actual needs and is not limited here. Furthermore, by determining whether the source storage cluster meets the preset resource scheduling conditions, it can be ensured that the source storage cluster will only undertake cloud disk creation requests based on snapshot data when resources allow, thereby avoiding excessive resource consumption and the risk of guaranteed returns. As cloud disk creation requests based on snapshot data are processed, the system automatically reduces the current resource scheduling quota for that snapshot data after each successful cloud disk creation, thus ensuring that the current resource scheduling quota for the snapshot data is less than or equal to the initial resource scheduling quota for the snapshot data. This dynamic adjustment mechanism ensures that quota consumption is reflected in the system in real time, enabling the scheduling system to accurately assess the remaining disk creation capacity to guide subsequent request allocation decisions. For example, if the initial quota for a large snapshot is 2, after successfully creating a cloud disk based on that snapshot, its quota will be reduced to 1. This means the number of cloud disks that can be further created from that snapshot in the source cluster is reduced, and the system will decide whether to allow subsequent disk creation requests to execute in that cluster based on the current quota. The quota mechanism also includes dynamic adjustment and management processes. When snapshot data is used to create a cloud disk, the system automatically reduces the quota for that snapshot in the source storage cluster to reflect actual resource consumption. If the snapshot data quota is exhausted, the system will automatically refuse to create any new disks based on that snapshot in the source cluster, thus avoiding overuse of source cluster resources. Furthermore, the quota is configurable; system administrators can adjust the quota threshold according to actual needs and cluster resource status to adapt to different application requirements and resource management strategies.
[0073] In the above embodiments of this disclosure, sending a cloud disk creation request to a source storage cluster or to a destination storage cluster includes: adding a first preset field to the cloud disk creation request to obtain a first creation request; sending the first creation request to the target cluster and receiving a first or second prompt message returned by the target cluster, wherein the first prompt message is returned by the target cluster after determining that the first resource level has not exceeded the preset level and based on the first creation request to create a cloud disk, and the second prompt message is returned when the first resource level exceeds the preset level; wherein the target cluster is a source storage cluster or a destination storage cluster, the first resource level is used to characterize the resource level of the target cluster after the cloud disk is created, the first prompt message is used to indicate that the cloud disk creation was successful, and the second prompt message is used to indicate that the cloud disk creation failed.
[0074] The first preset field mentioned above can be the request identifier field set in the cloud disk creation request for the soft rejection mechanism. For example, it can be the check_space:enable field. The specific form of the first preset field can be determined according to actual needs and is not limited here.
[0075] The first prompt message mentioned above can be the request for consent field replied by the target cluster in the soft rejection mechanism. The specific form of the first prompt message can be determined according to actual needs, and is not limited here.
[0076] The second prompt message mentioned above can be the request rejection field replied by the target cluster in the soft rejection mechanism. For example, it can be a soft reject. The specific form of the second prompt message can be determined according to actual needs and is not limited here.
[0077] In an optional embodiment, the system can generate a first creation request carrying a first preset field, such as `check_space: enable`. This first creation request instructs the target cluster to check its current resource level before processing the disk creation request, assessing whether it exceeds a preset upper limit. This ensures that disk creation activities are not performed blindly but are based on the actual resource status of the target cluster. The first creation request is sent to the target cluster, which can be either the source or destination storage cluster. Upon receiving the request, the target cluster first checks the first resource level—whether the cluster's resource level exceeds the preset level after creating the cloud disk. If it does not exceed the preset level, the target cluster can create the cloud disk based on the first creation request and return a first notification message to inform the system that the disk creation was successful. If the first resource level exceeds the preset level, the target cluster will return a second notification message indicating that the disk creation failed. In this process, through a soft rejection mechanism, the system can avoid meaningless disk creation attempts in resource-constrained clusters, thereby reducing ineffective resource consumption and improving the overall system resource utilization efficiency. The soft rejection mechanism ensures that disk creation is completed first when resource status allows. The soft rejection mechanism can predict the impact of disk creation activities on cluster resources in advance, and avoid the cluster's performance or capacity from reaching the critical point due to a large number of disk creation requests in a short period of time. The soft rejection mechanism prompts the cluster's physical level to be checked before disk creation by using a first preset field in the request. If it is found that disk creation will cause resource tension and reach the risk value, an error code is returned, thereby preventing this operation and avoiding performance degradation or service interruption.
[0078] In the above embodiments of this disclosure, the method further includes: in response to a second prompt message sent by the target cluster, determining a first storage cluster from among multiple storage clusters, excluding the target cluster; sending a first creation request to the first storage cluster; and in response to a second prompt message sent by the first storage cluster, repeatedly performing the steps of determining a new storage cluster from among multiple storage clusters, excluding the target cluster and the first storage cluster, and sending a first creation request to the new storage cluster, until a first prompt message is received from the new storage cluster, or a second prompt message is received from any one of the multiple storage clusters.
[0079] In one optional embodiment, the system can select a target cluster according to a resource scheduling strategy and send a first creation request carrying a first preset field to that cluster. Upon receiving the request, the target cluster checks its resource level. If it finds that the resource level will exceed the preset level after disk creation, it will return a response containing a second warning message, indicating disk creation failure. When the system receives the second warning message, it indicates that the target cluster cannot currently fulfill the request. At this point, the system will exclude the target cluster that received the second warning message from multiple storage clusters and select another cluster, namely the first storage cluster, as a backup. It will then send a first creation request carrying the same field to this backup cluster. This process uses a "retry" mechanism to attempt to complete disk creation in a cluster with sufficient resources, achieving multiple attempts. Whenever the first storage cluster also returns the second warning message, indicating disk creation failure, the system will add the first storage cluster to the target cluster list, meaning it also becomes a resource-constrained cluster. This mechanism continues to iterate. The system will attempt to send the first creation request to other storage clusters in the list that have not yet been tried, until it receives the first notification message from one cluster indicating successful disk creation, or all storage clusters return a second notification message, indicating that under the current resource constraints, none of the clusters can fulfill the disk creation request. Through multiple attempts and intelligent scheduling, the soft rejection "retry" mechanism significantly improves the success rate of disk creation requests, ensuring that user cloud disk creation requests can be processed even in environments with generally tight resources, thereby significantly improving the user experience. The introduction of the "retry" mechanism enhances the system's flexibility and fault tolerance. Even when a single cluster is under resource pressure, the system can maintain service continuity by scheduling to other clusters, improving the overall stability and reliability of the system.
[0080] In the above embodiments of this disclosure, the method further includes: upon receiving a second prompt message sent by multiple storage clusters, adding a second preset field to the cloud disk creation request to obtain a second creation request; and sending the second creation request to the target cluster so that the target cluster creates a cloud disk based on the second creation request.
[0081] The aforementioned second preset field can be a field that prevents the soft rejection retry request from being made when the soft rejection retry "fallback" mechanism is set in the cloud disk creation request. For example, it can be a field such as check_space:disable or check_space:false. The specific form of the second preset field can be determined according to actual needs and is not limited here.
[0082] In an optional embodiment, if all clusters return the second prompt message, the system can adopt a fallback strategy, selecting the target cluster initially judged as the better cluster and sending a second creation request carrying the second preset field, i.e., check_space:disable or check_space:false. That is, even if the cluster's resource level may be close to or reach a preset threshold, the disk creation operation will be performed directly to ensure that the user's request can be processed. Simultaneously, the frequency and timing of this operation can be strictly controlled to avoid excessive consumption of cluster resources. In the above process, by introducing a "fallback" mechanism, even if all clusters reject the disk creation request due to resource constraints, the system can still ensure that the request is eventually processed, avoiding service interruption and improving user experience. The "fallback" mechanism enhances the system's flexibility, enabling it to react according to real-time resource conditions and cluster status, while also improving system reliability, ensuring that core service functions are not affected even under extreme resource constraints. The synergy between the "retry" mechanism and the "fallback" mechanism finds a balance between resource management and user experience, and by automatically readjusting resource management strategies through recovery strategies, long-term resource optimization and system stability are ensured.
[0083] In the above embodiments of this disclosure, determining the target storage cluster from multiple storage clusters, excluding the source storage cluster, includes: determining the cloud disk type of the cloud disk to be created based on the cloud disk creation request; adjusting the preset scheduling coefficient based on the cloud disk type to obtain the target scheduling coefficient; and determining the target storage cluster from multiple storage clusters, excluding the source storage cluster, based on the target scheduling coefficient.
[0084] The cloud disk types mentioned above refer to the classification of different disks in cloud storage services. The main basis is the performance characteristics and storage capacity of the disk. Cloud disks to be created can be divided into performance-type cloud disks and capacity-type cloud disks. Performance-type cloud disks usually have a high number of input / output operations per second and low latency, and are suitable for scenarios that require frequent read and write operations, such as databases and high-performance computing. Capacity-type cloud disks can provide a large amount of storage space and are suitable for scenarios where a large amount of data is stored but the read and write frequency is not high, such as data backup and log storage.
[0085] The aforementioned preset scheduling coefficients refer to the original parameters used to evaluate and allocate cluster resources during cloud disk creation request processing. They reflect the default importance of cluster performance and capacity resources during scheduling. By default, the preset scheduling coefficients give performance and capacity resources equal weight in the scheduling system, meaning that when selecting a cluster, performance and capacity resources are considered equally by default.
[0086] In one optional embodiment, the system can identify the type of cloud disk, classifying it into performance-oriented and capacity-oriented cloud disks based on performance requirements and storage capacity. After identifying the cloud disk type, the system can adjust preset scheduling coefficients in its scheduling decisions based on the cloud disk type. These preset scheduling coefficients represent the default weights of performance and capacity resources during resource scheduling. For performance-oriented cloud disks, the system can adjust the weight of performance parameters to be greater than that of capacity parameters; for capacity-oriented cloud disks, the system can adjust the weight of capacity parameters to be greater than that of performance parameters, so that the target cluster is more biased towards meeting the specific type requirements of the cloud disk during selection. After obtaining the adjusted target scheduling coefficients, the system can then select a target storage cluster from multiple storage clusters, excluding the source storage cluster. This selection process can be based on the adjusted weights, comparing the performance and capacity resources of each cluster to select the cluster that best matches the cloud disk requirements. For example, for performance-oriented cloud disks, the system may tend to select clusters with sufficient performance resources and relatively less demanding capacity resources. In the above process, by adjusting the scheduling coefficient according to the cloud disk type, the solution can more accurately match the cloud disk demand with the cluster resources, avoid resource tilting and waste, and improve the overall utilization efficiency of resources. By intelligently selecting the target storage cluster, it ensures that different types of cloud disks are reasonably allocated to each cluster, avoids excessive performance or capacity resource burden between clusters, and achieves cluster load balance.
[0087] In the above embodiments of this disclosure, the cloud disk type to be created is determined based on the cloud disk creation request, including parsing the cloud disk creation request to obtain the capability indicators and performance indicators of the cloud disk to be created; if the cloud disk to be created has preset capabilities based on the capability indicators, the cloud disk type is determined to be the first type; if the cloud disk to be created does not have preset capabilities based on the capability indicators, the cloud disk type is determined based on the performance indicators.
[0088] The aforementioned preset capabilities may refer to the cloud disk's burst request processing capability. Burst capability can be the cloud storage system's ability to process a large number of requests in a short period of time. Preset capabilities may also be other capabilities used to determine the type of cloud disk, which can be determined according to actual needs and are not limited here.
[0089] In one optional embodiment, the system can parse the cloud disk creation request to determine the characteristics of the cloud disk to be created. These parameters may include information such as the size and performance requirements of the cloud disk, as well as the capability and performance metrics of the cloud disk to be created. Next, the system can check whether the cloud disk to be created has preset capabilities, which could be a capability that allows the disk to provide performance levels exceeding normal levels for a short period. For cloud disks with preset capabilities, they can be directly identified as the first type, i.e., performance-oriented cloud disks, which have high performance requirements for the cluster. For cloud disks without preset capabilities, the system can further analyze the performance metrics contained in the cloud disk creation request, such as the number of input / output operations per second, latency, and throughput. Based on these performance metrics, the system can determine whether the cloud disk is performance-intensive. If so, it can also be identified as a first-type cloud disk; if not, it can be identified as a second-type cloud disk, i.e., capacity-oriented cloud disk. By accurately determining the cloud disk type, the system can selectively choose the most suitable storage cluster, avoiding unbalanced resource usage and improving overall resource utilization. For performance-oriented cloud disks, clusters with sufficient performance resources are selected; for capacity-oriented cloud disks, clusters with abundant storage capacity are selected, thereby achieving appropriate resource allocation. In this process, different types of cloud disks are assigned to suitable clusters, which helps prevent excessive burden on performance or capacity resources between clusters, achieving load balancing between clusters. The dynamic parameter disk creation scheduling mechanism can intelligently allocate disk creation requests to suitable clusters based on cloud disk type and real-time resource status, thereby improving resource utilization efficiency and avoiding unbalanced resource consumption between clusters. Dynamic parameter disk creation scheduling ensures a reasonable distribution of performance-oriented and capacity-oriented cloud disks across clusters, helping to maintain load balance within clusters, preventing excessive load on certain clusters, and ensuring the stability and reliability of the overall system.
[0090] In the above embodiments of this disclosure, the cloud disk creation request is parsed to determine the performance indicators of the cloud disk to be created, and the cloud disk type is determined based on the performance indicators. This includes: parsing the cloud disk creation request to determine the standard performance indicators, storage capacity, performance increment, and performance threshold of the cloud disk to be created, wherein the standard performance indicators are used to characterize the performance of the cloud disk to be created under normal working conditions; determining the performance indicators based on the standard performance indicators, storage capacity, performance increment, and performance threshold; determining the cloud disk type as a first type if the performance indicators are greater than preset indicators; and determining the cloud disk type as a second type if the performance indicators are less than or equal to preset indicators, wherein the optimization objectives of the second type of cloud disk are different from those of the first type of cloud disk.
[0091] In one optional embodiment, the cloud disk creation request may include the basic specifications and performance requirements of the cloud disk, including standard performance indicators, storage capacity, performance increment, i.e., capacity-based performance improvement rate and performance threshold. The system can parse this information for subsequent type determination and resource scheduling. Based on the parsed standard performance indicators, storage capacity, performance increment, and performance threshold, the system can calculate the actual performance indicators of the cloud disk to be created to evaluate the expected performance of the cloud disk under the current capacity. After calculating the cloud disk's performance indicators, the system can determine the cloud disk type based on these indicators and preset performance thresholds. Specifically, if the actual performance indicators of the cloud disk are greater than the preset performance threshold, the cloud disk type is determined to be the first type, i.e., a performance-based cloud disk; if the actual performance indicators are less than or equal to the preset performance threshold, the cloud disk type is determined to be the second type, i.e., a capacity-based cloud disk. This facilitates the subsequent execution of a dynamic parameter disk creation and scheduling mechanism to ensure that the cloud disk is created on a suitable cluster, optimizing resource allocation; and ensures a reasonable distribution of different types of cloud disks among clusters, avoiding the problem of some clusters being overly concentrated on the same type of cloud disk, leading to resource skew and uneven load. This helps maintain load balancing across clusters, improving the overall stability and reliability of the system. The dynamic parameter disk scheduling mechanism can flexibly respond to changing application needs and system states, ensuring the accuracy and adaptability of resource scheduling by adaptively adjusting scheduling parameters.
[0092] In the above embodiments of this disclosure, the performance index is determined based on the standard performance index, storage capacity, performance increment, and performance threshold, including: obtaining the product of storage capacity and performance increment; obtaining the sum of the standard performance index and the product; obtaining the minimum value between the sum and the performance threshold to obtain the total performance; and obtaining the ratio of the total performance to the storage capacity to obtain the performance index.
[0093] In one optional embodiment, firstly, the system can obtain the storage capacity (size) and performance increment (step) of the cloud disk to be created from the cloud disk creation request. The performance increment refers to the amount by which the cloud disk performance will increase for each additional unit of storage capacity. Multiplying these two values yields step*size, reflecting the direct impact of cloud disk capacity on its performance requirements. Next, the system can calculate the sum of a standard performance metric (base) and the obtained product (step*size). The standard performance metric can refer to the basic performance level of the cloud disk. Adding these two values yields (base + step*size), a step that comprehensively considers both the basic performance requirements of the cloud disk and the additional performance requirements brought about by capacity expansion. Then, the system can obtain the minimum value between the sum and a performance threshold to obtain the total performance: between the obtained sum (base + step*size) and a preset performance threshold (maxPerformance), the system can select the minimum value as the total performance. The performance threshold can be a maximum performance limit set according to the cloud disk type and cluster capabilities, ensuring that the performance requirements of the cloud disk do not exceed the physical limits of the cluster. By selecting the minimum value as the total performance, the system can ensure that the actual performance requirements of the cloud disk are within a reasonable range, avoiding excessive consumption or waste of resources. Finally, the system can calculate the ratio of the total performance to the storage capacity, i.e., performance / size. This ratio can serve as a performance metric. Performance metrics measure the performance level corresponding to each unit of storage capacity and can be used as a basis for distinguishing between performance-oriented and capacity-oriented cloud disks.
[0094] The formula for calculating the overall performance can be expressed as follows: performance = min(base + step * size, maxPerformance);
[0095] Wherein, `performance` represents the total performance, or performance metric; `base` represents the basic performance, or standard performance metric; `size` represents the storage capacity; `step` represents the performance increment per GB of storage; and `maxPerformance` represents the maximum performance of this type of cloud disk, or performance threshold. This formula allows the system to assess the expected performance of the cloud disk at its current capacity.
[0096] In the above embodiments of this disclosure, the initial scheduling coefficient is adjusted based on the cloud disk type to obtain the target scheduling coefficient, including: increasing the target sub-coefficient corresponding to the cloud disk type in the initial scheduling coefficient, and keeping the other coefficients except the target sub-coefficient unchanged, to obtain the target scheduling coefficient.
[0097] In one optional embodiment, the system can parse the request information to determine the cloud disk type, and then obtain the initial scheduling coefficient, which is a set of pre-set parameters used to evaluate the resource status and allocation priority of the storage cluster. Next, the system can adjust the target sub-coefficient based on the cloud disk type. According to the parsed cloud disk type, the scheduling coefficient is adjusted by increasing the target sub-coefficient corresponding to the cloud disk type in the initial scheduling coefficient, while keeping other coefficients unchanged. For example, if it is a performance-type cloud disk, the system can adjust the performance parameter in the scheduling coefficient to twice the capacity parameter, meaning that the performance requirement weight of the performance-type cloud disk will significantly increase when evaluating cluster resources. If it is a capacity-type cloud disk, the system can adjust the capacity parameter to twice the performance parameter, ensuring a significant increase in the storage space weight of the capacity-type cloud disk, while keeping other coefficients unchanged, thus obtaining the target scheduling coefficient. While adjusting the target sub-coefficient, the system can ensure that other coefficients remain unchanged to maintain the comprehensiveness and stability of the scheduling strategy. By increasing the weight of performance parameters for performance-oriented cloud disks or the weight of capacity parameters for capacity-oriented cloud disks, the system can more accurately select the most suitable cluster based on the actual needs of the cloud disks, avoiding unbalanced resource consumption and improving resource allocation efficiency and cluster utilization. Adjusting the scheduling coefficient helps to distribute different types of cloud disks to various clusters, preventing a single cluster from becoming overloaded due to an excessive number of cloud disks of a certain type, maintaining load balance between clusters, improving the overall stability and response speed of the system, and providing strong technical support for resource scheduling in high-load and variable application scenarios in distributed cloud storage environments. This helps to build a more efficient, stable, and intelligent cloud storage resource allocation system.
[0098] The technical solution proposed in this disclosure is described below with reference to an optional embodiment. This disclosure proposes a method and apparatus for optimizing cloud storage capacity and performance guarantees. In a distributed cloud storage environment, an availability zone may have multiple storage clusters providing services to the outside world, and each storage cluster supports different cloud disk types. Since a storage cluster consists of multiple servers, there are certain physical limits, namely, bandwidth and capacity limits. Mapped to the storage environment, this means that the real-time throughput and total write volume of all user disks have a physical limit. To address the rigid guarantee of capacity and performance, this disclosure proposes optimization methods and apparatus for disk creation in three scenarios, based on a global availability zone perspective, to solve the technical problems existing in these three scenarios: rigid guarantee of traffic and capacity for cross-cluster snapshot disk creation, rigid guarantee of data write capacity for disk creation, and capacity and performance coordination between disk creation clusters. Among them, cross-cluster snapshot creation disk traffic capacity guarantee can solve the capacity guarantee problem brought by the main disk; creation disk data writing capacity guarantee can solve the capacity guarantee in non-main disk and main disk quota use scenarios; the capacity and performance coordination problem between creation disk clusters can solve how to guarantee performance guarantee while guaranteeing capacity guarantee. These three scenarios can be regarded as being in one system.
[0099] Currently, in related technologies, there are three scenarios: Scenario 1: Cross-cluster snapshot disk creation traffic and capacity rigidity. Creating disks based on cluster snapshots consumes cross-cluster traffic and causes the snapshot capacity of the target cluster to increase. Scenario 1 can be addressed by setting cross-cluster flow control values, but it is still difficult to avoid unnecessary cross-cluster traffic and capacity rigidity. Scenario 2: Creation disk data write capacity rigidity. When users create disks based on snapshots, a large amount of data is written to the disk instantly. Too much data can lead to the risk of cluster capacity rigidity. Scenario 2 can be addressed immediately by migrating large-capacity disks, but this increases operational complexity, has a certain time sensitivity, and introduces some uncertainty. Scenario 3: Capacity and performance coordination between disk creation clusters. In related technologies, the scheduling system does not predict disk usage and uniformly allocates clusters according to a set of scheduling parameters. This can easily lead to an excessive number of performance-oriented disks or capacity-oriented disks in a cluster. Scenario 3 can be addressed by adding machines to redundancy the cluster's capacity and performance limits, increasing the cluster's tolerance, but this adds significant additional costs.
[0100] This disclosure primarily applies to the cloud disk creation process, therefore it introduces the basic creation workflow, showing how storage disks are created through the following steps: Region-level management component -> Availability Zone-level management component -> Cluster-level management component -> Cluster-level distributed storage components. It focuses on optimizing the scheduling system within the Availability Zone-level management component. This disclosure mainly targets the Availability Zone-level management component service, which is responsible for allocating a suitable cluster during cloud disk creation; this cloud disk service is also the service provided by block storage.
[0101] Figure 5 is a schematic diagram of an optional cloud disk creation process according to an embodiment of the present disclosure. As shown in Figure 5, it includes multiple storage clusters such as storage cluster 1... storage cluster N. Each storage cluster includes a storage cluster-level management and control component and a storage cluster-level distributed storage component. The cloud disk creation process is as follows: storage region-level management and control component -> storage availability zone-level management and control component (scheduling system) -> storage cluster-level management and control component -> storage cluster-level distributed storage component.
[0102] Scenario 1: Cross-cluster snapshot disk creation with guaranteed traffic and capacity. When users create disks via cluster snapshots, the scheduling system selects a suitable cluster based on the "water level" of each cluster, tending to distribute the disks across multiple clusters. Cluster snapshot disk creation involves cross-cluster traffic and capacity consumption. The "water level" here is a comprehensive measure of multiple dimensions of the cluster, including physical data write volume, logical disk creation volume, real-time traffic volume, and the number of disks. Based on these water levels, a score is obtained after normalization. This score is then used to rank and select a superior cluster.
[0103] In the triggering scenario, when a user creates a disk using a cluster snapshot, the scheduling system selects a suitable cluster based on the available resources, and it's highly likely that the disk will not be in the same cluster as the snapshot. In the case of large snapshots, because they are cluster snapshots, to prevent the creation of multiple disks based on the snapshot, the conventional approach is to copy the snapshot data to the destination cluster. This mechanism leads to two problems in large snapshot scenarios: space impact—for 10TB-level snapshots, the space required for the destination cluster increases significantly, while the source cluster, due to hard links between disks, does not occupy much additional space; and traffic impact—the copying process for terabyte-level snapshots consumes bandwidth for both the destination and source clusters, and bandwidth, like physical space, is a rigid constraint.
[0104] Figure 6 is a schematic diagram of an optimized cross-cluster snapshot data creation cloud disk according to an embodiment of the present disclosure. As shown in Figure 6, it includes multiple storage clusters such as storage cluster 1... storage cluster N. Each storage cluster includes a storage cluster-level management and control component and a storage cluster-level distributed storage component. The cloud disk creation process is as follows: storage region-level management and control component -> storage availability zone-level management and control component (scheduling system) -> storage cluster-level management and control component -> storage cluster-level distributed storage component. During the cross-cluster snapshot data creation cloud disk process, the cluster snapshot data generated by storage cluster 1 can be loaded into storage cluster N, that is, cross-cluster snapshot data creation cloud disk is performed. In the optimized cross-cluster snapshot data creation cloud disk process, the scheduling system determines whether the cloud disk creation request is cluster snapshot data. If it is not cluster snapshot data, default scheduling is performed. If it is cluster snapshot data, it determines whether it is a large snapshot. If it is not a large snapshot, default scheduling is performed. If it is a large snapshot, it determines whether the source storage cluster meets the disk creation quota mechanism. If the source storage cluster does not meet the disk creation quota mechanism, default scheduling is performed. If the source storage cluster meets the disk creation quota mechanism, the snapshot data source storage cluster is used to create the disk.
[0105] The main function of this solution is in the scheduling system of the storage availability zone level management component. First, a standard is needed to define what constitutes a large disk. This can be achieved by obtaining statistical data based on historical online data. Actual disk creation requests are statistically analyzed across four dimensions: [0, 5T], [5T, 10T], [10T, 20T], and [20T, +]. Combined with real-world cases impacting space and traffic, it can be determined that the proportion of snapshots larger than 5T is 0.3%. Therefore, 5T is used as the standard for measuring a large disk. Thus, performing affinity scheduling on 0.3% of the disks has little impact on the actual default scheduling, while simultaneously mitigating the traffic and space issues caused by large snapshots creating disks across the cluster.
[0106] Secondly, the affinity strategy allocates disks to the cluster where the snapshot resides. Creating multiple disks based on snapshots without considering cluster affinity can lead to traffic congestion. For example, creating 30 large disks based on a snapshot would cause traffic bottlenecks if users immediately start using the service, impacting user experience. Therefore, a quota mechanism is introduced to determine whether the source cluster can create disks. Quota allocation is determined statistically. Analyzing the number of disks created from snapshots larger than 5TB reveals that more than two disks (1.05%) account for 387. Based on this, the quota principle designed in this publication is as follows: for disk creation based on snapshots, first obtain the size N of the cluster where the snapshot resides; if N is greater than or equal to 48, allocate 2 quotas; if N is less than 48, allocate 1 quota. Each disk creation based on this snapshot consumes one quota until no quota remains. In other words, for clusters smaller than 48, one quota is allocated for each large snapshot; for clusters larger than or equal to 48, two quotas are allocated for each large snapshot. Therefore, this solution can effectively avoid 99% of large snapshot disk creation, thereby reducing unnecessary cluster bandwidth performance and cluster capacity consumption, and avoiding the risk of performance and capacity rigidity. However, how to truly avoid the risk of space rigidity caused by snapshot disk creation and data loading is given in the following scenario two.
[0107] Scenario 2 involves guaranteed data write capacity for disk creation. While Scenario 1 resolves nearly 99% of the issues, various scenarios still arise where disk creation occurs across clusters. For example, the cluster hosting the snapshot might be shut down for sale for other reasons, or the quota for a single snapshot might be exhausted. Furthermore, even when creating disks based on non-cluster snapshots, a large amount of snapshot data can lead to a sudden surge in disk writes. This can impact cluster physical space and cross-cluster network traffic.
[0108] Figure 7 is an interactive diagram of snapshot data creation of a cloud disk according to an embodiment of the present disclosure of a soft rejection retry fallback mechanism. As shown in Figure 7, during the snapshot data creation of the cloud disk using the soft rejection retry fallback mechanism, the storage availability zone level management component, i.e., the scheduling system, interacts with multiple storage clusters, including storage cluster 1... storage cluster N, which have storage cluster level management components. Each storage cluster can perform soft rejection for disk creation. The scheduling system sends a first creation request carrying a first preset field to the storage cluster level management component in storage cluster 1. After receiving the feedback second prompt information... the scheduling system sends a first creation request carrying a first preset field to the storage cluster level management component in storage cluster N. After receiving the feedback second prompt information, the scheduling system can send a second creation request carrying a second preset field to the storage cluster level management component in storage cluster 1 to implement the soft rejection retry fallback mechanism.
[0109] This disclosure introduces a soft rejection retry mechanism. First, at the cluster level, a soft rejection function is provided for disk creation. Whether a check space judgment is needed is explicitly stated when the storage availability zone level management component sends a request to the storage cluster level management component. When the storage availability zone level management component schedules the system to create a disk for the cluster storage cluster level management component, the request protocol includes the `check_space: enable` field. Upon receiving the request, the storage cluster level management component performs a pre-judgment of the cluster's physical water level after successful disk creation. If it finds that adding this disk would cause the cluster's physical water level to reach a risk value, it returns a soft reject error code; otherwise, the disk creation is successful. Alternatively, when the storage availability zone level management component schedules the system to create a disk for the cluster storage cluster level management component, the request protocol includes the `check_space: disable` field. Upon receiving the request, the storage cluster level management component does not perform a water level pre-judgment and proceeds with the normal disk creation process. Second, the scheduling system is optimized to handle scenarios with bandwidth less than 5TB and scenarios with bandwidth greater than 5TB but where the quota has been exhausted. Each scheduling operation is based on the water level algorithm described above. The system selects a relatively optimal cluster (Cluster 1) and attempts to create a disk on Cluster 1 with the `check_space: enable` field. If a soft reject error code is returned, it will try other alternative clusters. If all alternative clusters have been tried and the attempt still fails, it will return to the previously selected optimal cluster (Cluster 1) and create the disk with the `check_space: false` field. This ensures global optimization, meaning that disk creation requests with capacity risks are directed to a suitable cluster. It also serves as a fallback, ensuring that disk creation activities are not hindered even if no suitable cluster is found globally, thus maintaining a smooth user experience.
[0110] Scenario 3 addresses capacity and performance coordination between disk creation clusters. While Scenario 1 and 2 effectively resolve physical space constraints and inter-cluster disk creation traffic issues, they don't consider expected read / write traffic and space consumption, i.e., they don't differentiate between traffic-intensive and capacity-intensive disks. This can lead to a cluster having too many traffic-intensive or too many capacity-intensive cloud disks, resulting in suboptimal availability at the global availability zone level. Some clusters reach capacity constraints but still have sufficient performance, while others reach performance constraints but still have sufficient space. The root cause is that despite current storage clusters supporting all types of disks, disk scheduling still uses uniform parameters. In a mixed deployment context, all disks are scheduled using uniform parameters, leading to situations where some clusters reach capacity constraints but still have sufficient performance, and vice versa.
[0111] Figure 8 is a schematic diagram of creating a cloud disk from snapshot data according to an embodiment of the present disclosure using a dynamic parameter mechanism. As shown in Figure 8, it includes multiple storage clusters such as storage cluster 1... storage cluster N. Each storage cluster includes a storage cluster-level management and control component and a storage cluster-level distributed storage component. The cloud disk creation process is as follows: storage region-level management and control component -> storage availability zone-level management and control component (scheduling system) -> storage cluster-level management and control component -> storage cluster-level distributed storage component. During the creation of the cloud disk from snapshot data using the dynamic parameter mechanism, the scheduling system confirms the disk type based on the cloud disk creation request. If it is capacity-type, the scheduling algorithm is executed based on the capacity-type scheduling parameters; if it is performance-type, the scheduling algorithm is executed based on the performance-type scheduling parameters.
[0112] This disclosure introduces a disk type confirmation mechanism and corresponding adaptive scheduling parameters. First, let's introduce the disk type confirmation mechanism, which categorizes capacity-type disks and performance-type disks based on this mechanism. Currently, cloud disks sold by various cloud vendors can be simply divided into two types: those with Burst capability and those without. Disks with Burst capability are directly classified as performance-type disks; for disks without Burst capability, the relationship between performance and capacity is generally as follows: performance = min(base + step * size, maxPerformance);
[0113] Wherein, performance is the total performance, also known as the performance metric; base is the basic performance, also known as the standard performance metric; size is the storage capacity; step is the performance increment corresponding to each GB of storage; and maxPerformance is the maximum performance of this type of cloud disk, also known as the performance threshold.
[0114] Based on this formula, by analyzing several typical cloud disk types, the relationship between capacity and performance can be obtained, and the trend chart of single GB performance of different types of disks can be derived from the relationship between capacity and performance.
[0115] Figure 9 is a schematic diagram of the relationship between capacity and performance of an optional cloud disk according to an embodiment of the present disclosure. As shown in Figure 9, the horizontal axis represents capacity and the vertical axis represents performance. Different curves in the figure represent different types of cloud disks. That is, five types of cloud disks are shown in the figure. The relationship between cloud disk performance and capacity can be seen from the figure.
[0116] Figure 10 is a schematic diagram of the relationship between single gigabyte performance and capacity of an optional cloud disk according to an embodiment of the present disclosure. As shown in Figure 10, the horizontal axis represents capacity and the vertical axis represents performance / capacity, which can be single gigabyte performance. Different curves in the figure represent different types of cloud disks. That is, five types of cloud disks are shown in the figure. The trend of cloud disk single gigabyte performance with capacity can be seen from the figure.
[0117] As shown in Figure 10, the single-GB performance trend chart reveals that for all disk types, the performance per GB gradually decreases as capacity increases. Therefore, a single-GB performance threshold N can be added to differentiate between two different types of cloud disks. Based on online data, a threshold of 0.5 can be used: disk single-GB performance greater than N indicates a performance-type cloud disk; disk single-GB performance less than or equal to N indicates a capacity-type cloud disk. Secondly, the scheduling parameters in the storage availability zone level management component's scheduling system module are adapted. For performance-type cloud disks, after determining whether it's a capacity-type or performance-type cloud disk in the scheduling coefficient, different scheduling coefficients are configured for different dimensions in the water level scheduling algorithm. For example, the performance parameter will be adjusted to twice the capacity parameter; for capacity-type cloud disks, the capacity parameter will be adjusted to twice the performance parameter. Both the single-GB performance threshold N and the performance-capacity ratio are configurable parameters and can be dynamically adjusted.
[0118] The technical solutions proposed in this disclosure address three scenarios: Scenario 1: Guaranteed capacity for cross-cluster snapshot disk creation. By analyzing online statistical data, scenarios prone to generating large online snapshots are identified. Through affinity scheduling and quota pre-allocation mechanisms, 99% of these scenarios can be resolved, reducing invalid cross-cluster disk creation and thus reducing data replication and inter-cluster traffic. Scenario 2: Guaranteed capacity for disk creation data writes. For disk creation data write scenarios, a soft rejection mechanism is used. Each disk creation attempt checks its impact on the cluster. If potential problems are found, the cluster is switched. If no suitable cluster is found across the entire availability zone, the disk creation attempt is not detected, thus ensuring overall optimization and a better user experience. Scenario 3: Capacity and performance coordination between disk creation clusters. This approach generalizes to all disk creation scenarios, focusing on the timing of disk creation. By introducing a single GB performance metric, non-Burst disks are categorized into performance-oriented and capacity-oriented disks, with Burst disks directly classified as performance-oriented disks. Different scheduling parameters are set for different types of disks to ensure a balance between performance-oriented and capacity-oriented disks within a single cluster, preventing skewed scenarios.
[0119] According to an embodiment of the present disclosure, a distributed storage system is also provided. FIG11 is a structural block diagram of a distributed storage system according to an embodiment of the present disclosure. As shown in FIG11, the system includes: multiple storage clusters 1102 and a scheduling system 1104.
[0120] The scheduling system is connected to multiple storage clusters and is used to execute the methods in the various embodiments of this disclosure.
[0121] Embodiments of this disclosure can provide an electronic device. FIG12 is a structural block diagram of an electronic device according to an embodiment of this disclosure. As shown in FIG12, the electronic device may include: an input / output device 1202; a memory 1204; and a processor 1206, wherein the processor 1206 is connected to the input / output device 1202 and the memory 1204 via a bus 1208.
[0122] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0123] The processor can invoke an executable program stored in memory via a transmission device to perform the following methods: applied to multiple storage clusters, responding to a cloud disk creation request for snapshot data, determining the amount of snapshot data and the source storage cluster storing the snapshot data; if the amount of data is greater than or equal to a preset amount of data and the source storage cluster meets preset resource scheduling conditions, sending a cloud disk creation request to the source storage cluster so that the source storage cluster creates a cloud disk based on the cloud disk creation request; if the amount of data is less than the preset amount of data or the source storage cluster does not meet the preset resource scheduling conditions, determining a destination storage cluster other than the source storage cluster from the multiple storage clusters, and sending a cloud disk creation request to the destination storage cluster so that the destination storage cluster creates a cloud disk based on the cloud disk creation request, and implementing the methods in the various embodiments of this disclosure during execution.
[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0125] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause the processing unit to execute the methods in the various embodiments of this disclosure.
[0127] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0128] Embodiments of this disclosure also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.
[0129] Optionally, in this embodiment, the storage medium may be located in a computing device.
[0130] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, which, when the executable program is running, controls the device where the computer-readable storage medium is located to execute the method described in any of the above embodiments.
[0131] Embodiments of this disclosure also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0132] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.
[0133] Embodiments of this disclosure also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0134] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.
[0135] Embodiments of this disclosure also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.
[0136] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.
[0137] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0138] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0141] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0142] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.
Claims
1. A resource scheduling method, applied to a scheduling system connected to multiple storage clusters, the method comprising: In response to a cloud disk creation request for snapshot data, determine the amount of snapshot data and the source storage cluster where the snapshot data is stored; When the data volume is greater than or equal to the preset data volume and the source storage cluster meets the preset resource scheduling conditions, the cloud disk creation request is sent to the source storage cluster so that the source storage cluster creates a cloud disk based on the cloud disk creation request. If the amount of data is less than the preset amount of data, or if the source storage cluster does not meet the preset resource scheduling conditions, a target storage cluster is determined from the multiple storage clusters, excluding the source storage cluster, and the cloud disk creation request is sent to the target storage cluster so that the target storage cluster can create a cloud disk based on the cloud disk creation request.
2. The method according to claim 1, wherein, Whether the source storage cluster meets the preset resource scheduling conditions is determined in the following manner: Obtain the current resource scheduling quota of the snapshot data; If the current resource scheduling quota is greater than or equal to the preset quota, it is determined that the source storage cluster meets the preset resource scheduling conditions; If the current resource scheduling quota is less than the preset quota, it is determined that the source storage cluster does not meet the preset resource scheduling conditions.
3. The method according to claim 2, wherein, The current resource scheduling quota of the snapshot data is less than or equal to the initial resource scheduling quota of the snapshot data, which is determined based on the total resource amount of the source storage cluster.
4. The method according to any one of claims 1 to 3, wherein, Sending the cloud disk creation request to the source storage cluster, or sending the cloud disk creation request to the destination storage cluster, includes: Add a first preset field to the cloud disk creation request to obtain the first creation request; Send the first creation request to the target cluster, and receive the first prompt message or the second prompt message returned by the target cluster. The first prompt message is returned by the target cluster when it determines that the first resource level has not exceeded the preset level and creates a cloud disk based on the first creation request. The second prompt message is returned by the target cluster when the first resource level exceeds the preset level. Wherein, the target cluster is the source storage cluster or the destination storage cluster, the first resource level is used to characterize the resource level of the target cluster after the cloud disk is created, the first prompt information is used to indicate that the cloud disk was created successfully, and the second prompt information is used to indicate that the cloud disk was created unsuccessfully.
5. The method according to claim 4, wherein, The method further includes: In response to the second prompt message sent by the target cluster, a first storage cluster is determined from the plurality of storage clusters, excluding the storage cluster that received the message from the target cluster; Send the first creation request to the first storage cluster; In response to the second prompt message sent by the first storage cluster, the steps of determining a new storage cluster from among the plurality of storage clusters, excluding the target cluster and the first storage cluster, and sending the first creation request to the new storage cluster are repeated until the first prompt message sent by the new storage cluster is received, or the second prompt message sent by any one of the plurality of storage clusters is received.
6. The method according to claim 5, wherein, The method further includes: Upon receiving the second prompt information sent by the multiple storage clusters, a second preset field is added to the cloud disk creation request to obtain a second creation request; Send the second creation request to the target cluster so that the target cluster creates a cloud disk based on the second creation request.
7. The method according to any one of claims 1 to 3, wherein, Determining the destination storage cluster from the plurality of storage clusters, excluding the source storage cluster, includes: Based on the cloud disk creation request, determine the type of cloud disk to be created; Based on the cloud disk type, the preset scheduling coefficient is adjusted to obtain the target scheduling coefficient; Based on the target scheduling coefficient, the destination storage cluster is determined from the plurality of storage clusters and the storage clusters other than the source storage cluster.
8. The method according to claim 7, wherein, The process of determining the type of cloud disk to be created based on the cloud disk creation request includes: The cloud disk creation request is parsed to obtain the capability and performance metrics of the cloud disk to be created. If the cloud disk to be created is determined to have preset capabilities based on the capability indicators, the cloud disk type is determined to be the first type; If it is determined based on the capability indicators that the cloud disk to be created does not have the preset capabilities, the cloud disk type is determined based on the performance indicators.
9. The method according to claim 8, wherein, The process of parsing the cloud disk creation request, determining the performance metrics of the cloud disk to be created, and determining the cloud disk type based on the performance metrics includes: The cloud disk creation request is parsed to determine the standard performance indicators, storage capacity, performance increment, and performance threshold of the cloud disk to be created. The standard performance indicators are used to characterize the performance of the cloud disk to be created under normal working conditions. The performance metrics are determined based on the standard performance metrics, the storage capacity, the performance increment, and the performance threshold. If the performance index is greater than the preset index, the cloud disk type is determined to be the first type; If the performance index is less than or equal to the preset index, the cloud disk type is determined to be the second type, wherein the optimization objective of the second type of cloud disk is different from that of the first type of cloud disk.
10. The method according to claim 9, wherein, The process of determining the performance metrics based on the standard performance metrics, the storage capacity, the performance increment, and the performance threshold includes: Obtain the product of the storage capacity and the performance increment; Obtain the sum of the product of the standard performance index and the product; The minimum value between the sum and the performance threshold is obtained to get the total performance. The ratio of the total performance to the storage capacity is used to obtain the performance index.
11. The method according to claim 7, wherein, The adjustment of the initial scheduling coefficient based on the cloud disk type to obtain the target scheduling coefficient includes: The target scheduling coefficient is obtained by increasing the target sub-coefficient corresponding to the cloud disk type in the initial scheduling coefficient and keeping the other coefficients unchanged.
12. A distributed storage system, comprising: Multiple storage clusters; The scheduling system, connected to the plurality of storage clusters, is configured to perform the following method: responding to a cloud disk creation request for snapshot data, determining the amount of snapshot data and the source storage cluster storing the snapshot data; and, if the amount of data is greater than or equal to a preset amount of data and the source storage cluster meets preset resource scheduling conditions, sending the cloud disk creation request to the source storage cluster so that the source storage cluster creates a cloud disk based on the cloud disk creation request. If the amount of data is less than the preset amount of data, or if the source storage cluster does not meet the preset resource scheduling conditions, a target storage cluster is determined from the multiple storage clusters, except for the source storage cluster, and the cloud disk creation request is sent to the target storage cluster so that the target storage cluster creates a cloud disk based on the cloud disk creation request. This method is applied to scheduling systems that connect to multiple storage clusters.
13. The distributed storage system according to claim 12, wherein, The scheduling system is also configured to perform the following method: obtain the current resource scheduling quota of the snapshot data; if the current resource scheduling quota is greater than or equal to a preset quota, determine that the source storage cluster meets the preset resource scheduling conditions; If the current resource scheduling quota is less than the preset quota, it is determined that the source storage cluster does not meet the preset resource scheduling conditions.
14. The distributed storage system according to claim 12, wherein, The current resource scheduling quota of the snapshot data is less than or equal to the initial resource scheduling quota of the snapshot data, which is determined based on the total resource amount of the source storage cluster.
15. The distributed storage system according to any one of claims 12 to 14, wherein, The scheduling system is further configured to execute the following method: adding a first preset field to the cloud disk creation request to obtain a first creation request; sending the first creation request to the target cluster, and receiving a first prompt message or a second prompt message returned by the target cluster, wherein the first prompt message is returned by the target cluster after determining that the first resource level has not exceeded the preset level and based on the first creation request to create the cloud disk, and the second prompt message is returned by the target cluster when the first resource level exceeds the preset level; wherein the target cluster is the source storage cluster or the destination storage cluster, the first resource level is used to characterize the resource level of the target cluster after the cloud disk is created, the first prompt message is used to indicate that the cloud disk creation was successful, and the second prompt message is used to indicate that the cloud disk creation failed.
16. The distributed storage system according to claim 15, wherein, The scheduling system is also configured to perform the following method: in response to the second prompt information sent by the target cluster, determine a first storage cluster from the plurality of storage clusters, excluding the storage cluster that received the target cluster; Send the first creation request to the first storage cluster; In response to the second prompt message sent by the first storage cluster, the steps of determining a new storage cluster from among the plurality of storage clusters, excluding the target cluster and the first storage cluster, and sending the first creation request to the new storage cluster are repeated until the first prompt message sent by the new storage cluster is received, or the second prompt message sent by any one of the plurality of storage clusters is received.
17. The distributed storage system according to claim 16, wherein, The scheduling system is also configured to perform the following method: upon receiving the second prompt information sent by the plurality of storage clusters, add a second preset field to the cloud disk creation request to obtain a second creation request; Send the second creation request to the target cluster so that the target cluster creates a cloud disk based on the second creation request.
18. The distributed storage system according to any one of claims 12 to 14, wherein, The scheduling system is also configured to perform the following method: based on the cloud disk creation request, determine the cloud disk type of the cloud disk to be created; based on the cloud disk type, adjust the preset scheduling coefficient to obtain the target scheduling coefficient; Based on the target scheduling coefficient, the destination storage cluster is determined from the plurality of storage clusters and the storage clusters other than the source storage cluster.
19. An electronic device comprising: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 11.
20. A computer-readable storage medium comprising a stored executable program, wherein, When the executable program is executed, it controls the device containing the storage medium to perform the method described in any one of claims 1 to 11.