An intelligent scheduling and allocation method for container local storage applied to a cloud platform
By introducing a combination solution of storage scheduling plug-in, disk device manager and distributed database, the optimal matching algorithm and Cgroup blkio technology are used to solve the problem of low storage utilization in container local storage scheduling and allocation, and efficient storage resource management and isolation between containers are achieved, and the read and write needs of high-performance containers are met.
Patent Information
- Application Number
- CN202211260655.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-10-14
AI Technical Summary
The existing container local storage scheduling and allocation solutions cannot effectively support diversified storage usage methods, and cannot achieve storage isolation and bandwidth limitations, resulting in low storage utilization and cannot meet the read and write performance requirements of high-performance containers.
Components composed of storage scheduling plug-in, disk device manager and distributed database are used to perform container scheduling and storage allocation through the optimal matching algorithm, supporting bare disk exclusiveness and sharing, and quota restrictions. IO bandwidth control is carried out in combination with Cgroup blkio technology to ensure efficient scheduling and isolation of storage resources.
A diversified storage allocation form is realized, the storage utilization rate is improved, data loss is prevented during container restart, storage isolation between containers and IO bandwidth limitations are ensured, and storage allocation speed and overall performance are improved.
Smart Images

Figure CN115756726B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer cloud computing, and particularly to an intelligent scheduling and allocation method for container local storage applied to a cloud platform. Background Art
[0002] In a container cloud platform, the resources relied on by container scheduling mainly include CPU, memory, GPU, and storage. Among them, CPU, memory, and GPU are stateless resources. The internal differences of these resources are not perceived by containers. Containers that only use these resources can drift between eligible nodes in the cluster. Storage resources are stateful resources. In the field of cloud computing, according to the relative position between storage and containers, it can be divided into network storage and local storage. Network storage is suitable for the scenario of container drift. Due to its high dependence on the unreliable infrastructure of the network, and the performance of existing network storage solutions is lower than that of local storage, containers with performance and high-reliability storage requirements can only use node local storage. Containers using local storage resources will save data in the disk of the node where the container runs. To prevent data loss, such containers need to control drift.
[0003] Currently, the use of local storage resources by containers mainly includes two methods:
[0004] Host Path Mode: Mount a specified directory on the host directly into the container. Writing operations on the mounted directory in the container will save data under the host directory. The advantages of this method are simple implementation and high read and write efficiency. Its disadvantages include the following three points: First, container isolation cannot be achieved. If the same directory is mounted into different containers, they will have the same storage view, and the operation results in one container will be perceived by another container, posing a data security risk; second, container drift cannot be prevented (drift means data loss); third, storage quotas and IO bandwidth cannot be restricted, and the capacity of the mounted directory cannot be limited, which may cause the disk space to be exhausted by a single container.
[0005] (2) Local storage mode: Based on the host path mode, the storage is abstracted into virtual resources and an adaptation layer is added, and the storage scheduling is completed by the adaptation layer. This method solves the problems of storage isolation and container drift in the host path mode, but also introduces other problems: First, it is still impossible to limit the quota and IO bandwidth, and the container will write to the disk without limit until the space is exhausted, and it is impossible to ensure the allocation of IO bandwidth according to the container priority; Second, the storage is abstracted into virtual resources, and when performing storage scheduling and allocation, virtual resources need to be created, which increases the complexity of container scheduling and startup, resulting in slow storage allocation speed and difficult data synchronization; Third, most implementations of this mode merge physical volumes (PVs, Physical Volumes) into one or more volume groups (VGs, Volume Groups) and then divide them into logical volumes (LVs, Logical Volumes) according to container requirements. After merging multiple disks, LVM masks the differences between different disks. On the one hand, it makes it impossible to limit the IO bandwidth of a certain disk, and on the other hand, it limits the diversified use methods of disks and does not support allocation in the form of exclusive and shared use of raw disks. For containers with higher performance requirements, raw disks are needed. In addition, making LVM will initialize the disk, resulting in the loss of original data and making it impossible to perform in-place upgrades of container local storage.
[0006] In summary, there are obvious defects in the two current mainstream local storage scheduling and allocation schemes. These defects limit the storage control ability and storage utilization rate of lightweight cloud platforms and cannot meet the read and write performance requirements of high-performance containers for storage. There is an urgent need for a container local storage scheduling and allocation scheme that can support diversified container local storage use methods and has functions such as storage quota, storage isolation, and bandwidth limitation. Summary of the Invention
[0007] The present invention aims to solve the problems existing in the current lightweight cloud platform, and proposes a local storage scheduling and allocation method applicable to lightweight cloud platforms, which supports multiple allocation forms of container local storage, realizes efficient scheduling of local storage resources, prevents drift during container restart, ensures storage isolation between containers, and realizes container storage quota limitation and IO bandwidth limitation.
[0008] The present invention adopts the following technical solutions to solve the above technical problems:
[0009] A method for intelligent scheduling and allocation of container local storage in a cloud platform specifically includes three processes: disk information collection and reporting, container scheduling and storage allocation, and container destruction and storage recovery, which are completed by three components: a storage scheduling plugin, a disk device manager, and a distributed database;
[0010] Among them, the storage scheduling plugin, as an extended plugin of the container cloud engine scheduling center, is used to complete the intelligent scheduling of containers based on container storage requests, the topological structure between containers, the storage resources of each node, and the storage binding situation of scheduled containers;
[0011] The disk device manager running on each node is responsible for collecting storage mount information, volume preparation, volume mounting and recycling, and limiting the number of bytes per second and the number of I / O operations per second for container reading and writing;
[0012] The distributed database is used to store the status information of the cluster, including the total storage of each node, container and binding information, and the affinity and anti-affinity requirements of containers, providing a decision for allocation and scheduling;
[0013] The disk information collection and reporting process: The disk device manager running on each node collects disk information and reports it to the container cloud engine, and the engine stores the reported node disk information in the distributed database;
[0014] The scheduling and storage allocation process of containers:
[0015] After the container creation command is issued, the container is in the pending scheduling container queue. The containers in the queue are sorted according to priority. The scheduler takes out the pending scheduling containers from the queue for scheduling, and initially filters the nodes based on the CPU, memory, ports, and labels of each node. The filtered node list is passed to the storage scheduling plugin;
[0016] The storage scheduling plugin first checks whether the current container has been scheduled. If there is binding information that matches the container being scheduled, it uses this binding information to complete the scheduling. This scheduling behavior maintains the container storage state and ensures that the storage is not lost when the container is restarted multiple times;
[0017] If there is no matching binding information, it dynamically obtains the storage information of each node in the cluster, the storage information of the already running containers, and the affinity and anti-affinity requirements from the distributed database, and combines the disk requests, affinity and anti-affinity requirements of the container being scheduled. After comprehensive calculation, it obtains the disk allocation method, usage amount, remaining amount of each node, and the topological relationship between containers with storage requests;
[0018] The container destruction and storage recycling process:
[0019] Filter and sort the nodes by analyzing the affinity and anti-affinity between the container being scheduled and the already scheduled and undeleted containers;
[0020] Among them, for containers with anti-affinity requirements, the storage scheduling plugin filters out the nodes where the containers with anti-affinity conflicts are running;
[0021] For containers with affinity requirements, the scheduler scores and sorts the nodes according to the affinity scoring strategy.
[0022] As a further preferred solution of the intelligent scheduling and allocation method for container local storage applied to the cloud platform in the present invention, in the container destruction and storage recycling process, an optimal matching algorithm is used for storage allocation, specifically as follows:
[0023] The input of the optimal matching algorithm is the container local storage request, the node list, and the disk information on each node;
[0024] The output is the name of the node to which the container is to be bound and the disk list.
[0025] As a further preferred solution of the intelligent scheduling and allocation method for container local storage applied to the cloud platform in the present invention, the storage requests in the input of the optimal matching algorithm are divided into two categories: raw disks and quotas according to whether they are for raw disk allocation. The disk screening conditions for both include disk type, disk quantity, capacity, and bandwidth;
[0026] Raw disk allocation is further divided into raw disk sharing and raw disk exclusivity. Raw disk sharing means that multiple containers share the capacity of a certain disk, and the storage between these containers is isolated from each other;
[0027] The upper limit of the number of shared containers that can be configured for each disk;
[0028] Raw disk exclusivity is exclusive. The disk that has been allocated to a certain container in an exclusive form cannot be bound to other containers for use. The additional screening conditions for the container with a raw disk request include the number of disks and the minimum disk size; the container with a quota request needs to add the requested quota size;
[0029] If the container local storage request being scheduled is for raw disk exclusivity, the node disk information is taken out from the node list in order, and it is judged whether it meets the container's disk type, disk number, and minimum disk size requests. If it meets, the node and the disk number are returned;
[0030] If the request is for raw disk sharing, the node list is re-sorted from largest to smallest according to the number of shared but not yet reaching the sharing limit disks that meet the container requirements, and the disks on each node are sorted according to the sharing times;
[0031] Check whether the node and the disk meet the container request in order. If it meets, the node name and the disk list are returned; if the request is for quota allocation, the disk lists after excluding whole disk allocation for each node are sorted from smallest to largest according to the remaining amount;
[0032] First, check the nodes in sequence. Analyze whether there is a disk that has been allocated according to the quota and meets the requests of the containers being scheduled. If there is, return the node name and disk number; if after checking all nodes, a satisfactory node and disk still cannot be obtained, then during the second check of the nodes, obtain a disk that meets the requests and has not been used for allocating raw disks.
[0033] The logic of the allocation process is as follows:
[0034] For raw disk sharing, prefer disks that have been shared before.
[0035] For quota allocation, prefer disks that have been allocated quotas and have the smallest remaining margin.
[0036] Among them, the output result of the optimal matching algorithm is the node name and disk list. This result, combined with the basic information of the containers being scheduled and the disk requests, forms binding information and is submitted to the distributed database.
[0037] The binding information will only be cleared when the container is completely deleted.
[0038] When the container restarts, the scheduling and storage allocation of the container are completed according to the binding information recorded in the distributed database to maintain the storage state of the container.
[0039] As a further preferred solution of the intelligent scheduling and allocation method for container local storage applied to the cloud platform in the present invention, the disk device manager running on each node, as the implementer of storage management, undertakes the responsibilities of disk collection and reporting, allocation and recycling, mounting and unmounting, and container IO bandwidth limitation; specifically as follows:
[0040] When the node is incorporated into the container cloud platform, the disk device manager collects the type, capacity, and remaining amount of the node disks and reports them to the container cloud engine, and the latter records them in the distributed database.
[0041] When a container with a local storage request completes scheduling and writes the binding information into the distributed database, this binding information will be pushed to the disk device manager.
[0042] The disk device manager first obtains the corresponding disk according to the disk number in the binding information, prepares to mount the volume according to the local storage usage method of the container. If it is a raw disk, create a mounting directory; if it is a quota, use the quota technology to isolate the corresponding size of storage.
[0043] Mount the storage volume to the container.
[0044] Complete the IO bandwidth configuration based on the Cgroup blkio subsystem.
[0045] Recycling is the reverse process of allocation. When the disk device manager detects a storage unbinding event, it recycles the storage and unmounts it. When the recycling is completed, the disk device manager notifies the Container Cloud Engine to return the resources, and the Container Cloud Engine deletes the binding information of the container from the distributed database.
[0046] As a further preferred solution of a method for intelligent scheduling and allocation of container local storage applied to a cloud platform according to the present invention, the implementation of the storage scheduling plugin specifically includes the following steps:
[0047] Step 1, receiving and processing a scheduling request from the Container Cloud Engine scheduler:
[0048] The request received by the storage scheduling plugin includes a list of available nodes and container information. The storage scheduling plugin first obtains the current storage topology information of the cluster, node storage information, and binding information of the running containers from the distributed database; then, queries whether there is a scheduling record for the current container. If found, it returns the binding information. If not found, it completes disk allocation according to the optimal matching algorithm and writes the binding information into the distributed database; finally, returns the scheduling result to the scheduler of the container engine.
[0049] Step 2, implementing the scheduling process:
[0050] Perform topology calculation between the container being scheduled and the scheduled and undeleted containers to obtain a list of nodes sorted by priority, and then perform scheduling according to the optimal matching algorithm.
[0051] Step 3, reading and writing to the distributed database:
[0052] After the scheduling is completed, write the scheduling information into the distributed database for storage reservation. At this time, the storage is not completely allocated. The reservation can prevent resource competition. If the subsequent allocation fails, resource return is required; if the container enters the deletion process, the database content needs to be cleared to complete resource return.
[0053] As a further preferred solution of a method for intelligent scheduling and allocation of container local storage applied to a cloud platform according to the present invention, the implementation of the storage scheduling plugin specifically includes the following steps:
[0054] Step 1, receiving and processing a scheduling request from the Container Cloud Engine scheduler:
[0055] The request received by the storage scheduling plugin includes a list of available nodes and container information. The storage scheduling plugin first obtains the current storage topology information of the cluster, node storage information, and binding information of the running containers from the distributed database; then, queries whether there is a scheduling record for the current container. If found, it returns the binding information. If not found, it completes disk allocation according to the optimal matching algorithm and writes the binding information into the distributed database; finally, returns the scheduling result to the scheduler of the container engine.
[0056] Step 2, implement the scheduling process:
[0057] Perform topology calculation between the containers being scheduled and the scheduled containers that have not been deleted, obtain the node list sorted by priority, and then develop the optimal matching algorithm according to the process shown below; Figure 3 Develop the optimal matching algorithm according to the process shown.
[0058] Step 3, read and write to the distributed database:
[0059] After the scheduling is completed, write the scheduling information to the distributed database for storage reservation. At this time, the storage is not fully allocated. The reservation can prevent resource competition. If the subsequent allocation fails, resource return is required; if the container enters the deletion process, the database content needs to be cleared to complete resource return.
[0060] As a further preferred solution of the intelligent scheduling and allocation method for container local storage applied to the cloud platform in the present invention, the implementation of the disk device manager StorageAgent specifically includes the following steps:
[0061] Step 1, report storage information:
[0062] When the disk device manager StorageAgent starts, it needs to push the disk information of the node to the container cloud engine, and this information is stored in the distributed database for use by the storage scheduling plugin during container scheduling;
[0063] Step 2, listen for container and storage binding events:
[0064] When a storage binding event is detected, it is necessary to implement the creation of the mounted volume, the mounting of the container directory and the host directory, the limitation of the container IO bandwidth, and report the results of the allocation and mounting to the container cloud engine;
[0065] Step 3, listen for container and storage unbinding events:
[0066] When an unbinding event is detected, it is necessary to implement the recycling of the mounted volume, the unmounting of the container directory and the host directory, and report the results of the recycling and unmounting.
[0067] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0068] 1. In the storage scheduling plugin of the present invention, through the optimal matching algorithm, it supports diverse disk allocation requests. Disk scheduling and allocation can be performed according to quotas, exclusive use of raw disks, and shared use of raw disks. On the basis of the above usage methods, multiple screening methods can be added to meet more refined requirements;
[0069] 2. In the storage scheduling plugin of the present invention, through the optimal matching algorithm, the scheduling process is simplified, the scheduling speed is accelerated, and the fragmentation of node storage resources can be effectively reduced.
[0070] 3. The storage status information of the present invention is stored in a distributed database and dynamically calculated according to the overall storage status of the cluster during storage allocation, preventing the phenomenon of resource out-of-sync, making this solution applicable to large-scale clusters.
[0071] 4. Through the storage quota limit technology of the present invention, a fixed-size storage space is allocated to the container to prevent the container from overusing disk space.
[0072] 5. Through the Cgroup blkio technology of the present invention, according to the priority of the container, IO bandwidth is allocated to the container to prevent low-priority containers from occupying too much bandwidth and affecting the operation of high-priority containers.
[0073] 6. The present invention supports various allocation forms of container local storage, realizes the efficient scheduling of local storage resources, prevents drift during container restart, ensures storage isolation between containers, and realizes container storage quota limit and IO bandwidth limit. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is a schematic diagram of the scheduling and storage allocation of a container with a local storage request in the present invention;
[0075] Figure 2 It is a schematic diagram of the storage scheduling process of the present invention;
[0076] Figure 3 It is a schematic diagram of the optimal matching algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings:
[0078] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0079] A local storage scheduling and allocation method applicable to lightweight cloud platforms. By introducing a storage scheduling plugin, storage resources can participate in container scheduling like other resources (CPU, memory, GPU). The storage allocation algorithm adopts an optimal matching algorithm, which on the one hand supports diverse local storage usage methods, and on the other hand can reduce disk fragmentation and improve disk utilization; a disk device manager (StorageAgent) deployed on each node is introduced for disk reporting, allocation, IO bandwidth limitation, and recycling.
[0080] As Figure 1 shown, it mainly includes three processes, namely disk information collection and reporting, container scheduling and storage allocation, and container destruction and storage recycling. The above three processes are completed by three components: a storage scheduling plugin, a disk device manager StorageAgent, and a distributed database.
[0081] Among them, the storage scheduling plugin, as an extended plugin of the container cloud engine scheduling center, completes the intelligent scheduling of containers based on container storage requests, the topological structure between containers, the storage resources of each node, and the storage binding situation of scheduled containers.
[0082] The disk device manager StorageAgent running on each node is responsible for collecting storage mount information, volume preparation, volume mounting and recycling, and limiting the number of bytes per second (byte per second, bps) and the number of I / O operations per second (ioper second, iops) for container read and write.
[0083] The storage status information of the distributed database cluster, such as the total storage of each node, container and binding information, and the affinity and anti-affinity requirements of containers, is stored in the distributed database to provide a decision for storage allocation scheduling.
[0084] In the reporting process, the disk device manager StorageAgent running on each node collects disk information and reports it to the container cloud engine, and the engine stores the reported node disk information in the distributed database.
[0085] This disk reporting method is dynamic. During system operation, each node can expand or contract the disk as needed, and the disk device manager StorageAgent will dynamically perceive storage change information and report it correctly.
[0086] The scheduling and storage allocation process of containers with local storage requests is as Figure 1 shown.
[0087] After the container creation command is issued, the container is in the pending scheduling container queue. The containers in the queue are sorted according to their priorities. The scheduler takes out the containers to be scheduled from the queue and performs scheduling. Based on the resource information of each node such as CPU, memory, ports, and labels, a preliminary screening of the nodes is carried out. The list of screened nodes is passed to the storage scheduling plugin.
[0088] The storage scheduling plugin first checks whether the current container has been scheduled. If there is binding information that matches the container being scheduled, then this binding information is used to complete the scheduling. This behavior can maintain the container's storage state and ensure that the storage is not lost when the container is restarted multiple times.
[0089] If there is no matching binding information, then based on the storage information of each node in the cluster dynamically obtained from the distributed database, the storage information of the already running containers, and the affinity and anti-affinity requirements, combined with the disk requests, affinity, and anti-affinity requirements of the container being scheduled, information such as the allocation method, usage amount, remaining amount of disks on each node, and the topological relationship between the containers with storage requests is obtained through comprehensive calculation for the storage scheduling plugin to work with.
[0090] Considering the complexity of the container topology, the diversity of storage usage methods, and the possible resource waste caused by storage resource allocation, the following process and algorithm are used to complete the scheduling of containers with local storage requests.
[0091] In the container topology structure analysis, the nodes are mainly filtered and sorted by analyzing the affinity and anti-affinity between the container being scheduled and the already scheduled and not deleted containers.
[0092] For containers with anti-affinity requirements, the storage scheduling plugin filters out the nodes where the containers with anti-affinity conflicts are running; for containers with affinity requirements, the scheduler scores and sorts the nodes according to the affinity scoring strategy.
[0093] To support the diverse local storage usage methods of containers and solve the problem of low disk utilization caused by the frequent creation and deletion of containers with local storage requests, it is necessary to use Figure 3 the optimal matching algorithm shown for storage allocation.
[0094] Topology-aware storage allocation, topology domain, according to topology affinity, rack, network, computer room, region.
[0095] The input of the optimal matching algorithm is the container's local storage request, the node list, and the disk information on each node.
[0096] The output is the names of the nodes to be bound by the container and the disk list. The storage requests in the algorithm input can be divided into two categories: raw disks and quotas according to whether they are allocated for raw disks. The disk screening conditions for both include disk types (such as SATA, SAS, and SSD), the number of disks, capacity, bandwidth, etc. (supporting fuzzy matching, best matching within a specified range, and full disk matching algorithm).
[0097] Raw disk allocation is further divided into raw disk sharing and raw disk exclusivity. Raw disk sharing means that multiple containers jointly use the capacity of a certain disk, and the storage between these containers is isolated from each other.
[0098] The upper limit of the number of shared containers can be configured for each disk. Raw disk exclusivity is exclusive. The disk that has been allocated to a certain container in an exclusive form cannot be used for other purposes.
[0099] The screening conditions that a container with a raw disk request can attach include the number of disks, the minimum disk size, etc. A container with a quota request needs to attach the requested quota size.
[0100] If the local storage request of the container being scheduled is for raw disk exclusivity, the node disk information is taken out from the node list in order, and it is judged whether it meets the container's disk type, the number of disks, and the minimum disk size request. If it meets, the node and the disk number are returned. If the request is for raw disk sharing, the node list is re-sorted in descending order according to the number of shared but not reaching the sharing upper limit disks that meet the container requirements, and the disks on each node are sorted according to the sharing times. Check whether the node and the disk meet the container request in order. If it meets, the node name and the disk list are returned.
[0101] If the request is for quota allocation, the disk list after excluding full disk allocation for each node is sorted in ascending order according to the remaining amount.
[0102] First, check the nodes in order and analyze whether there are disks on the nodes that have been allocated according to the quota and meet the request of the container being scheduled. If there are, the node name and the disk number are returned; if the nodes and disks that meet the requirements still cannot be obtained after checking all the nodes, when checking the nodes for the second time, obtain the disks that meet the request and have not been used for raw disk allocation.
[0103] The logic of the above allocation process is as follows:
[0104] For raw disk sharing, disks that have been shared are preferred.
[0105] For quota allocation, disks that have been allocated quotas and have the smallest remaining amount are preferred.
[0106] The optimal matching algorithm can obtain the optimal allocation result with relatively small computational cost, can reduce disk fragmentation to a certain extent, increase the utilization rate of the disk resource pool, and make the algorithm applicable to large-scale clusters. Compared with the method of virtualizing disk resources in other open-source solutions, this algorithm has a faster scheduling speed.
[0107] The output result of the above algorithm is the node name and the disk list. This result, combined with the basic information of the container being scheduled and the disk requests, forms binding information and is submitted to the distributed database.
[0108] The binding information will only be cleared when the container is completely deleted. When the container restarts, the scheduling and storage allocation of the container are completed according to the binding information recorded in the distributed database to maintain the storage state of the container.
[0109] The distributed database StorageAgent running on each node, as the implementer of storage management, is responsible for disk collection and reporting, allocation and recycling, mounting and unmounting, and container IO bandwidth limitation.
[0110] When a node is incorporated into the container cloud platform, the distributed database StorageAgent collects the type, capacity, and remaining amount of the node's disks and reports them to the container cloud engine, which records them in the distributed database.
[0111] When a container with a local storage request completes scheduling and writes the binding information to the distributed database, this binding message will be pushed to the distributed database StorageAgent.
[0112] The distributed database StorageAgent first obtains the corresponding disk according to the disk number in the binding information, prepares the mounted volume according to the local storage usage method of the container. If it is a raw disk, it creates a mount directory. If it is a quota, it uses the quota technology to isolate the corresponding size of storage. Then it mounts the storage volume to the container.
[0113] Finally, the IO bandwidth configuration is completed based on the Cgroup blkio subsystem. Compared with network storage, the container storage allocation scheme based on quota and raw disk in this solution has a faster disk read and write speed. In the same physical environment and the same container, the read and write speed of raw disk mounting is 20% faster than that of network storage, and the read and write speed based on quota isolation is 15% faster than that of network storage.
[0114] Recycling is the reverse process of allocation. When the distributed database StorageAgent monitors a storage unbinding event, it recovers the storage and unmounts it. When the recovery is completed, the distributed database StorageAgent notifies the container cloud engine to return the resources, and the container cloud engine deletes the binding information of the container from the distributed database.
[0115] The implementation process includes the implementation of the storage scheduling plugin and the implementation of the node distributed database StorageAgent.
[0116] The implementation of the storage scheduling plugin has 3 steps:
[0117] Step 1. Receive and process scheduling requests from the container cloud engine scheduler:
[0118] The requests received by the storage scheduling plugin include the list of available nodes and container information. The storage scheduling plugin first obtains the current storage topology information of the cluster, node storage information, and the binding information of the running containers from the distributed database. Then, query whether there is a scheduling record for the current container. If found, return the binding information. If not found, complete the disk allocation according to the optimal matching algorithm and write the binding information into the distributed database. Finally, return the scheduling result to the scheduler of the container engine.
[0119] Step 2. Implement Figure 2 the scheduling process shown in
[0120] First, perform the topology calculation between the container being scheduled and the scheduled but not deleted containers to obtain the list of nodes sorted by priority. Then, develop the optimal matching algorithm according to the Figure 3 process shown in
[0121] Step 3. Read and write to the distributed database:
[0122] After completing the scheduling, write the scheduling information into the distributed database for storage reservation. At this time, the storage is not completely allocated. The reservation can prevent resource competition. If the subsequent allocation fails, resource return is required. If the container enters the deletion process, the database content needs to be cleared to complete the resource return.
[0123] The implementation of StorageAgent has 3 steps:
[0124] Step 1, report storage information:
[0125] When StorageAgent starts, it needs to push the disk information of the node to the container cloud engine. This information is stored in the distributed database for use by the storage scheduling plugin during container scheduling.
[0126] Step 2, listen for container and storage binding events:
[0127] When a storage binding event is listened for, it is necessary to implement the creation of the mounted volume, the mounting of the container directory and the host directory, the limitation of the container IO bandwidth, and report the results of the allocation and mounting to the container cloud engine.
[0128] Step 3, listen for container and storage unbinding events:
[0129] When the unbinding event is monitored, it is necessary to implement the recycling of the mounted volume and the unmounting of the container directory and the host directory. And report the results of the recycling and unmounting.
[0130] In the storage scheduling plugin of the present invention, through the optimal matching algorithm, it supports diverse disk allocation requests. Disk scheduling and allocation can be performed according to quotas, exclusive use of raw disks, and shared use of raw disks. On the basis of the above usage methods, multiple screening methods can be added to meet more refined requirements; through the optimal matching algorithm in the storage scheduling plugin, the scheduling process is simplified and the scheduling speed is accelerated, which can effectively reduce the fragmentation of node storage resources; the storage status information is stored in a distributed database and dynamically calculated according to the overall storage status of the cluster during storage allocation to prevent the phenomenon of resource out-of-sync, making this solution applicable to large-scale clusters; through the storage quota limit technology, a fixed-size storage space is allocated to the container to prevent the container from overusing disk space; through the Cgroup blkio technology, according to the priority of the container, IO bandwidth is allocated to the container to prevent low-priority containers from occupying too much bandwidth and affecting the operation of high-priority containers; it supports various allocation forms of container local storage, realizes the efficient scheduling of local storage resources, prevents drift when the container restarts, ensures storage isolation between containers, and realizes container storage quota limit and IO bandwidth limit.
[0131] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0132] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention. The above has made a detailed description of the embodiments of the present invention, but the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A method for intelligent scheduling and allocation of container local storage applied to a cloud platform, characterized in that: Specifically, it includes three processes of disk information collection and reporting, container scheduling and storage allocation, and container destruction and storage recycling, which are completed by three components: the storage scheduling plugin, the disk device manager, and the distributed database; Among them, the storage scheduling plugin, as an extended plugin of the container cloud engine scheduling center, is used to complete the intelligent scheduling of containers based on container storage requests, the topological structure between containers, the storage resources of each node, and the storage binding situation of scheduled containers; The disk device manager running on each node is responsible for collecting storage mount information, volume preparation, volume mounting and recycling, and limiting the number of bytes per second and the number of I / O operations per second for container reading and writing; The distributed database is used to store the status information of the cluster, including the total storage of each node, container and binding information, and the affinity and anti-affinity requirements of containers, providing a decision-making basis for allocation and scheduling; Disk information collection and reporting process: The disk device manager running on each node collects disk information and reports it to the container cloud engine, and the engine stores the reported node disk information in the distributed database; Container scheduling and storage allocation process: After the container creation command is issued, the container is in the pending scheduling container queue. The containers in the queue are sorted according to priority. The scheduler takes out the pending scheduling containers from the queue for scheduling, and initially filters the nodes based on the CPU, memory, ports, and labels of each node. The filtered node list is passed to the storage scheduling plugin; The storage scheduling plugin first checks whether the current container has been scheduled. If there is binding information that matches the container being scheduled, it uses this binding information to complete the scheduling. This scheduling behavior maintains the container storage state and ensures that the storage is not lost when the container is restarted multiple times; If there is no matching binding information, it dynamically obtains the storage information of each node in the cluster, the storage information of the running containers, and the affinity and anti-affinity requirements from the distributed database, and combines the disk requests, affinity and anti-affinity requirements of the container being scheduled. After comprehensive calculation, it obtains the disk allocation method, usage, remaining amount of each node, and the topological relationship between containers with storage requests; Container destruction and storage recycling process: Filter and sort the nodes by analyzing the affinity and anti-affinity between the container being scheduled and the scheduled and undeleted containers; Among them, for containers with anti-affinity requirements, the storage scheduling plugin filters out the nodes where the containers with anti-affinity conflicts are running; For containers with affinity requirements, the scheduler scores and sorts the nodes according to the affinity scoring strategy.
2. The intelligent scheduling and allocation method for container local storage applied to a cloud platform according to claim 1, characterized in that: In the container destruction and storage recycling process, an optimal matching algorithm is used for storage allocation, which is as follows: The input of the optimal matching algorithm is the container local storage request, the node list, and the disk information on each node; The output is the name of the node to which the container is to be bound and the disk list.
3. The intelligent scheduling and allocation method for container local storage applied to a cloud platform according to claim 2, wherein: The storage requests in the input of the optimal matching algorithm are divided into two categories: bare disks and quotas according to whether they are allocated for bare disks. The disk screening conditions for both include disk type, disk quantity, capacity, and bandwidth; Bare disk allocation is further divided into bare disk sharing and bare disk exclusivity. Bare disk sharing means that multiple containers share the capacity of a certain disk, and the storage between these containers is isolated from each other; The upper limit of the number of shared containers that can be configured for each disk; The exclusive use of raw disks is exclusive. Disks that have been allocated to a certain container in an exclusive form cannot be bound to other containers for use. The filtering conditions attached to the container with a raw disk request include the number of disks and the minimum disk size; the container with a quota request needs to attach the requested quota size; If the local storage request for the container being scheduled is exclusive use of raw disks, the disk information of the nodes is retrieved from the ordered list of nodes in sequence, and it is judged whether the disk type, the number of disks, and the minimum disk size request of the container are met. If they are met, the node and the disk number are returned; If the request is for shared raw disks, the list of nodes is re-sorted in descending order according to the number of shared but not reaching the sharing limit disks that meet the container requirements, and the disks on each node are sorted according to the sharing times; Check whether the nodes and disks meet the container request in sequence. If they meet, return the node name and the disk list; If the request is for quota allocation, the disk list after excluding whole disk allocation for each node is sorted in ascending order according to the remaining amount; First, check the nodes in sequence and analyze whether there are disks on the nodes that have been allocated according to the quota and meet the request of the container being scheduled. If there are, return the node name and the disk number; if the nodes and disks that meet the requirements still cannot be obtained after checking all the nodes, then when checking the nodes for the second time, obtain the disks that meet the request and have not been used for raw disk allocation; The logic of the allocation process is as follows: For shared raw disks, preferentially select disks that have been shared; For quota allocation, preferentially select disks that have been allocated quotas and have the smallest remaining amount; Among them, the output result of the optimal matching algorithm is the node name and the disk list. This result combines the basic information of the container being scheduled and the disk request to form binding information and submit it to the distributed database; The binding information will only be cleared when the container is completely deleted; When the container restarts, the scheduling and storage allocation of the container are completed according to the binding information recorded in the distributed database to maintain the storage state of the container.
4. The intelligent scheduling and allocation method for container local storage applied to a cloud platform according to claim 1, characterized in that: The disk device manager running on each node, as the implementer of storage management, undertakes the responsibilities of disk collection and reporting, allocation and recycling, mounting and unmounting, and container IO bandwidth limitation; specifically as follows: When the node is incorporated into the container cloud platform, the disk device manager collects the type, capacity, and remaining amount of the node disks and reports them to the container cloud engine, and the latter records them in the distributed database; When the container with a local storage request completes scheduling and writes the binding information into the distributed database, this binding information will be pushed to the disk device manager; The disk device manager first obtains the corresponding disk according to the disk number in the binding information, prepares to mount the volume according to the local storage usage method of the container. If it is a raw disk, create a mount directory. If it is a quota, use the quota technology to isolate the corresponding size of storage; Mount the storage volume to the container; Complete the IO bandwidth configuration based on the Cgroup blkio subsystem; Recovery is the reverse process of allocation. When the disk device manager detects a storage unbinding event, it recovers the storage and unmounts it. When the recovery is completed, the disk device manager notifies the container cloud engine to return the resources, and the container cloud engine deletes the binding information of the container from the distributed database.
5. The intelligent scheduling and allocation method for container local storage applied to a cloud platform according to claim 2, wherein: The implementation of the storage scheduling plugin specifically includes the following steps: Step 1, receive and process the scheduling request from the container cloud engine scheduler: The requests received by the storage scheduling plugin include the list of available nodes and container information. The storage scheduling plugin first obtains the current storage topology information of the cluster, node storage information, and the binding information of the running containers from the distributed database. Then, it queries whether the current container has a scheduling record. If found, it returns the binding information. If not found, it completes the disk allocation according to the optimal matching algorithm and writes the binding information into the distributed database. Finally, it returns the scheduling result to the scheduler of the container engine. Step 2, implement the scheduling process: Perform topology calculation between the container being scheduled and the scheduled but undeleted containers to obtain a list of nodes sorted by priority, and then perform scheduling using the optimal matching algorithm. Step 3, read and write to the distributed database: After the scheduling is completed, write the scheduling information into the distributed database for storage reservation. At this time, the storage is not fully allocated. The reservation can prevent resource competition. If the subsequent allocation fails, resource return is required. If the container enters the deletion process, the database content needs to be cleared to complete the resource return.
6. The intelligent scheduling and allocation method for container local storage applied to a cloud platform according to claim 1, characterized in that: The implementation of the disk device manager StorageAgent specifically includes the following steps: Step 1, report storage information: When the disk device manager StorageAgent starts, it needs to push the disk information of the node to the container cloud engine. This information is stored in the distributed database and used by the storage scheduling plugin during container scheduling. Step 2, listen for container and storage binding events: When a storage binding event is detected, it is necessary to implement the creation of the mounted volume, the mounting of the container directory and the host directory, and the limitation of the container IO bandwidth, and report the results of the allocation and mounting to the container cloud engine. Step 3, listen for container and storage unbinding events: When an unbinding event is detected, it is necessary to implement the recovery of the mounted volume, the unmounting of the container directory and the host directory, and report the results of the recovery and unmounting.
Citation Information
Patent Citations
Resource scheduling method and device based on kubernetes, equipment and storage medium
CN112199194A
Virtual resource object component
WO2014036717A1