Container scheduling method and apparatus in cloud environment, electronic device, and storage medium
By introducing sub-cluster managers and resource cluster objects in cloud data centers, the scheduling status of containers can be monitored in real time and cross-cluster migration can be performed. This solves the resource bottleneck problem caused by the decentralized management of Kubernetes sub-clusters, improves the availability and resource utilization of containers, and ensures the quality of cloud services.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-28
AI Technical Summary
In cloud data centers, the decentralized management and resource bottlenecks of Kubernetes sub-clusters lead to inefficient container scheduling and resource allocation, affecting the quality and availability of cloud services.
By introducing a sub-cluster manager and a resource cluster object, the scheduling status of containers is monitored in real time, and information about containers that have not been successfully allocated system resources is stored in the resource cluster object. The main cluster manager enables cross-cluster resource awareness and balancing, and containers can be quickly migrated to nodes with sufficient resources.
It improves container availability and resource utilization, reduces container downtime, and ensures the quality and efficiency of cloud services.
Smart Images

Figure CN2025136457_28052026_PF_FP_ABST
Abstract
Description
A container scheduling method, apparatus, electronic device, and storage medium in a cloud environment
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411683856.3, filed on November 22, 2024, entitled "A container scheduling method, apparatus, electronic device and storage medium in a cloud environment", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of cloud computing technology, and in particular to a container scheduling method, apparatus, electronic device and storage medium in a cloud environment. Background Technology
[0004] With the rapid development of cloud computing technology, cloud service providers' cloud data centers are becoming increasingly large in scale. For example, the number of Kubernetes sub-clusters in cloud data centers is increasing, and the scale of nodes and containers on a single Kubernetes sub-cluster is constantly growing.
[0005] This leads to increasingly prominent performance bottlenecks in the management of Kubernetes sub-clusters. At the same time, due to factors such as disaster recovery in different locations, cost budgets, hybrid cloud, multi-cloud, and traffic access proximity, cloud service providers choose to deploy various Kubernetes sub-clusters in cloud data centers in different geographical locations and in data centers of different cloud vendors. This results in the dispersion of Kubernetes sub-clusters, the continuous increase in the number of Kubernetes sub-clusters, and the increasing difficulty in management. Summary of the Invention
[0006] This application discloses a container scheduling method, apparatus, electronic device, and storage medium in a cloud environment.
[0007] Firstly, this application discloses a container scheduling method in a cloud environment, applied to a sub-cluster manager in a sub-cluster within a cloud data center. The method includes:
[0008] The scheduling status of containers on each node in the sub-cluster is monitored in real time. The scheduling status of containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container.
[0009] If the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container, the rescheduling field of the resource cluster object in the sub-cluster should at least store the container information of the target container and the node information of the first node, which is located in the sub-cluster.
[0010] Secondly, this application discloses a container scheduling method in a cloud environment, applied to the master cluster manager in a master cluster of a cloud data center, the method including:
[0011] Real-time monitoring of whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed;
[0012] When the information in the rescheduling field of the resource cluster object in the first sub-cluster of the cloud data center changes, the container information of the target container and the node information of the first node where the target container is located are obtained from the rescheduling field. The container information of the target container and the node information of the first node are stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container. The first node is located in the first sub-cluster.
[0013] Based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster in the cloud data center.
[0014] Thirdly, this application discloses a container scheduling device in a cloud environment, applied to a sub-cluster manager in a sub-cluster of a cloud data center, the device comprising:
[0015] The first monitoring module is used to monitor the scheduling status of containers on each node in the sub-cluster in real time. The scheduling status of the containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container.
[0016] The first storage module is used to store, in the rescheduling field of the resource cluster object in the sub-cluster, at least the container information of the target container and the node information of the first node, where the first node is located in the sub-cluster, when the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container.
[0017] Fourthly, this application discloses a container scheduling device in a cloud environment, applied to the master cluster manager in a master cluster of a cloud data center. The device includes:
[0018] The second monitoring module is used to monitor in real time whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed.
[0019] The third acquisition module is used to acquire the container information of the target container and the node information of the first node where the target container is located in the rescheduling field of the resource cluster object in the first sub-cluster in the cloud data center when the information in the rescheduling field of the resource cluster object in the first sub-cluster changes. The container information of the target container and the node information of the first node are stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container. The first node is located in the first sub-cluster.
[0020] The migration module is used to migrate the target container to a node in the second sub-cluster in the cloud data center based on the target container's container information and the node information of the first node.
[0021] Fifthly, this application discloses an electronic device comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to perform the method as described in any of the foregoing aspects.
[0022] Sixthly, this application discloses a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods described in any of the above aspects.
[0023] In a seventh aspect, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the methods as described in any of the above aspects.
[0024] The technical solution provided in this application may include the following beneficial effects:
[0025] In this application, the sub-cluster manager in the sub-cluster of the cloud data center monitors the scheduling status of containers on each node in the sub-cluster in real time. The scheduling status of the containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container. If the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container, the rescheduling field in the resource cluster object in the sub-cluster stores at least the container information of the target container and the node information of the first node, where the first node is located in the sub-cluster.
[0026] This allows the master cluster manager in the main cluster of the cloud data center to monitor in real time whether the information in the rescheduled field of the resource cluster object in each sub-cluster of the cloud data center has changed. If the information in the rescheduled field of the resource cluster object in the first sub-cluster of the cloud data center changes, it obtains the container information of the target container and the node information of the first node where the target container is located from the rescheduled field. The container information of the target container and the node information of the first node are stored in the rescheduled field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated according to the actual needs of the target container. The first node is located in the first sub-cluster. Based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster of the cloud data center.
[0027] This application introduces a sub-cluster manager and a resource cluster object in the sub-cluster, and a master cluster manager in the master cluster. The sub-cluster manager, resource cluster object, and master cluster manager work together to achieve real-time resource awareness and resource balancing capabilities for each sub-cluster in a federated scenario within a cloud data center. This improves the real-time performance of detecting whether containers are normal / available, ensures that containers can be immediately scheduled to other nodes or clusters when they cannot obtain the system resources they actually need, reduces container downtime, avoids affecting the quality of cloud services provided by containers, and improves container availability and overall resource utilization of each sub-cluster in a federated scenario within a cloud data center. Attached Figure Description
[0028] Figure 1 is a structural block diagram of a cloud data center according to this application.
[0029] Figure 2 is a flowchart of the steps of a container scheduling method in a cloud environment according to this application.
[0030] Figure 3 is a flowchart of the steps of a container scheduling method in a cloud environment according to this application.
[0031] Figure 4 is a structural block diagram of a container scheduling device in a cloud environment according to this application.
[0032] Figure 5 is a structural block diagram of a container scheduling device in a cloud environment according to this application.
[0033] Figure 6 is a block diagram of an electronic device according to this application.
[0034] Figure 7 is a block diagram of an electronic device according to this application. Detailed Implementation
[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] To unify the management of these Kubernetes sub-clusters in cloud data centers, the Kubernetes community introduced cluster federation technology, which aims to unify the management of Kubernetes sub-clusters distributed in various locations through the management plane of cluster federation.
[0037] In Kubernetes sub-cluster federation technology, cross-cluster scheduling of containers is provided. When a container needs to be created, the scheduler in the cloud data center can identify the system resource availability of each Kubernetes sub-cluster in the cloud data center and schedule the container to a Kubernetes sub-cluster with more system resource availability. This allows the container to be created on a Kubernetes sub-cluster with more system resource availability, and in turn, the Kubernetes sub-cluster with more system resource availability allocates the system resources needed by the container for the container to use.
[0038] However, the system resources in the Kubernetes sub-cluster may change subsequently. For example, nodes in the Kubernetes sub-cluster may crash, the network of nodes in the Kubernetes sub-cluster may fail, the storage function of nodes in the Kubernetes sub-cluster may fail, or the load on nodes in the Kubernetes sub-cluster may be too high.
[0039] If the system resources in a Kubernetes sub-cluster undergo the aforementioned changes, it may result in the containers hosted on the Kubernetes sub-cluster not having enough system resources available, leading to abnormal container status. Therefore, it is necessary to reschedule the containers to other Kubernetes sub-clusters in the cloud data center with sufficient system resources.
[0040] Therefore, it is necessary to check the status of the containers hosted on each Kubernetes sub-cluster.
[0041] The current Kubernetes cluster federation technology checks the status of containers hosted on each Kubernetes sub-cluster every 5 minutes, and triggers cross-cluster scheduling of containers when an anomaly is detected.
[0042] However, excessively long check intervals can make it difficult to schedule abnormal containers across clusters in a timely manner. The unavailability of containers can last up to 5 minutes. Excessive unavailability of containers leads to poor availability and affects the quality of cloud services provided by containers.
[0043] Therefore, this application is proposed. Referring to FIG1, a structural block diagram of a cloud data center according to this application is shown.
[0044] A cloud data center includes a main cluster and multiple sub-clusters. Figure 1 illustrates this with an example of one main cluster and two sub-clusters, but this is not intended to limit the scope of protection of this application.
[0045] The two sub-clusters are sub-cluster 1 and sub-cluster 2.
[0046] Each sub-cluster has its own sub-cluster manager.
[0047] Sub-clusters include Kubernetes sub-clusters, etc.
[0048] For example, sub-cluster managers can be deployed separately in each sub-cluster using Kubernetes Deployment workloads.
[0049] Sub-cluster managers can include components such as resource-cluster-manager. A sub-cluster can have one sub-cluster manager, and it is not necessary to deploy a sub-cluster manager on every node in a sub-cluster. Sub-cluster managers in a sub-cluster can communicate and interact with each node in the sub-cluster, allowing them to monitor the status of containers on each node.
[0050] Each sub-cluster also has a resource cluster object, which can include Kubernetes custom resources such as ResourceCluster, to represent abnormal conditions of the sub-cluster.
[0051] Each sub-cluster also has its own nodes, which can be physical nodes or virtual nodes. Physical nodes can be physical machines, such as servers, while virtual nodes can be virtual machines set up on physical machines. The nodes can also communicate with each other.
[0052] Each sub-cluster also has a cluster federation agent service, which is used to interface with the cluster federation management plane service in the main cluster.
[0053] The master cluster manager can be deployed in the master cluster using Kubernetes Deployment workloads.
[0054] The master cluster manager can include components such as resource-hub-manager. A master cluster can have only one master cluster manager, or it doesn't need to be deployed on every single node in the master cluster. The master cluster manager in the master cluster can communicate and interact with each node in the master cluster.
[0055] The primary cluster includes the Kubernetes primary cluster.
[0056] The master cluster manager can monitor each resource cluster object in each sub-cluster manager, for example, to monitor whether the content of each resource cluster object in each sub-cluster manager has changed.
[0057] The primary cluster also includes a cluster federation management plane service, which can interact with the primary cluster manager.
[0058] Referring to Figure 2, a flowchart of the steps of a container scheduling method in a cloud environment according to this application is shown, which is applied to a sub-cluster manager in a sub-cluster in a cloud data center.
[0059] The method includes:
[0060] In step S101, the scheduling status of containers on each node in the sub-cluster is monitored in real time. The scheduling status of containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container.
[0061] System resources include ports, IP (Internet Protocol) addresses, CPU (Central Processing Unit), memory, bandwidth, and disks, etc.
[0062] The sub-cluster manager includes a Kubernetes Client, which has a built-in watch function. The sub-cluster manager can monitor the scheduling status of containers on each node in the sub-cluster in real time via HTTP Watch, thereby enabling real-time monitoring of whether the system resources on any node have been successfully allocated to any container on any node in the sub-cluster according to its actual needs.
[0063] For any node in the sub-cluster, if after multiple attempts it fails to allocate system resources to a container on that node according to the container's actual needs, the node can store the container information in a specific field on that node. The container information may include the container's metadata, which may include the container's name and the namespace where the container's name is located. In this application, different containers have different names.
[0064] Thus, for any node in the sub-cluster, the sub-cluster manager can access specific fields in that node in real time, extract container information from those fields, and determine whether the system resources on the node where the container resides have been successfully allocated according to the actual needs of that container.
[0065] In step S102, if the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container, the rescheduling field of the resource cluster object in the sub-cluster stores at least the container information of the target container and the node information of the first node, where the first node is located in the sub-cluster.
[0066] Node information can include the node's IP address, and different nodes have different node information.
[0067] In this application, the sub-cluster manager in the sub-cluster of the cloud data center monitors the scheduling status of containers on each node in the sub-cluster in real time. The scheduling status of the containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container. If the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container, the rescheduling field in the resource cluster object in the sub-cluster stores at least the container information of the target container and the node information of the first node, where the first node is located in the sub-cluster.
[0068] This allows the master cluster manager in the main cluster of the cloud data center to monitor in real time whether the information in the rescheduled field of the resource cluster object in each sub-cluster of the cloud data center has changed. If the information in the rescheduled field of the resource cluster object in the first sub-cluster of the cloud data center changes, it obtains the container information of the target container and the node information of the first node where the target container is located from the rescheduled field. The container information of the target container and the node information of the first node are stored in the rescheduled field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated according to the actual needs of the target container. The first node is located in the first sub-cluster. Based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster of the cloud data center.
[0069] This application introduces a sub-cluster manager and a resource cluster object in the sub-cluster, and a master cluster manager in the master cluster. The sub-cluster manager, resource cluster object, and master cluster manager work together to achieve real-time resource awareness and resource balancing capabilities for each sub-cluster in a federated scenario within a cloud data center. This improves the real-time performance of detecting whether containers are normal / available, ensures that containers can be immediately scheduled to other nodes or clusters when they cannot obtain the system resources they actually need, reduces container downtime, avoids affecting the quality of cloud services provided by containers, and improves container availability and overall resource utilization of each sub-cluster in a federated scenario within a cloud data center.
[0070] In this application, according to statistics, there are many reasons why the system resources on the first node where the target container is located were not successfully allocated to the target container according to the actual needs of the target container. These reasons include: the IP network of the first node where the target container is located is not connected; the port of the first node where the target container is located is not connected; the idle CPU resources of the first node where the target container is located cannot meet the actual needs of the target container; the idle memory resources of the first node where the target container is located cannot meet the actual needs of the target container; the idle disk resources of the first node where the target container is located cannot meet the actual needs of the target container; and the first node where the target container is located crashes.
[0071] In order to enable the master cluster manager in the main cluster of the cloud data center to perceive the reason why the system resources on the first node where the target container is located were not successfully allocated to the target container according to the actual needs of the target container, in some other embodiments of this application, the rescheduling field in the resource cluster object in the sub-cluster includes at least an information field.
[0072] Thus, in step S102, when storing at least the container information of the target container and the node information of the first node in the rescheduling field of the resource cluster object in the sub-cluster, the container information of the target container and the node information of the first node can be stored in the information field of the rescheduling field.
[0073] Secondly, the rescheduling field in the resource cluster object in the sub-cluster also includes a reason field.
[0074] Furthermore, it is also possible to obtain the reason why the system resources on the first node were not successfully allocated to the target container according to the actual needs of the target container, and store the reason in the reason field of the rescheduling field.
[0075] In some embodiments of this application, if the sub-cluster manager detects that the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container, it can detect whether the IP network of the first node where the target container is located is connected. If the IP network of the first node where the target container is located is not connected, it can be determined as one of the reasons why the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container.
[0076] In some embodiments of this application, if the sub-cluster manager detects that system resources on the first node have not been successfully allocated to the target container according to its actual needs, it can detect whether the port of the first node where the target container resides is connected. If the port of the first node where the target container resides is not connected, this can be determined as one of the reasons why system resources on the first node have not been successfully allocated to the target container according to its actual needs. The port of the first node can be all ports in the first node, or it can be a portion of the ports in the first node, such as the Kubernetes kubelet port 10250.
[0077] In some embodiments of this application, when the sub-cluster manager detects that the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container, it can detect whether the idle CPU resources of the first node where the target container is located meet the actual CPU needs of the target container. If the idle CPU resources of the first node where the target container is located cannot meet the actual CPU needs of the target container, it can be determined as one of the reasons why the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container.
[0078] In some embodiments of this application, when the sub-cluster manager detects that the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container, it can detect whether the free memory resources of the first node where the target container is located meet the actual memory needs of the target container. If the free memory resources of the first node where the target container is located cannot meet the actual memory needs of the target container, it can be determined as one of the reasons why the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container.
[0079] In some embodiments of this application, when the sub-cluster manager detects that the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container, it can detect whether the free disk resources of the first node where the target container is located meet the actual disk needs of the target container. If the free disk resources of the first node where the target container is located cannot meet the actual disk needs of the target container, it can be determined as one of the reasons why the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container.
[0080] In some embodiments of this application, if the sub-cluster manager detects that the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container, it can detect whether the first node where the target container is located has crashed. If the first node where the target container is located crashes, the crash of the first node where the target container is located can be determined as one of the reasons why the system resources on the first node have not been successfully allocated to the target container according to the actual needs of the target container.
[0081] In order to enable the main cluster manager in the main cluster of the cloud data center to perceive the duration for which the system resources on the first node where the target container is located were not successfully allocated according to the actual needs of the target container, in some other embodiments of this application, the rescheduling field in the resource cluster object in the sub-cluster includes at least an information field.
[0082] Thus, in step S102, when storing at least the container information of the target container and the node information of the first node in the rescheduling field of the resource cluster object in the sub-cluster, the container information of the target container and the node information of the first node can be stored in the information field of the rescheduling field.
[0083] Secondly, the rescheduling field in the resource cluster object in the sub-cluster also includes a duration field.
[0084] Furthermore, it is also possible to obtain the duration during which system resources on the first node were not successfully allocated to the target container according to its actual needs.
[0085] If the duration exceeds the preset tolerance duration of the sub-cluster manager, the duration is stored in the duration field of the rescheduling field in the resource cluster object in the sub-cluster.
[0086] The preset tolerance time of the sub-cluster manager can be set in advance by technicians according to the actual situation. The specific value of the preset tolerance time of the sub-cluster manager can be determined according to the actual situation, and this application does not limit it.
[0087] The preset tolerance time for the sub-cluster manager can include, for example, 4 seconds, 5 seconds, or 6 seconds.
[0088] Specifically, when retrieving the duration during which system resources on the first node were not successfully allocated to the target container according to its actual needs, the current time of the sub-cluster can be obtained. In the scenario of allocating system resources on the first node to the target container, the creation time of the target container's metadata information in the first node's memory can be retrieved.
[0089] In the process of allocating system resources on the first node to the target container, the first node first needs to establish the metadata information of the target container in the memory of the first node, and then execute other steps related to "allocating system resources on the first node to the target container". Secondly, the first node can also store the establishment time when the metadata information of the target container is established in memory.
[0090] For example, the binding relationship between storing the target container's metadata information in memory and the creation time when the target container's metadata information is created.
[0091] Subsequently, after the first node reclaims the system resources allocated to the target container on the first node, the metadata information of the target container and the binding relationship between the creation time of the metadata information of the target container and the creation time are deleted from the memory of the first node.
[0092] In this way, the sub-cluster manager in the sub-cluster can obtain the creation time of the target container's metadata information in the memory of the first node in the scenario of allocating system resources on the first node for the target container, based on the target container's metadata information and the binding relationship.
[0093] Then, the difference between the current time and the establishment time can be calculated, and the duration can be obtained based on this difference. For example, the difference can be used as the duration.
[0094] In other embodiments of this application, the sub-cluster manager in the sub-cluster may also store the creation time of the metadata information of the target container in the memory of the first node in the duration field of the rescheduling field in the resource cluster object in the sub-cluster.
[0095] Referring to Figure 3, a flowchart illustrating the steps of a container scheduling method in a cloud environment according to this application is shown, applied to the master cluster manager in a master cluster of a cloud data center. The method includes:
[0096] In step S201, the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center is monitored in real time to see if there are any changes.
[0097] The master cluster manager includes a Kubernetes Client, which has a built-in watch function. The master cluster manager can use HTTP Watch to monitor in real time whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed.
[0098] If the information in the rescheduling field of the resource cluster object in the first sub-cluster in the cloud data center changes, in step S202, the container information of the target container and the node information of the first node where the target container is located are obtained from the rescheduling field.
[0099] The container information of the target container and the node information of the first node are: stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container. The first node is located in the first sub-cluster.
[0100] Alternatively, if the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has not changed, return to the execution step S201 in real time: monitor whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed.
[0101] In step S203, based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster in the cloud data center.
[0102] In this application, the master cluster manager in the master cluster of the cloud data center monitors in real time whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed; if the information in the rescheduling field of the resource cluster object in the first sub-cluster of the cloud data center changes, the container information of the target container and the node information of the first node where the target container is located are obtained from the rescheduling field; the container information of the target container and the node information of the first node are: stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated according to the actual needs of the target container; the first node is located in the first sub-cluster; based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster of the cloud data center.
[0103] This application introduces a sub-cluster manager and a resource cluster object in the sub-cluster, and a master cluster manager in the master cluster. The sub-cluster manager, resource cluster object, and master cluster manager work together to achieve real-time resource awareness and resource balancing capabilities for each sub-cluster in a federated scenario within a cloud data center. This improves the real-time performance of detecting whether containers are normal / available, ensures that containers can be immediately scheduled to other nodes or clusters when they cannot obtain the system resources they actually need, reduces container downtime, avoids affecting the quality of cloud services provided by containers, and improves container availability and overall resource utilization of each sub-cluster in a federated scenario within a cloud data center.
[0104] In this application, the main cluster has a cluster federation management plane service, which includes the Kubernetes cluster federation management plane service, etc.
[0105] In the cluster federation management plane service of the main cluster, when allocating the target container to a node in a sub-cluster in the cloud data center in advance, it can first select a sub-cluster and a node in that sub-cluster based on the status of each node in each sub-cluster in the cloud data center. For example, it can select the first sub-cluster and the first node in the first sub-cluster. Then, in the cluster federation management plane service of the main cluster, it stores the first binding relationship between the container information of the target container, the node information of the first node, and the sub-cluster information of the first sub-cluster. Then, according to the first binding relationship, it allocates the target container to the first node in the first sub-cluster, so that the first node can allocate system resources on the first node where the target container is located according to the actual needs of the target container.
[0106] Thus, in step S203, the first binding relationship between the target container's container information, the first node's node information, and the first sub-cluster's sub-cluster information can be deleted in the cluster federation management plane service in the main cluster. A cross-cluster scheduling request is then sent to the cluster federation management plane service. Based on the cross-cluster scheduling request and the status of each node in the sub-clusters other than the first sub-cluster in the cloud data center, the cluster federation management plane service selects a sub-cluster and a node within that sub-cluster. For example, it selects a second sub-cluster and a second node within that sub-cluster. Then, in the cluster federation management plane service in the main cluster, the second binding relationship between the target container's container information, the second node's node information, and the second sub-cluster's sub-cluster information is stored. Based on this second binding relationship, the target container is migrated to the second node in the second sub-cluster in the cloud data center, allowing the second node to allocate system resources on the second node according to the target container's actual needs.
[0107] The system resources of the second node in the selected second sub-cluster are sufficient to meet the actual needs of the target container, such as the actual needs for starting and running the target container, so that the target container can provide cloud services to the outside world normally.
[0108] This application does not limit the specific selection method for "selecting the second node in the second sub-cluster among the sub-clusters other than the first sub-cluster in the cloud data center", and existing selection methods can also be used.
[0109] In other embodiments of this application, the rescheduling field in the resource cluster object in the first sub-cluster includes at least an information field. The container information of the target container and the node information of the first node are located in the information field of the rescheduling field in the resource cluster object in the first sub-cluster.
[0110] The rescheduling field in the resource cluster object of the first sub-cluster also includes a reason field.
[0111] The reason field is used to store the reason why the system resources on the node where the container is located were not successfully allocated to the container according to the actual needs of the container.
[0112] Thus, before performing step S203, the method further includes:
[0113] Determine if the reason field contains content.
[0114] If the reason field contains content, then proceed to step S203: based on the container information of the target container and the node information of the first node, migrate the target container to a node in the second sub-cluster in the cloud data center.
[0115] Since the reason field is used to store the reason why the system resources on the node where the container is located were not successfully allocated to the container according to the actual needs of the container, if the reason field has content, it means that the reason field stores the reason why the system resources on the first node where the target container is located were not successfully allocated to the target container according to the actual needs of the target container, and thus explains why the system resources on the first node where the target container is located were not successfully allocated to the target container according to the actual needs of the target container.
[0116] Therefore, in order to successfully allocate system resources on a node to the target container according to its actual needs, so that the target container can start and run, and thus enable the target container to provide cloud services to the outside world normally, step S203 can be executed directly: based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster in the cloud data center.
[0117] In other embodiments of this application, the rescheduling field in the resource cluster object in the first sub-cluster includes at least an information field. The container information of the target container and the node information of the first node are located in the information field of the rescheduling field in the resource cluster object in the first sub-cluster.
[0118] The rescheduling field in the resource cluster object of the first sub-cluster also includes a duration field.
[0119] The duration field is used to store the duration during which system resources on the node where the container resides were not successfully allocated to the container according to the container's actual needs.
[0120] Thus, before performing step S203, the method further includes:
[0121] Determine if the duration field contains content.
[0122] If the duration field contains content, determine whether the timeout check policy of the master cluster manager has been enabled.
[0123] Since the duration field is used to store the duration during which system resources on the node where the container resides were not successfully allocated according to the actual needs of the container, if the duration field contains content, the explanation field stores the duration during which system resources on the first node where the target container resides were not successfully allocated according to the actual needs of the target container, thus indicating that system resources on the first node where the target container resides were not successfully allocated according to the actual needs of the target container for a period of time.
[0124] Therefore, in order to successfully allocate system resources on a node to the target container according to its actual needs, so that the target container can start and run, and thus enable the target container to provide cloud services to the outside world normally, it is necessary to determine whether the timeout check policy of the main cluster manager has been enabled.
[0125] Whether the timeout check policy of the main cluster manager is enabled or not can be set by the technicians in advance according to the situation. The timeout check policy of the main cluster manager can be enabled or disabled. If the timeout check policy of the main cluster manager is enabled, the technicians also need to set the preset tolerance time of the main cluster manager.
[0126] The preset tolerance time of the main cluster manager can be set in advance by technicians according to the actual situation. The specific value of the preset tolerance time of the main cluster manager can be determined according to the actual situation, and this application does not limit it.
[0127] The default tolerance time for the master cluster manager can include, for example, 9 seconds, 10 seconds, or 11 seconds.
[0128] The default tolerance duration of the master cluster manager and the default tolerance duration of each sub-cluster manager may be the same or different in value.
[0129] If the timeout check policy of the main cluster manager is not enabled, step S203 can be executed directly: based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster in the cloud data center.
[0130] Alternatively, if the timeout check policy of the master cluster manager is enabled, the duration during which system resources on the first node were not successfully allocated to the target container according to the actual needs of the target container can be obtained.
[0131] If the duration exceeds the preset tolerance time of the main cluster manager, proceed to step S203: based on the container information of the target container and the node information of the first node, migrate the target container to a node in the second sub-cluster in the cloud data center.
[0132] Alternatively, if the duration does not exceed the preset tolerance time of the main cluster manager, step S203 can be skipped.
[0133] Specifically, when retrieving the duration during which system resources on the first node were not successfully allocated to the target container according to its actual needs, the current time of the main cluster can be obtained. In the scenario where system resources are allocated to the target container, the creation time of the target container's metadata information in the memory of the first node in the first sub-cluster can also be obtained.
[0134] In some embodiments of this application, the sub-cluster manager in the first sub-cluster stores the creation time of the target container's metadata information in the memory of the first node in the duration field of the rescheduling field of the resource cluster object in the first sub-cluster. Thus, the master cluster manager can extract the creation time of the target container's metadata information in the memory of the first node from the duration field of the rescheduling field of the resource cluster object in the first sub-cluster.
[0135] Then, the difference between the current time and the establishment time can be calculated, and the duration can be obtained based on this difference. For example, the difference can be used as the duration.
[0136] In one embodiment, the preset tolerance time of the master cluster manager in the master cluster has higher priority than the preset tolerance time of the sub-cluster managers in each sub-cluster. For example, if the preset tolerance time of the master cluster manager in the master cluster is 10 seconds, while the preset tolerance time of each sub-cluster manager in each sub-cluster is 5 seconds, then whether the master cluster manager executes step S203 needs to be determined based on the preset tolerance time of 10 seconds. This ensures that when there are too many sub-clusters and the preset tolerance times of each sub-cluster are inconsistent, the master cluster manager in the master cluster can decide whether to perform cross-cluster scheduling of containers based on a unified preset tolerance time.
[0137] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that some of the actions involved in the embodiments described in the specification are not necessarily essential to this application.
[0138] Referring to Figure 4, a container scheduling device in a cloud environment according to this application is shown, which is applied to a sub-cluster manager in a sub-cluster in a cloud data center. The device includes:
[0139] The first listening module 11 is used to monitor the scheduling status of containers on each node in the sub-cluster in real time. The scheduling status of the containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container.
[0140] The first storage module 12 is used to store at least the container information of the target container and the node information of the first node, which is located in the sub-cluster, in the rescheduling field of the resource cluster object in the sub-cluster when the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container.
[0141] The rescheduling field includes at least an information field;
[0142] The first storage module includes:
[0143] The first storage unit is used to store the container information of the target container and the node information of the first node in the information field of the rescheduling field.
[0144] Secondly, the rescheduling field also includes a reason field;
[0145] Accordingly, the device also includes:
[0146] The first acquisition module is used to obtain the reason why the system resources on the first node were not successfully allocated to the target container according to the actual needs of the target container.
[0147] The second storage module is used to store the reason in the reason field of the rescheduling field.
[0148] The rescheduling field includes at least an information field;
[0149] The first storage module includes:
[0150] The first storage unit is used to store the container information of the target container and the node information of the first node in the information field of the rescheduling field.
[0151] Secondly, the rescheduling field also includes a duration field;
[0152] Accordingly, the device also includes:
[0153] The second acquisition module is used to acquire the duration during which system resources on the first node were not successfully allocated to the target container according to the actual needs of the target container.
[0154] The third storage module is used to store the duration in the duration field of the rescheduling field when the duration exceeds the preset tolerance duration of the sub-cluster manager.
[0155] In this application, the sub-cluster manager in the sub-cluster of the cloud data center monitors the scheduling status of containers on each node in the sub-cluster in real time. The scheduling status of the containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container. If the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container, the rescheduling field in the resource cluster object in the sub-cluster stores at least the container information of the target container and the node information of the first node, where the first node is located in the sub-cluster.
[0156] This allows the master cluster manager in the main cluster of the cloud data center to monitor in real time whether the information in the rescheduled field of the resource cluster object in each sub-cluster of the cloud data center has changed. If the information in the rescheduled field of the resource cluster object in the first sub-cluster of the cloud data center changes, it obtains the container information of the target container and the node information of the first node where the target container is located from the rescheduled field. The container information of the target container and the node information of the first node are stored in the rescheduled field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated according to the actual needs of the target container. The first node is located in the first sub-cluster. Based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster of the cloud data center.
[0157] This application introduces a sub-cluster manager and a resource cluster object in the sub-cluster, and a master cluster manager in the master cluster. The sub-cluster manager, resource cluster object, and master cluster manager work together to achieve real-time resource awareness and resource balancing capabilities for each sub-cluster in a federated scenario within a cloud data center. This improves the real-time performance of detecting whether containers are normal / available, ensures that containers can be immediately scheduled to other nodes or clusters when they cannot obtain the system resources they actually need, reduces container downtime, avoids affecting the quality of cloud services provided by containers, and improves container availability and overall resource utilization of each sub-cluster in a federated scenario within a cloud data center.
[0158] Referring to Figure 5, a container scheduling device in a cloud environment according to this application is shown, which is applied to the master cluster manager in the master cluster of a cloud data center. The device includes:
[0159] The second monitoring module 21 is used to monitor in real time whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed.
[0160] The third acquisition module 22 is used to acquire the container information of the target container and the node information of the first node where the target container is located in the rescheduling field of the resource cluster object in the first sub-cluster in the cloud data center when the information in the rescheduling field of the resource cluster object in the first sub-cluster changes; the container information of the target container and the node information of the first node are: stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container; the first node is located in the first sub-cluster.
[0161] Migration module 23 is used to migrate the target container to a node in the second sub-cluster in the cloud data center based on the container information of the target container and the node information of the first node.
[0162] The rescheduling field includes at least an information field; the container information of the target container and the node information of the first node are located in the information field of the rescheduling field.
[0163] The rescheduling field also includes a reason field;
[0164] The reason field is used to store the reason why the system resources on the node where the container is located were not successfully allocated to the container according to the actual needs of the container;
[0165] The device also includes:
[0166] The first determination module is used to determine whether the reason field contains content;
[0167] The migration module is also used to: if the reason field has content, migrate the target container to a node in the second sub-cluster in the cloud data center based on the container information of the target container and the node information of the first node.
[0168] The rescheduling field includes at least an information field; the container information of the target container and the node information of the first node are located in the information field of the rescheduling field.
[0169] The rescheduling field also includes a duration field;
[0170] The duration field is used to store the duration during which system resources on the node where the container resides were not successfully allocated to the container according to the container's actual needs;
[0171] The device also includes:
[0172] The second determination module is used to determine whether the duration field contains content;
[0173] The third determination module is used to determine whether the timeout check policy of the main cluster manager has been enabled if the duration field contains content.
[0174] The migration module is also used to: migrate the target container to a node in the second sub-cluster in the cloud data center, based on the container information of the target container and the node information of the first node, when the timeout check policy of the main cluster manager is not enabled.
[0175] or,
[0176] The fourth acquisition module is used to acquire the duration of timeout checks that failed to allocate system resources on the first node to the target container according to the actual needs of the target container, provided that the timeout check policy of the main cluster manager is enabled.
[0177] The migration module is also used to migrate the target container to a node in the second sub-cluster in the cloud data center, based on the container information of the target container and the node information of the first node, if the duration exceeds the preset tolerance time of the main cluster manager.
[0178] In this application, the master cluster manager in the master cluster of the cloud data center monitors in real time whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed; if the information in the rescheduling field of the resource cluster object in the first sub-cluster of the cloud data center changes, the container information of the target container and the node information of the first node where the target container is located are obtained from the rescheduling field; the container information of the target container and the node information of the first node are: stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated according to the actual needs of the target container; the first node is located in the first sub-cluster; based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster of the cloud data center.
[0179] This application introduces a sub-cluster manager and a resource cluster object in the sub-cluster, and a master cluster manager in the master cluster. The sub-cluster manager, resource cluster object, and master cluster manager work together to achieve real-time resource awareness and resource balancing capabilities for each sub-cluster in a federated scenario within a cloud data center. This improves the real-time performance of detecting whether containers are normal / available, ensures that containers can be immediately scheduled to other nodes or clusters when they cannot obtain the system resources they actually need, reduces container downtime, avoids affecting the quality of cloud services provided by containers, and improves container availability and overall resource utilization of each sub-cluster in a federated scenario within a cloud data center.
[0180] In some embodiments of this application, an electronic device is also provided, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here.
[0181] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may include read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0182] Figure 6 is a block diagram of an electronic device 800 according to this application. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.
[0183] Referring to FIG6, the electronic device 800 may include one or more of the following components: processing component 802, memory 804, power supply component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0184] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0185] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, images, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), Magnetic Storage, Flash Memory, Disk, or Optical Disk.
[0186] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0187] Multimedia component 808 includes a screen that provides an output interface between electronic device 800 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also monitor the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When device 800 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0188] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0189] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0190] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 may monitor the on / off state of device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, the orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0191] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast operation information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.
[0192] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0193] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.
[0194] Figure 7 is a block diagram of an electronic device 1900 according to this application. For example, the electronic device 1900 can be provided as a server.
[0195] Referring to FIG7, the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0196] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0197] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0199] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0200] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0201] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0202] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0203] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0204] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0205] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0206] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A container scheduling method in a cloud environment, characterized in that, The method, applied to a sub-cluster manager in a sub-cluster within a cloud data center, includes: The scheduling status of containers on each node in the sub-cluster is monitored in real time. The scheduling status of containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container. If the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container, the rescheduling field in the resource cluster object in the sub-cluster shall at least store the container information of the target container and the node information of the first node, which is located in the sub-cluster.
2. The method of claim 1, wherein, The rescheduling field includes at least an information field; The rescheduling field in the resource cluster object in the sub-cluster stores at least the container information of the target container and the node information of the first node, including: The information field in the rescheduling field stores the container information of the target container and the node information of the first node; Secondly, the rescheduling field also includes a reason field; Accordingly, the method further includes: The reason why the system resources on the first node were not successfully allocated to the target container according to the actual needs of the target container; The reason is stored in the reason field of the rescheduling field.
3. The method of claim 1, wherein, The rescheduling field includes at least an information field; The rescheduling field in the resource cluster object in the sub-cluster stores at least the container information of the target container and the node information of the first node, including: The information field in the rescheduling field stores the container information of the target container and the node information of the first node; Secondly, the rescheduling field also includes a duration field; Accordingly, the method further includes: The duration during which the allocation of system resources on the first node to the target container according to the actual needs of the target container was unsuccessful; If the duration exceeds the preset tolerance duration of the sub-cluster manager, the duration is stored in the duration field of the rescheduling field.
4. A container scheduling method in a cloud environment, characterized by, The method, applied to the master cluster manager in a master cluster within a cloud data center, includes: Real-time monitoring of whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed; When the information in the rescheduling field of the resource cluster object in the first sub-cluster of the cloud data center changes, the container information of the target container and the node information of the first node where the target container is located are obtained from the rescheduling field. The container information of the target container and the node information of the first node are stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container. The first node is located in the first sub-cluster. Based on the container information of the target container and the node information of the first node, the target container is migrated to a node in the second sub-cluster in the cloud data center.
5. The method of claim 4, wherein, The rescheduling field includes at least an information field; the container information of the target container and the node information of the first node are located in the information field of the rescheduling field; The rescheduling field also includes a reason field; The reason field is used to store the reason why the system resources on the node where the container is located were not successfully allocated to the container according to the actual needs of the container. The method further includes: Determine whether the reason field contains content; If the reason field contains content, the step of migrating the target container to a node in the second sub-cluster in the cloud data center based on the container information of the target container and the node information of the first node is executed.
6. The method of claim 4, wherein, The rescheduling field includes at least an information field; the container information of the target container and the node information of the first node are located in the information field of the rescheduling field; The rescheduling field also includes a duration field; The duration field is used to store the duration during which system resources on the node where the container is located were not successfully allocated to the container according to the container's actual needs; The method further includes: Determine whether the duration field contains content; If the duration field contains content, determine whether the timeout check policy of the main cluster manager has been enabled; Without enabling the timeout check policy of the main cluster manager, perform the step of migrating the target container to a node in the second sub-cluster in the cloud data center based on the container information of the target container and the node information of the first node. or, With the timeout check policy enabled in the master cluster manager, obtain the duration during which the system resources on the first node were not successfully allocated to the target container according to the actual needs of the target container. If the duration exceeds the preset tolerance time of the main cluster manager, the step of migrating the target container to a node in the second sub-cluster in the cloud data center based on the container information of the target container and the node information of the first node is executed. 7.A container scheduling apparatus in a cloud environment, characterized by comprising: A sub-cluster manager for use in a sub-cluster within a cloud data center, the device comprising: The first monitoring module is used to monitor the scheduling status of containers on each node in the sub-cluster in real time. The scheduling status of the containers is used to indicate whether the system resources on the node where the container is located have been successfully allocated to the container according to the actual needs of the container. The first storage module is used to store, in the rescheduling field of the resource cluster object in the sub-cluster, at least the container information of the target container and the node information of the first node, where the first node is located in the sub-cluster, when the scheduling status indicates that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container. 8.A container scheduling apparatus in a cloud environment, characterized by comprising: A master cluster manager used in a master cluster in a cloud data center, the device comprising: The second monitoring module is used to monitor in real time whether the information in the rescheduling field of the resource cluster object in each sub-cluster of the cloud data center has changed. The third acquisition module is used to acquire the container information of the target container and the node information of the first node where the target container is located in the rescheduling field of the resource cluster object in the first sub-cluster in the cloud data center when the information in the rescheduling field of the resource cluster object in the first sub-cluster changes. The container information of the target container and the node information of the first node are stored in the rescheduling field of the resource cluster object in the first sub-cluster when the sub-cluster manager in the first sub-cluster detects that the system resources on the first node where the target container is located have not been successfully allocated to the target container according to the actual needs of the target container. The first node is located in the first sub-cluster. The migration module is used to migrate the target container to a node in the second sub-cluster in the cloud data center based on the target container's container information and the node information of the first node.
9. An electronic device, comprising: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the container scheduling method in a cloud environment as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the container scheduling method in a cloud environment as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-cluster exception processing method and device
CN112463535A
Service recovery method, service deployment method, server and storage medium
CN115314363A
Dispatching method and system based on distributed container cluster
CN118349352A
Container scheduling method and device in cloud environment, electronic equipment and storage medium
CN119668797A
Container cluster construction method and system
WO2023071576A1