Cluster network scheduling control method, apparatus, device, medium, and program product
Patent Information
- Application Number
- CN202610894032.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]相关技术中,调度器基于节点的资源容量等属性进行调度决策,导致将负载调度至节点后,出现运行故障等问题,进而使得数据面任务执行中断或数据丢失
[0011] As described above, the provided cluster network scheduling and control method, device, equipment, media, and program products first impose scheduling restrictions on the first node before allocating load to it. This prevents the container cluster management system from directly allocating load to the first node, which could lead to subsequent operational failures due to insufficient network capabilities of the first node. By assessing the network resource status of the first node and confirming that it has the capability to communicate across private networks via auxiliary networks, the scheduling restrictions are lifted, allowing the container cluster management system to allocate cluster workloads to the first node. This ensures that the container cluster management system schedules cluster workloads to nodes with cross-private network communication capabilities, thus avoiding abnormal cluster workload operation due to insufficient network readiness. Simultaneously, it eliminates the need for post-processing repairs such as error path troubleshooting and network reconfiguration caused by node mismatches, significantly reducing the frequency and cost of operational interventions.
Smart Images

Figure CN122783508A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a cluster network scheduling and control method, apparatus, equipment, medium, and program product. Background Technology
[0002] Kubernetes (Container Cluster Management System, or K8s for short) is an open-source container orchestration platform designed to automate the deployment, scaling, and management of containerized applications. In a Kubernetes environment, a node is the sole physical entity in the cluster that provides physical or virtual computing resources such as memory and storage; a workload must attach to a node to obtain resources and actually execute.
[0003] In related technologies, the scheduler makes scheduling decisions based on attributes such as the resource capacity of nodes, which can lead to problems such as operational failures after the load is scheduled to the nodes, resulting in interruption of data plane task execution or data loss. Summary of the Invention
[0004] In view of this, the purpose is to propose a cluster network scheduling and control method, device, equipment, medium and program product to solve or partially solve the above-mentioned technical problems.
[0005] Based on the above objectives, a cluster network scheduling and control method is proposed, including:
[0006] Obtain the first node and impose scheduling restrictions on the first node to prevent the container cluster management system from scheduling cluster workloads to the first node, wherein the first node is not assigned any workload; Obtain the network resource status corresponding to the first node, and in response to the network resource status being ready to complete, determine that the first node has the function of communicating across private networks through the auxiliary network. Remove the scheduling restrictions to allow the container cluster management system to allocate cluster workloads to the first node.
[0007] Based on the same concept, a cluster network scheduling and control device is also proposed, comprising: The data acquisition module is configured to acquire the first node and impose scheduling restrictions on the first node to prevent the container cluster management system from scheduling cluster workloads to the first node, wherein the first node is not assigned any workload. The status determination module is configured to obtain the network resource status corresponding to the first node, and in response to the network resource status being ready to complete, determine that the first node has the function of communicating across private networks through the auxiliary network. The restriction removal module is configured to remove the scheduling restrictions to allow the container cluster management system to allocate cluster workloads to the first node.
[0008] Based on the same concept, an electronic device is also proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described above.
[0009] Based on the same concept, a non-transitory computer-readable storage medium is also proposed, which stores computer instructions for causing a computer to perform the methods described above.
[0010] Based on the same concept, a computer program product is also proposed, comprising computer program instructions that, when run on a computer, cause the computer to perform the method described above.
[0011] As described above, the provided cluster network scheduling and control method, device, equipment, media, and program products first impose scheduling restrictions on the first node before allocating load to it. This prevents the container cluster management system from directly allocating load to the first node, which could lead to subsequent operational failures due to insufficient network capabilities of the first node. By assessing the network resource status of the first node and confirming that it has the capability to communicate across private networks via auxiliary networks, the scheduling restrictions are lifted, allowing the container cluster management system to allocate cluster workloads to the first node. This ensures that the container cluster management system schedules cluster workloads to nodes with cross-private network communication capabilities, thus avoiding abnormal cluster workload operation due to insufficient network readiness. Simultaneously, it eliminates the need for post-processing repairs such as error path troubleshooting and network reconfiguration caused by node mismatches, significantly reducing the frequency and cost of operational interventions. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram illustrating an application scenario for an example. Figure 2 A flowchart of a cluster network scheduling and control method as an example; Figure 3This is a schematic diagram of the architecture of a cluster network scheduling and control system as an example. Figure 4 A modular schematic diagram of the cluster network scheduling and control system as an example; Figure 5 This is a schematic diagram of the fixed node expansion process in an embodiment. Figure 6 This is a schematic diagram of the fixed node scaling-down process in an embodiment; Figure 7 This is a schematic diagram of the elastic node expansion process in an embodiment. Figure 8 This is a schematic diagram of the flexible node scaling-down process in an embodiment. Figure 9 This is a structural block diagram of a cluster network scheduling and control device according to an optional embodiment; Figure 10 This is a schematic diagram of the structure of an electronic device according to an optional embodiment. Detailed Implementation
[0014] It is understandable that the data involved in the plan (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.
[0015] The principles and spirit of the solution will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement the solution, and are not intended to limit the scope of the solution in any way. Rather, these embodiments are provided to make the solution more thorough and complete, and to fully convey the scope of the solution to those skilled in the art.
[0016] It is understood that before using the technical solutions of each embodiment in the solution, the user will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and the user's authorization will be obtained.
[0017] For example, upon receiving a user's proactive request, a prompt message can be sent to the user, explicitly informing them that the requested operation will require the acquisition and use of their personal information. This allows the user to choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the technical solution.
[0018] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the solution. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the solution.
[0020] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0021] The definitions of the terms involved are as follows: Kubernetes: Container Cluster Management System, or K8s for short, is an open-source container orchestration platform designed to automate the deployment, scaling, and management of containerized applications.
[0022] Operator: A custom Kubernetes controller that encapsulates domain-specific knowledge (such as creating blockchain nodes and networking) to achieve full lifecycle management of custom resources.
[0023] CRD: Custom Resource Definition (CRD), is the most popular API extension mechanism in Kubernetes. It allows developers to declare new resource types without modifying the core code.
[0024] CR: Custom Resource (CR) is an extension interface of the Kubernetes API that allows users to define and manage their own new resource types in addition to the default built-in resources (such as Pods and Deployments).
[0025] Celeborn typically consists of a Master node and Worker nodes, serving as an independent distributed system. The Master node manages the overall cluster status and resource allocation, supporting high availability (HA), while the Worker nodes handle the actual data read / write requests and storage, hosting Shuffle data in remote storage, thus freeing the computing cluster from dependence on local large-capacity disks.
[0026] ENI: Elastic Network Interface (ENI). In cloud computing, ENI is a virtual network interface card that can be freely bound and unbound between cloud servers (ECS) to achieve a highly available network architecture or manage traffic in different subnets.
[0027] Apache Celeborn (Apache Cluster) is an open-source, general-purpose intermediate data service system positioned as a unified intermediate data service (Remote Shuffle Service) for big data computing engines. Celeborn addresses the three major pain points of traditional Shuffle solutions in terms of performance, stability, and elasticity. Specifically, in terms of performance, it employs Push-based Shuffle, Aggregated Shuffle, Partition Split, and asynchronous read / write mechanisms to significantly reduce read / write latency for Shuffle (redistributed) data. Celeborn can reduce the amount of Shuffle data by approximately 40% and computational resource costs by approximately 30%. In terms of stability, the Master node achieves high availability (HA) based on the Raft (Ratis implementation) consensus protocol, supporting rolling upgrades and rapid failover. Worker nodes support a Revive mechanism; when a push fails, it does not immediately determine that a Worker has been lost, but instead pushes data to other Workers, achieving transparent fault recovery. It also supports graceful retirement (Decommission mode) and graceful shutdown (Graceful mode) of Shuffle data, ensuring no data loss during scaling down or rolling upgrades. In terms of elasticity, worker nodes can scale up or down elastically according to the load, while compute nodes (executors) do not need to rely on large-capacity local disks, realizing a storage-compute separation architecture. Celeborn has a built-in HPA (Horizontal Pod Autoscaling) mechanism, and the Master automatically adjusts the dynamic number of workers through a built-in traffic prediction algorithm.
[0028] Celeborn supports multiple computing engines such as Spark, Flink, and MapReduce. Celeborn's core components include Master, Worker, and Client. The Master is responsible for cluster resource allocation, Worker registration, and state management, using Raft for state synchronization and high availability, and provides an HTTP management interface (port 9098). The Worker handles read and write requests for Shuffle data, supports data merging and partition management, and provides RPC service ports (9090), Push port (9091), Fetch port (9092), Replicate port (9093), and Metrics port (9096). The Client includes LifecycleManager (control plane, managing Shuffle metadata) and ShuffleClient (data plane, responsible for data push and read), and supports Revive fault tolerance mechanisms.
[0029] Kubernetes Operator is a cloud-native application management model based on Kubernetes Custom Resource Definitions (CRDs) and custom controllers. The CNCF Operator White Paper defines the core capabilities of the Operator, specifically including encapsulating domain-specific operational knowledge and implementing full application lifecycle management through declarative APIs (Application Programming Interfaces). The core working principle of the Operator is the reconciliation loop, where the controller continuously listens for changes to CRD resources, compares the user-defined desired state (Spec) with the actual cluster running state (Status), and performs reconciliation operations to converge the actual state to the desired state. This model inherently possesses idempotency.
[0030] Common operator development frameworks include Kubebuilder and Controller Runtime. This solution uses the Controller Runtime framework, combined with Kubernetes' Informer-Watch mechanism, to implement event-driven programming.
[0031] Kubernetes (Container Cluster Management System, or K8s for short) is an open-source container orchestration platform designed to automate the deployment, scaling, and management of containerized applications. In a Kubernetes environment, a node is the sole physical entity in the cluster that provides physical or virtual computing resources such as CPU, memory, and storage. A workload must be attached to a node to obtain resources and actually execute.
[0032] Currently, the scheduler makes scheduling decisions based on node resource capacity and other attributes, completely unaware of the node's network capabilities. This results in the Celeborn workload being scheduled to a node that hasn't completed network preparation for cross-VPC (Virtual Private Network) operations. Celeborn's data plane network cannot establish stable communication with other services within the VPC, manifesting as the workload Pod being in a Running state but unable to perform cross-VPC data read / write operations. This leads to Shuffle task failures, data loss, or persistent connection timeouts. In other words, workload malfunctions occur, ultimately causing data plane task execution interruptions or data loss.
[0033] Based on the above description, the principles and spirit of the solution will be explained in detail below with reference to several representative implementation methods.
[0034] refer to Figure 1 , Figure 1 This is a schematic diagram of an application scenario for an optional embodiment. The application scenario includes a scheduler 101 of a container cluster management system and at least one node 102. Figure 1(Taking a single node as an example), a communication connection is established between scheduler 101 and node 102. Scheduler 101 is used to manage node 102 and schedule and assign job tasks to run on node 102. Node 102 is used to receive scheduling information or assigned job tasks sent by scheduler 101 and run the job tasks assigned by scheduler 101.
[0035] The following is combined with Figure 1 The above application scenarios are used to describe the cluster network scheduling and control method of exemplary implementation. It should be noted that the above application scenarios are only shown for the purpose of understanding the spirit and principle of the solution, and the implementation of the solution is not limited in any way. On the contrary, the implementation of the solution can be applied to any applicable scenario.
[0036] In one scenario, the solution provides a cluster network scheduling and control method applied to the container cluster controller (Operator), such as... Figure 2 As shown, it includes: Step 201: Obtain the first node and impose scheduling restrictions on the first node to prevent the container cluster management system from scheduling cluster workloads to the first node, wherein the first node is not assigned any workload.
[0037] In practice, the cluster network contains at least one node, which is the only actual carrier in the cluster network that provides physical or virtual computing resources such as CPU, memory, and storage. Cluster workloads need to be attached to a certain node to obtain resources and actually execute.
[0038] Obtain the first node in the cluster network, where the first node has not been assigned cluster workloads. Impose a scheduling constraint on the first node indicating that the network is not ready, to prevent the container cluster management system (Kubernetes) from scheduling cluster workloads to the first node.
[0039] Optionally, scheduling restrictions can be implemented using node taints or other equivalent scheduling constraint mechanisms. A node taint is a label placed on a node, which prevents cluster workloads from being scheduled to that node by default. The scheduling restrictions can be used to characterize at least one of the following states: the first node has not completed the mounting of auxiliary network resources; the auxiliary network resources have been mounted but routing rules have not yet been configured; or the target data plane link has not yet been verified as available.
[0040] Before the container cluster management system schedules the cluster workload to the first node, scheduling restrictions are imposed on the first node to prevent the container cluster management system's scheduler (Kubernetes scheduler) from scheduling the cluster workload to the first node. This avoids service instances running in an incorrect network environment that does not have the ability to communicate across private networks, which could lead to problems such as operational failures of the cluster workload.
[0041] Step 202: Obtain the network resource status corresponding to the first node. In response to the network resource status being ready to complete, determine that the first node has the function of communicating across private networks through the auxiliary network.
[0042] In specific implementation, the network resource status corresponding to the first node is obtained, wherein the network resource status is used to describe whether the first node has the ability to communicate across private networks. Specifically, the network resource status includes "ready to complete" or "not ready to complete". If the network resource status is determined to be "not ready to complete", it means that the first node does not have the ability to communicate across private networks (across VPCs), and scheduling restrictions on the first node are maintained until the network resource status becomes "ready to complete".
[0043] If the network resource status is determined to be ready, it is determined that the first node has the function of communicating across private networks through the auxiliary network. That is, the first node can communicate with the target private network through the auxiliary network after startup, avoiding subsequent post-repair work such as load eviction and rescheduling, and reducing operation and maintenance intervention costs.
[0044] Step 203: Remove the scheduling restriction to allow the container cluster management system to allocate cluster workloads to the first node.
[0045] In practice, once the network resource status corresponding to the first node is determined to be ready, that is, once it is determined that the first node has the function of communicating across private networks through the auxiliary network, the scheduling restrictions imposed on the first node are lifted. At this time, the container cluster management system can schedule the cluster workload to the first node, that is, allow the container cluster management system to allocate the cluster workload to the first node.
[0046] The above solution applies scheduling restrictions to the first node before allocating load to it. This prevents the container cluster management system from directly allocating load to the first node, which could lead to subsequent cluster workload failures due to insufficient network capabilities on the first node. The network resource status of the first node is assessed. Once it is determined that the first node has the capability to communicate across private networks via an auxiliary network, the scheduling restrictions are lifted, allowing the container cluster management system to allocate cluster workloads to it. This ensures that the container cluster management system schedules cluster workloads to nodes with cross-private network communication capabilities, thus avoiding cluster workload malfunctions due to insufficient network readiness. Furthermore, it eliminates the need for post-implementation repairs such as error path troubleshooting and network reconfiguration caused by node mismatches, significantly reducing the frequency and cost of operational interventions.
[0047] In another scenario, since both network resource mounting and routing configuration are indispensable during network preparation, the network resource status of the first node is considered complete only after the auxiliary network resources are mounted to the first node and the routing of the network resources mounted on the first node is configured. Specifically, obtaining the network resource status of the first node in step 202 includes: Step 301: Obtain the node attributes corresponding to the first node, obtain the auxiliary network resources according to the node attributes, and attach the auxiliary network resources to the first node so that the first node has a communication interface for cross-private network communication. Step 302: Configure routing for the network resources mounted on the first node. In response to the completion of routing configuration, determine that the network resource status corresponding to the first node is ready to be completed.
[0048] In specific implementation, the node attributes corresponding to the first node are obtained. The node attributes represent the inherent characteristics or metadata of the first node, which are used to uniquely determine what auxiliary network resources the first node needs in the target private network and how to configure the routing. The node attributes include infrastructure topology information, network identifier, VPC identifier, node identity identifier, etc.
[0049] The target network configuration corresponding to the first node is determined based on the node attributes, and auxiliary network resources are obtained based on the target network configuration. For example, the target network configuration includes which type of auxiliary elastic network interface card (NIC) the first node needs to create or mount, which VPC IP address to assign, and which security groups to bind, etc.
[0050] After determining the auxiliary network resources, the auxiliary network resources are mounted to the first node. This means that the network card is attached to the node's operating system layer. In other words, the auxiliary network card is provided to the first node through the mounting operation. At this time, the first node has a communication interface for cross-private network communication.
[0051] Routing is configured for the network resources attached to the first node to define the logical path of the data plane. If only auxiliary network resources are attached to the first node, issues such as unassigned IP addresses, disabled link states, or default traffic routes not pointing to the network interface card may exist, and actual data plane communication capabilities are not yet available. Therefore, after determining the routing configuration, it is crucial to ensure that data packets originating from the Celeborn workload can be correctly forwarded through the auxiliary network according to the expected strategy, thus confirming that the network resource status corresponding to the first node is ready.
[0052] The above scheme provides the basic interface for mounting and defines the logical path of the data plane through routing configuration. Only after it is determined that the auxiliary network resources are mounted to the first node and the routing configuration of the network resources mounted on the first node is completed can the status of the network resources corresponding to the first node be determined as ready to complete, ensuring that the Celeborn workload has complete and correct data plane communication capabilities from the node to the target VPC.
[0053] In another scenario, after obtaining the auxiliary network resource based on the node attributes, it is first necessary to determine whether the auxiliary network resource already exists in the cluster. If it does exist, the auxiliary network resource is directly attached to the first node. If it does not exist, the auxiliary network resource needs to be created first. That is, attaching the auxiliary network resource to the first node as described in step 301 specifically includes: Step 401: Check whether the auxiliary network resources are configured in the cluster network; Step 402: In response to the existence of the auxiliary network resource, mount the auxiliary network resource to the first node; or, Step 403: In response to the absence of the auxiliary network resource, create the auxiliary network resource and attach the target auxiliary network resource to the first node.
[0054] In practice, it is checked whether the auxiliary network resource is configured in the cluster network, that is, whether an auxiliary network resource that meets the node attribute requirements already exists in the cluster network. Here, the auxiliary network resource is relative to the node's default main network, specifically referring to an independent network plane prepared for a particular workload. The auxiliary network resource provides a dedicated Shuffle data transmission channel, and its readiness status is perceived by the Kubernetes scheduling mechanism through node taints, thereby ensuring that workloads only run on nodes where the auxiliary network is fully ready, achieving precise scheduling awareness and network resource isolation.
[0055] If the auxiliary network resource already exists and is not currently occupied, it can be directly mounted to the first node to avoid resource waste and overhead caused by repeatedly creating the auxiliary network resource.
[0056] If the auxiliary network resource does not exist, the auxiliary network resource must first be created according to the target network resource configuration determined by the node attributes. After creation, the mounting operation is performed, that is, the auxiliary network resource is created first, and then the target auxiliary network resource is mounted to the first node.
[0057] Optionally, when creating secondary network resources, it is also necessary to associate corresponding security control policies. These policies define which source IPs, ports, or protocols can access the secondary network interface card (NIC), and which target network resources the NIC can access externally. For example, the security control policy may be binding a specified security group or setting network access control list rules.
[0058] Understandably, without associated security control policies, data plane communication for Celeborn workloads might be blocked due to overly strict default security rules, such as the inability to receive shuffle data from the target VPC. Furthermore, insecure default open policies could introduce unexpected network access risks. Therefore, by associating security control policies when creating secondary network resources, it ensures that these resources have access boundaries that match Celeborn's business needs and security baseline from the outset, thus avoiding the brief exposure period or configuration omissions that can occur if associated separately after creation.
[0059] The above scheme first determines whether auxiliary network resources exist in the cluster. If they exist, they are directly mounted to the first node. If they do not exist, the auxiliary network resources are created first and then mounted, ensuring the idempotency of the operation. That is, no matter how many times the controller restarts, retries, or concurrently calls the mounting process, the final result is that the auxiliary network resource is correctly mounted and mounted only once, avoiding resource leaks or mounting conflicts caused by repeated creation. At the same time, it improves execution efficiency. When the auxiliary network resource already exists, the time-consuming creation step is skipped and it is reused directly, significantly shortening the overall time from the start of preparation to network readiness of the first node.
[0060] In another scenario, when configuring routing for the network resources attached to the first node, if the communication path of the network resources matches the corresponding interface, the routing configuration is considered complete. Specifically, step 302, which involves configuring routing for the network resources attached to the first node, includes: Step 501: Obtain the communication path of the network resource and the type identifier corresponding to the communication path; Step 502: Obtain the first interface based on the type identifier, match the communication path with the first interface, and complete the routing settings.
[0061] In specific implementation, for network resources that have been mounted to the first node, the communication path of the network resources is obtained. The communication path represents the network logical link that a data packet needs to pass through from the Celeborn Pod on the first node to the peer address in the target VPC or other Celeborn components.
[0062] Obtain the type identifier corresponding to the communication path, whereby the type identifier is used to distinguish different VPCs. Determine the first interface corresponding to the communication path based on the type identifier; the first interface is the network interface card (NIC) corresponding to the network resource. Then, match the communication path with the first interface, that is, match and bind the communication path to the first interface to ensure that traffic from a specific VPC can be correctly routed to the corresponding interface.
[0063] Once the matching is complete, the routing rules take effect. At this point, a correct forwarding relationship is established between the communication path of the network resource and the first interface, thus completing the routing settings for the network resources mounted on the first node.
[0064] Through the above scheme, the essence of routing is to establish a mapping relationship between the target network segment and the outgoing interface. When the communication path matches the first interface, the operating system kernel has obtained the minimum necessary information for forwarding data packets. That is, for any traffic destined for this communication path, it can uniquely determine which interface should be sent out. Once the matching is complete and the routing rules take effect, data packets sent by the Celeborn workload can be correctly guided to the secondary network resources, and then communicate with the target VPC. Therefore, it can be considered that the routing settings have met the basic data plane requirements.
[0065] In another scenario, when determining the first interface corresponding to the communication path, the network resource is identified as either data plane traffic or control plane traffic based on the type identifier, thus determining whether to bind it to the secondary network interface or the primary network interface. By retaining the control flow within the original cluster network and allowing the data flow to enter the target VPC via the secondary network, cross-VPC deployment is achieved. Specifically, obtaining the first interface based on the type identifier in step 502 includes: Step 601: In response to the type identifier being a user-side identifier, identify the auxiliary network interface corresponding to the auxiliary network resource, configure the address of the auxiliary network interface, and designate the auxiliary network interface as the first interface; or, Step 602: In response to the type identifier being a cluster-side identifier, the main network interface is designated as the first interface, wherein the main network interface is the cluster network interface.
[0066] In practice, the type is determined by whether it is a user-side identifier or a cluster-side identifier. If the type identifier is a user-side identifier, it indicates that the communication path belongs to data plane-oriented service traffic, i.e., the path for data transmission of the Celeborn workload. At this point, the auxiliary network interface corresponding to the auxiliary network resource is identified, such as by searching for the auxiliary network interface corresponding to the auxiliary network resource based on the logical interface name.
[0067] Configure the address of the auxiliary network interface to ensure it has a valid communication address. After completing the address configuration, officially bind the auxiliary network interface as the primary interface of the current communication path, thus making the auxiliary network interface the first interface.
[0068] For example, configuring the address of a secondary network interface can include assigning or confirming a secondary IP address, setting network layer parameters such as a subnet mask and a broadcast address.
[0069] If the type identifier is determined to be a cluster-side identifier, it indicates that the communication path belongs to the cluster control plane communication path. At this time, the first interface corresponding to the communication path is set as the main network interface, that is, the first interface is the cluster network interface, ensuring that the control plane traffic goes through the default main network card.
[0070] The above approach results in high network bandwidth consumption due to the large amount of cross-node transmission of shuffle data involved in Celeborn's data read / write path. If this traffic shares the main network interface with management traffic such as accessing the Kubernetes control plane, pulling images, and reporting heartbeats, sudden surges in data plane traffic can easily lead to control plane request timeouts or unavailability of basic services. By forcing data plane traffic identified by user VPCs to use the secondary network interface, while management plane traffic identified by cluster VPCs uses the main network interface, physical or logical path separation is achieved at the node level. This avoids interference or congestion in management plane communication caused by data plane issues, and also reduces the risk of data plane network failures, such as VPC routing jitter or secondary network card malfunctions, affecting the core management link between the node and the Kubernetes control plane.
[0071] In another scenario, when the number of nodes required to handle the cluster workload exceeds the number of unallocated nodes in the current cluster's node pool, new nodes can be added to the node pool. However, scheduling restrictions must first be imposed, and these restrictions are lifted once the network resource status of the new node is confirmed to be ready. Specifically, the method further includes: Step 701: Determine the target number of nodes corresponding to the cluster workload and obtain the number of second nodes, wherein the second node is not assigned any load and the first node and the second node belong to the same node pool; Step 702: In response to the target number of nodes being greater than the sum of the number of the first node and the second node, a third node is added to the node pool, wherein the third node is not assigned any load. Step 703: Apply scheduling restrictions to the third node; in response to the network resource status corresponding to the third node being ready to complete, remove the scheduling restrictions.
[0072] In specific implementation, the target number of nodes corresponding to the cluster workload is determined, where the target number of nodes is the number of nodes required for the cluster workload. The number of second nodes contained in the node pool to which the first node belongs is obtained, where the second nodes are other nodes in the node pool that are not assigned load, excluding the first node; that is, the second nodes are not assigned load, and the first node and the second node belong to the same node pool.
[0073] The sum of the number of the first node and the number of the second node is the total number of nodes in the node pool that are not assigned load. The target number of nodes is then compared to the sum of the numbers of the first and second nodes.
[0074] If the target number of nodes is less than the sum of the number of the first node and the second node, it means that the number of nodes required for the cluster workload is less than the number of nodes in the node pool that can be allocated workload. That is, the number of nodes in the node pool that can be allocated workload is sufficient to allocate the cluster workload, so there is no need to expand the node pool.
[0075] If the target number of nodes is greater than the sum of the numbers of the first and second nodes, it means that the number of nodes required to handle the cluster workload is greater than the number of nodes in the node pool that can already allocate workload. To ensure the smooth allocation of cluster workload, a third node is added to the node pool. It is understood that since the purpose of adding a third node to the node pool is to ensure the smooth allocation of cluster workload, the third node is also a node that has not yet been allocated workload.
[0076] A scheduling restriction indicating network insecurity is imposed on the third node to prevent the Kubernetes (K8s) cluster management system from scheduling cluster workloads to the third node. The network resource status corresponding to the third node is obtained, where the network resource status includes ready or not ready. If the network resource status is determined to be not ready, it means that the third node does not have the ability to communicate across private networks (across VPCs), and the scheduling restriction on the third node is maintained until the network resource status becomes ready.
[0077] If the network resource status is determined to be ready, it is determined that the third node has the function of communicating across private networks through the auxiliary network. That is, the third node can communicate with the target private network through the auxiliary network after startup, avoiding subsequent repair work such as load eviction and rescheduling, and reducing operation and maintenance intervention costs.
[0078] Once the network resource status of the third node is determined to be ready, i.e., once it is determined that the third node has the function of communicating across private networks through the auxiliary network, the scheduling restrictions imposed on the third node are lifted. At this time, the container cluster management system can schedule the cluster workload to the third node, i.e., allow the container cluster management system to allocate the cluster workload to the third node.
[0079] The above solution adds new nodes to the node pool when the number of nodes required by the Celeborn workload exceeds the number of available nodes in the current node pool. Simultaneously, scheduling restrictions are applied to the newly added nodes, and these restrictions are lifted once the auxiliary network resources are ready. This avoids runtime failures such as incorrect network paths and cross-VPC unreachability caused by new nodes prematurely receiving cluster workloads due to unprepared auxiliary networks. It completely eliminates the post-scaling costs in scaling scenarios and ensures that the expanded nodes, after the restrictions are lifted, have stable data plane communication capabilities completely equivalent to existing nodes. This decouples and coordinates the elastic scaling of the node pool with Celeborn's strong dependence on the target VPC network, meeting the dynamic node requirements of the workload without sacrificing reliability.
[0080] In another scenario, to avoid exceeding the capacity of the underlying infrastructure, when the target number of nodes corresponding to the cluster workload is greater than the sum of the numbers of the first and second nodes, it is first necessary to determine the total number of nodes in the node pool. Only when the total number of nodes is determined to be less than a preset threshold can a new node be directly added to the node pool. This ensures the effectiveness of the expansion operation and that each node in the node pool ultimately has stable and usable target data plane communication capabilities. Specifically, adding a third node to the node pool in step 702 includes: Step 801: Obtain the node pool to which the first node belongs, and find the total number of all nodes in the node pool; Step 802: In response to the total number of nodes being less than a preset threshold, add a third node to the node pool; or, Step 803: In response to the total number of nodes being greater than a preset threshold and a node addition instruction being received, a third node is added to the node pool.
[0081] In specific implementation, the node pool to which the first node belongs is obtained, and the total number of nodes in the node pool is found. That is, the total number of nodes in the node pool to which the first node belongs is obtained, where all nodes include both nodes that have been assigned load and nodes that have not been assigned load. The total number of nodes is compared with a preset threshold, where the preset threshold is a pre-set threshold for the number of nodes in the node pool, and the preset threshold corresponds to the carrying capacity of the underlying infrastructure.
[0082] If the total number of nodes is less than the preset threshold, it means that the number of nodes in the node pool has not yet reached the preset maximum value, and expansion can be carried out directly, that is, a third node can be added directly to the node pool.
[0083] If the total number of nodes exceeds a preset threshold, it means the node pool has reached the set threshold. Since cloud platforms typically have hard quotas for auxiliary network resources (such as the number of elastic network interfaces and IP addresses) under each VPC, adding more nodes will cause the node pool size to exceed the capacity of the underlying network resources, resulting in a failure where new nodes cannot complete auxiliary network preparation due to quota exhaustion. Furthermore, the processing capacity of the Kubernetes control plane components and the external controller responsible for preparing auxiliary networks and route tuning for Celeborn is not unlimited. When the node pool size exceeds a certain critical value, operations such as node heartbeats, status updates, and network resource creation requests will significantly increase the load on these components, potentially leading to response delays, timeouts, or even controller crashes.
[0084] Therefore, in order to effectively control infrastructure costs and cloud resource quotas, and at the same time maintain the stability of the cluster control plane and network preparation process, it is prohibited to directly add a third node to the node pool when the total number of nodes exceeds a preset threshold.
[0085] If a node addition instruction is received when the total number of nodes exceeds the preset threshold, it indicates that the administrator has explicitly initiated an active expansion. In other words, even though the total number of nodes exceeds the preset threshold, the original threshold can still be exceeded to add a third node to the node pool.
[0086] If the total number of nodes exceeds the preset threshold and no node addition instruction is received, it means that the number of nodes is greater than the upper limit and no active expansion instruction has been received. Therefore, adding nodes to the node pool is prohibited at this time.
[0087] The above scheme ensures that the number of nodes in the node pool must not reach a certain threshold before expansion can begin. This guarantees that the node pool size remains within the infrastructure's capacity, specifically including the number of subnet IP addresses, the quota limit for auxiliary network resources, and the processing capacity of the network controller. If expansion continues once the threshold is reached, the newly added nodes may fail to allocate auxiliary network resources or complete route optimization, resulting in them remaining in a state of scheduling constraints and unable to truly handle Celeborn's workload. This wastes expansion operations and fails to alleviate resource shortages.
[0088] In another scenario, all nodes in the node pool are divided according to their node attributes, including elastic nodes and fixed nodes. Fixed nodes are used to carry a stable baseline load; they are added during cluster creation or through static configuration of the node pool and do not automatically scale up or down with transient workload demands. Elastic nodes support automatic scaling up and down; they are dynamically added when cluster resources are insufficient and automatically reclaimed when the load decreases. Therefore, when the target number of nodes exceeds the sum of the first and second nodes, a third node is added to the node pool, where the node attribute of the third node is elastic node.
[0089] Optionally, if it is determined that an elastic node should be added to the node pool, the number of existing elastic nodes in the node pool is obtained. If the number of existing elastic nodes is less than a preset elastic node threshold, an elastic node is added to the node pool; or, if the number of existing elastic nodes is greater than the preset elastic node threshold, a prompt message is output to suggest adjusting the elastic node threshold.
[0090] After outputting the prompt message, the adjusted elastic node threshold is obtained. In response to the adjusted elastic node threshold being greater than the existing number of elastic nodes, an elastic node is added to the node pool.
[0091] Specifically, when deciding to add an elastic node to the node pool, the number of existing elastic nodes in the pool must first be determined. If the number of existing elastic nodes is less than a preset elastic node threshold, an elastic node is added to the node pool. If the number of existing elastic nodes is greater than the preset elastic node threshold, a prompt message is output to suggest adjusting the elastic node threshold. Once the adjusted elastic node threshold is greater than the number of existing elastic nodes, an elastic node is added to the node pool.
[0092] Optionally, adding a fixed node can be achieved through a fixed node add command. That is, upon receiving a fixed node add command, the node to be added corresponding to the command is determined, and scheduling restrictions are imposed on the node to be added. When the network resource status corresponding to the node to be added is ready, the scheduling restrictions are lifted to allow the container cluster management system to allocate cluster workloads to the node to be added.
[0093] The above solution addresses the issue of adding new nodes to the node pool when the cluster workload requires too many nodes. Since elastic nodes can be created on demand and automatically scaled down and released after the load decreases, the cost waste of retaining a large number of nodes for short-term high-load tasks is avoided. Therefore, adding a third node with the attribute of an elastic node to the node pool ensures the stable operation of Celeborn's workload while achieving both flexibility and economy in resource utilization.
[0094] In another scenario, if a scaling-down command is received to reduce the number of nodes in the node pool, the reduced number of nodes must first be compared with the number of nodes already allocated load to determine whether to directly delete nodes. Specifically, the method further includes: Step A01: In response to receiving a scaling-down instruction, obtain the target number corresponding to the scaling-down instruction, wherein the scaling-down instruction is used to reduce the number of nodes in the node pool; Step A02: Obtain the number of fourth nodes, and shrink the node pool according to the number of fourth nodes and the target number to obtain a new node pool, wherein the fourth nodes have been assigned load.
[0095] In specific implementation, if a scaling-down instruction is received, the scaling-down instruction is used to reduce the number of nodes in the node pool, that is, it indicates that the nodes in the node pool need to be deleted. The target number corresponding to the scaling-down instruction is obtained, where the target number is the number of nodes in the node pool after the reduction. Optionally, the scaling-down instruction includes a scaling-down instruction for fixed nodes or a scaling-down instruction for elastic nodes.
[0096] Obtain the number of fourth nodes, which is the number of nodes in the node pool that have been assigned load. Compare the number of fourth nodes with the target number, and then shrink the node pool according to the comparison result to obtain a new node pool.
[0097] With the above scheme, when a scaling-down instruction is received to reduce the number of nodes in the node pool, the reduced number of nodes is first compared with the number of nodes that have been allocated load. This avoids directly executing the scaling-down instruction and deleting nodes from the node pool. It also avoids the interruption of critical tasks, data loss, or unrecoverable residual resources caused by the forced deletion of nodes, which could lead to system-level failures and high repair costs.
[0098] In another scenario, if the reduced number of nodes is less than the number of nodes already allocated workload, it is necessary to first ensure that the cluster workload jobs scheduled on the nodes to be deleted are completed before deleting the nodes. Specifically, step A02, which involves scaling down the node pool based on the number of the fourth node and the target number to obtain a new node pool, includes: Step B01: In response to the target number being less than the number of the fourth node, obtain the node to be deleted; Step B02: Apply scheduling restrictions to the node to be deleted to prevent the container cluster management system from scheduling new cluster workloads to the node to be deleted; Step B03: After the cluster workload job corresponding to the node to be deleted is completed, the node to be deleted is deleted, and a new node pool is obtained.
[0099] In practice, upon receiving a scaling-down command, the reduced number of nodes is compared with the number of nodes already assigned load. If the target number is less than the number of the fourth node (i.e., the reduced number of nodes is less than the number of nodes already assigned load), then the nodes already assigned load must be deleted. The nodes to be deleted are then identified, where the nodes to be deleted are those that need to be removed when the scaling-down command is executed.
[0100] Optionally, the process of determining the node to be deleted includes: In response to receiving a scaling-down command, the scaling-down amount is determined based on the target number corresponding to the scaling-down command and the total number of nodes, wherein the scaling-down amount is the number of nodes to be deleted. The identification information corresponding to each node in the node pool is obtained, and based on the identification information, the node with the index corresponding to the reciprocal of the scaling-down amount is selected as the node to be deleted. Specifically, each node in the node pool has corresponding identification information; for example, the identification information is a label. Upon receiving a scaling-down instruction, the scaling-down amount is determined based on the target number corresponding to the scaling-down instruction and the total number of nodes. The scaling-down amount is the number of nodes to be deleted. This is obtained by subtracting the target number from the total number of nodes. Based on the identification information, the node corresponding to the reciprocal of the scaling-down amount is selected as the node to be deleted.
[0101] For example, if the nodes in the node pool are numbered from 1 to 20, and the number of nodes to be deleted is determined to be 10, then nodes numbered from 11 to 20 are selected as the nodes to be deleted.
[0102] A scheduling restriction is imposed on the node to be deleted to prevent the container cluster management system from scheduling new cluster workloads to the node to be deleted. That is, the node to be deleted is only allowed to complete its corresponding cluster workload, while the container cluster management system is prohibited from allocating new cluster workloads to the node to be deleted.
[0103] Because Celeborn is a Shuffle service, its nodes typically cache the shards or metadata of the currently executing Shuffle data. Deleting a node before the job is complete will cause temporary data on that node to become inaccessible instantly. Upstream tasks will fail because they cannot read intermediate results, potentially triggering a recalculation or crash of the entire computation job. Furthermore, if Celeborn lacks fault tolerance for such forced deletions, such as an insufficient replication mechanism, data loss may occur, making the job unrecoverable.
[0104] Therefore, once the cluster workload job corresponding to the node to be deleted is completed, there are no unfinished cluster workload tasks on the node to be deleted, and the node to be deleted is deleted to obtain a new node pool.
[0105] The above approach ensures that when scaling down the node pool, the cluster workload jobs on the node to be deleted are completed before the node is deleted, guaranteeing complete job execution and preventing data loss or computational failure. Furthermore, after job completion, the cluster workload is either idle or completed; deleting the node at this point releases resources and avoids issues such as connection leaks and temporary file remnants caused by forcibly killing processes.
[0106] In another scenario, if the reduced number of nodes exceeds the number of nodes already allocated load, the nodes to be deleted can be directly deleted. Specifically, step A02 involves scaling down the node pool based on the number of the fourth node and the target number to obtain a new node pool, which includes: Step C01: In response to the target number being greater than the number of the fourth node, obtain the node to be deleted; Step C02: Delete the node to be deleted to obtain a new node pool.
[0107] In practice, upon receiving the scaling-down instruction, the number of nodes reduced is compared with the number of nodes already assigned load. If the target number is greater than the number of the fourth node, meaning the number of nodes reduced is greater than the number of nodes already assigned load, then after the scaling-down, there are still other nodes besides those already assigned load.
[0108] In other words, even if the number of nodes in the node pool is reduced, it will not affect the nodes that have already been assigned load. Therefore, after obtaining the node to be deleted, the node to be deleted can be deleted directly to obtain a new node pool.
[0109] Using the above approach, if the number of nodes reduced is greater than the number of nodes already assigned a load, the nodes to be deleted can be directly deleted. This reduces the number of nodes while ensuring the normal operation of the nodes already assigned a load, thereby lowering cloud resource costs and reducing the idle waste of computing and network resources.
[0110] In another scenario, upon receiving a cluster deletion command, all network resources associated with the cluster are first cleared. After this process is complete, the cluster is deleted. Specifically, the method further includes: Step D01: Receive a cluster deletion instruction, wherein the cluster deletion instruction is used to indicate the deletion of the cluster where the first node is located; Step D02: Obtain all network resources configured with the cluster and clear all network resources; Step D03: After confirming that all network resources have been cleared, delete the cluster.
[0111] In practice, upon receiving a cluster deletion command, it indicates that the cluster corresponding to the node pool to which the first node belongs needs to be deleted. At this point, all network resources configured for the cluster are retrieved, and then all network resources are cleared. Once it is confirmed that all network resources have been cleared, the cluster is deleted.
[0112] Specifically, to capture deletion events and interrupt the default deletion process at the earliest stage of Kubernetes object deletion, preventing the system from removing the object from the API Server (gateway server) before external resources have been cleaned up, the system first checks whether the CelebornCluster (cluster resource) has entered the deletion phase and confirms whether the object still retains a finalizer. The presence of the finalizer allows the controller to execute custom cleanup logic after the object is marked for deletion but before it is actually deleted. CelebornCluster refers to the overall description and configuration object of the Celeborn cluster defined through Kubernetes Custom Resources (CR).
[0113] Next, because load balancing typically spans multiple Kubernetes services or nodes, Kubernetes' default Service deletion behavior may only unbind the backend associated with the local cluster, but will not actively delete the load balancer instance itself created externally. Instead of relying on cascading deletion, explicit cleanup is performed on external load balancers and their associated resources (such as listeners, backend server groups, forwarding rules, etc.). Explicit cleanup means that the Celeborn controller directly calls the cloud API or underlying management interface to release the load balancer resources in the correct order (e.g., unbind the backend first, then delete the listener, and finally delete the instance), thereby avoiding remnants.
[0114] After confirming that relevant network resources (including but not limited to load balancers, secondary network interface cards, VPC routing policies, etc.) have been released from occupation, the corresponding external network resources should be released. For example, ensure that no cluster workloads are using a particular ENI, and that no ongoing Shuffle data streams are referencing a particular VPC routing entry. Only after the occupation relationship has been safely released can the actual resource release operation (such as deleting the ENI or revoking the routing table entry) be performed, preventing network anomalies or data loss caused by resources being forcibly released while still being used.
[0115] The finalizer is removed only after it has been confirmed that critical dependent resources, namely all the aforementioned external network resources, have been safely cleaned up. Once the finalizer is removed, Kubernetes will actually delete the CelebornCluster object from the API Server and automatically cascade the deletion of other related resources within the cluster (such as Pods, StatefulSets, ConfigMaps, etc.). This ensures that the cleanup of all external network resources is completed before the deletion of internal cluster resources, making the deletion process orderly, reversible, and observable.
[0116] The above solution eliminates the remnants of external network resources, preventing resource leaks and quota waste. It also prevents cluster rebuilding failures due to network resource conflicts, improving the success rate and reliability of cluster rebuilding. Furthermore, by combining finalizers and explicit cleanup, the previously uncontrollable process of deleting external dependencies is incorporated into Kubernetes' native lifecycle management paradigm. This makes the deletion of large-scale infrastructure idempotent and transactional, ensuring that unfinished cleanup steps continue even if the controller restarts during deletion. Finally, it reduces operational risks. If problems are discovered during cleanup (e.g., the load balancer is still processing traffic), deletion can be stopped by retaining the finalizer, providing a safe window for manual intervention or rollback.
[0117] In another scenario, the solution provides a cluster network scheduling and control method applied to a cluster network scheduling and control system, enabling cluster deployment based on a container cluster controller (Operator) in a Kubernetes environment. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the architecture of a cluster network scheduling and control system.
[0118] like Figure 3As shown, the cluster network scheduling and control system includes a management and control layer, an infrastructure orchestration layer, and a data service layer. The management and control layer consists of a package manager (Helm) and a container cluster controller (Operator). Helm is the standard package management tool in the Kubernetes ecosystem, used to package, version, and distribute deployment templates (Charts) of Kubernetes applications. The Helm Chart parameterizes the application's Kubernetes resource inventory, providing customizable configuration items through values.yaml, and supports installation, upgrades, and rollbacks with a single command. Helm supports lifecycle hooks (pre-install, post-install, pre-upgrade, post-upgrade, pre-delete, post-delete) to automate operations at specific deployment stages. The two-tier architecture of Helm hosting Operators and Operators managing CR (Custom Resource) instances is a best practice for cloud-native application distribution provided by the CNCF (Cloud Native Computing Foundation) community.
[0119] The package manager (Helm) manages the Operator's installation, upgrades, and rollbacks, as well as the registration and updates of Celeborn CRDs (custom resource definitions). The Operator uses a tuning loop based on the Controller Runtime, monitoring CelebornCluster CRs (CelebornCluster custom resources) and Node objects. CelebornCluster specifically refers to an independent cluster of multiple servers dedicated to providing high-performance and highly available Shuffle services. The Operator includes two types of controllers: CelebornClusterReconciler and NodeReconciler, responsible for cluster-level resource orchestration and node-level auxiliary network interface card lifecycle management, respectively.
[0120] Optionally, the API group for the CelebornCluster CRD is celeborn.celeborn.apache.org, version v1alpha1, and scoped as Namespaced. The API group is used to logically categorize Kubernetes API resources.
[0121] The core fields of the user-defined expected state (Spec) include Image (Celeborn image address), Version (version number), AccountId (cloud account identifier), Config (Celeborn configuration key-value pairs), Network (network configuration), Master, Worker, and DynamicWorker. Master, Worker, and DynamicWorker define the replicas, resources, node pools, affinity, tolerance, volume configuration, and graceful exit time of the Master, fixed Worker, and elastic Worker, respectively. AccountId is used to call VPC / VKE related interfaces. Config contains parameters such as Master HA, number of fixed Workers, and upper and lower limits for elastic Workers. Network includes at least the user-side VPC, subnet mapping, security groups, and control plane network segment.
[0122] The core fields of the cluster's actual status include State (lifecycle stage), MasterReplicas, WorkerReplicas (actually ready replicas), AttachedENIs (the binding relationship between nodes and secondary network interfaces), AppliedDynamicConfig (the dynamically configured cursor that has been actually applied), and PendingDeleteNodes (the set of nodes awaiting scaling down or retirement checks). Among these, Network and AttachedENIs together constitute the declarative and factual description of the multi-VPC network connectivity status. Config and AppliedDynamicConfig together form the basis for idempotent control of scaling up or down fixed or elastic nodes.
[0123] Using Helm Charts to manage the Operator itself and all its required permissions and resources is a recommended practice for delivering Operators as standardized platform capabilities. This Chart not only contains the core CelebornOperator Deployment but also manages its dependent CRD definitions, RBAC permissions formed by ClusterRole and ClusterRoleBinding, ServiceAccounts, and auxiliary resources such as health checks, Leader Election, and network policies. Unified management via Helm enables standardized installation of the Operator, allowing platform capabilities to be deployed with a single click, just like ordinary applications. Secondly, the centralized definition of CRD, RBAC, and controller image versions within the Chart facilitates unified upgrades and avoids component version mismatches. When deployment issues arise, Helm's native rollback capability allows the Operator to quickly revert to the previous stable version, significantly reducing fault recovery time. Furthermore, because the Helm Chart encapsulates all configuration and permission declarations, users can easily replicate the same Celeborn automation management capabilities across multiple clusters, ensuring consistent operational behavior across different environments.
[0124] Optionally, use `kubectl apply static YAML` instead of Helm. Use Helm or an equivalent package management method to uniformly publish, upgrade, and rollback the Operator ontology, CRD, and its permission resources, thereby improving the delivery consistency and operational stability of this type of automated deployment capability in the Kubernetes environment.
[0125] The infrastructure orchestration layer consists of a Kubernetes (Kubernetes) container cluster management system and cloud resources. The Operator translates the desired state of CelebornCluster into resources such as ConfigMap, Service, RBAC (Role-Based Access Control), StatefulSet / Advanced StatefulSet, PLB / CLB, node taints, node tags, finalizers, and secondary network interface (ENI) resources. For nodes that need to access the user-side VPC, the Operator first applies scheduling restrictions to the node and automatically creates or mounts secondary network interfaces on the corresponding node. After confirming network readiness, the scheduling restrictions are lifted, and the node enters a schedulable state, allowing the Kubernetes container cluster management system to allocate cluster workloads to the node.
[0126] The data service layer consists of nodes, including a cluster management node (Celeborn Master), fixed nodes (fixed Workers), and elastic nodes (elastic Workers). The cluster management node (Celeborn Master) is responsible for metadata, resource coordination, and receiving scaling parameters. Fixed nodes (fixed Workers) carry basic capacity and stable storage and service capabilities. Elastic nodes (elastic Workers) carry burst traffic and short-term peaks, and achieve elastic scaling through the linkage of node pool and Celeborn parameters.
[0127] The above scheme establishes a pre-configuration gate for the auxiliary network interface card (NIC) configuration required for multi-VPC connectivity. Only after a node completes auxiliary NIC binding, route optimization, and cross-VPC reachability establishment is the network insecurity taint removed, allowing cluster workloads to be scheduled to that node. This ensures that the Celeborn data plane naturally runs on the network paths connected by the auxiliary NICs, enabling multi-VPC connectivity deployment based on auxiliary NICs. Furthermore, by incorporating CelebornCluster target state resolution, node network preparation, cluster resource orchestration, fixed worker and elastic worker collaborative scaling, dynamic configuration updates, and resource release into a single Operator tuning loop, the deployment complexity, coordination costs, and operational uncertainties associated with distributed manual operations and maintenance can be reduced.
[0128] In another scenario, the solution provides a cluster network scheduling and control method, applied to a cluster network scheduling and control system, for reference. Figure 4 , Figure 4 This is a modular schematic diagram of a cluster network scheduling and control system. (Example) Figure 3 As shown, the cluster network scheduling and control system includes a resource receiving and parsing module, a tuning control module, a network preparation module, a cluster resource orchestration module, a fixed node control module, a flexible node control module, a configuration consistency control module, and a resource release and reclamation module.
[0129] The resource receiving and parsing module is used to receive configuration information such as images, versions, networks, fixed nodes (fixed workers), elastic nodes (elastic workers), and running parameters defined in the custom resources of CelebornCluster, and parse them to obtain the target status of the Celeborn cluster.
[0130] The tuning control module is used to monitor changes in CelebornCluster resources and related node resources, and triggers the tuning process based on the difference between the expected state and the actual state, so that the cluster running state converges to the target state.
[0131] The network preparation module is used to perform auxiliary network resource preparation, mounting, route optimization and connectivity verification on the target node, and to control whether the node is allowed to carry Celeborn workloads based on the network preparation results.
[0132] The cluster resource orchestration module is used to generate or update ConfigMap, Service, permission resources, and workload resources corresponding to the cluster management node (Celeborn Master), fixed nodes (fixed Worker), and elastic nodes (elastic Worker) according to the target state of CelebornCluster.
[0133] The fixed node control module is used to control the expansion and contraction process of fixed workers, and to perform decommissioning control, data security exit checks and status recovery in the contraction scenario.
[0134] The elastic node control module is used to coordinate the relationship between the target upper limit of elastic workers, the upper limit of node pool capacity, and the actual running replicas, so as to achieve safety control during the elastic scaling process.
[0135] The configuration consistency control module is used to record the dynamic configuration status that has actually taken effect, and decides whether to perform configuration update based on the difference between the target configuration and the effective configuration during the tuning process, so as to ensure the idempotency of dynamic configuration update.
[0136] The resource release and recycling module is used to control the release process of external dependent resources, auxiliary network resources and related finalizers when the cluster is deleted or a node exits, so as to avoid resource leakage or blockage during the deletion phase.
[0137] It is understandable that the above modules can be implemented collaboratively by one or more controllers, or they can be implemented by breaking them down into multiple logical components according to their functions. The solution does not limit its specific deployment method; as long as the above-mentioned functional collaboration can be achieved, it should fall within the scope of protection.
[0138] Optionally, the workflow of the tuning control module specifically includes: The system retrieves the CelebornCluster object. If the object does not exist, the current tuning round ends. If the object exists but has entered the deletion phase, resource release and termination control logic is executed. If the object exists but has not entered the deletion phase, the network preparation module executes the node network preparation process. This process includes identifying the node pools involved in the relevant workloads, imposing or maintaining scheduling restrictions on nodes that are not yet ready for network deployment, and creating, mounting, and configuring auxiliary network resources. Once network connectivity requirements are met, scheduling restrictions are lifted, and scheduling is enabled.
[0139] After the node network preparation process is completed, the tuning control module tunes the ConfigMap, Services, permission resources, and load balancing-related basic resources. It tunes the workload resources corresponding to the cluster management node (Celeborn Master), fixed nodes (fixed Workers), and elastic nodes (elastic Workers). Once the cluster management node is ready, it executes the scaling coordination logic for fixed and elastic nodes. The cluster status is refreshed, recording the current network binding relationships, configuration effectiveness status, and cluster operation stage.
[0140] By coordinating the tuning control module and the network preparation module, a closed-loop sequence is formed, which first completes the node network preparation, then allows the workload to be deployed, and then performs capacity scaling. This avoids data path errors, cross-VPC connectivity issues, or data loss during scaling down caused by scheduling business instances first and then supplementing the network.
[0141] Optionally, the network preparation module executes the specific process of the node network preparation procedure, including: When the Operator identifies a node as belonging to the Celeborn target node pool and that node needs to undertake cross-VPC data communication tasks, it first imposes a scheduling constraint on the node indicating that the network is not ready. For example, the scheduling constraint can be implemented using node taints or other equivalent scheduling constraint mechanisms. The scheduling constraint can be used to characterize at least one of the following states: the first node has not completed the mounting of auxiliary network resources, the auxiliary network resources have been mounted but the routing rules have not yet been configured, or the target data plane link has not yet been verified as available.
[0142] During the period of this scheduling restriction, the container cluster management system (K8s) is prohibited from scheduling cluster workloads to the first node to avoid service instances running in an incorrect network environment that does not have the ability to communicate across private networks, which could lead to operational failures and other problems for the cluster workloads.
[0143] The network configuration corresponding to the node is determined based on the node attributes, and the auxiliary network resources are determined based on the network configuration. It is then checked whether the auxiliary network resources are configured in the cluster network. If they exist, the auxiliary network resources are directly attached to the node; otherwise, the auxiliary network resources are first created and associated with security control policies, and then attached to the node.
[0144] While waiting for auxiliary network resources to become available, the mapping between nodes and auxiliary network resources is recorded in the cluster state to form repeatable idempotent checkpoints. Through this mechanism, the originally decentralized cloud-side network resource operations are incorporated into the Operator control loop, making the node's ability to access the target VPC observable, retryable, and recoverable.
[0145] The network resources of the nodes are optimized by routing, which involves switching the communication path facing the data plane to the auxiliary network interface, while continuing to direct the communication paths of the cluster control plane and the preset reserved address range to the main network interface. Specifically, the auxiliary network interface corresponding to the auxiliary network resource is identified, its address is configured, and the communication path is switched to the auxiliary network interface. Persistent rules are configured to ensure that the routing relationship can be restored after the node restarts. This scheme enables connection to the Kubernetes control plane, image repository, and basic services through the main network interface, and connection to the target VPC or multiple business VPCs through the auxiliary network interface, handling the Celeborn data read and write paths. By using multiple network interfaces and routing splitting, the control flow and data flow of Celeborn are decoupled. The control flow remains in the original cluster network, while the data flow enters the target VPC via the auxiliary network, thus achieving cross-VPC deployment.
[0146] Once it's confirmed that the auxiliary network resource has been created or exists, successfully mounted to the node, and the node's operating system-side routing rules have been completed as preset, the node possesses the ability to communicate with the target VPC. The scheduling restrictions are then lifted, and the Kubernetes scheduler can schedule Celeborn workloads to this node. Since the node now has cross-VPC data plane capabilities, Celeborn Workers can communicate with the target VPC via the auxiliary network immediately after startup, avoiding the need for subsequent network corrections.
[0147] Optionally, when a new node is added to the target node pool, the controller detects the new node event and triggers the auxiliary network resource preparation process. When a node is deleted or enters the recycling phase, the controller triggers the dismantling and recycling of auxiliary network resources based on the node exit event. Node objects can indicate their Celeborn cluster affiliation through tags, associations, or other identifiers, thereby establishing a link between the lifecycle of network resources and the lifecycle of the cluster and nodes. In other words, instead of configuring the network in a one-time batch, the creation, mounting, and recycling of auxiliary network resources are embedded into the entire lifecycle of the node through a node event-driven mechanism. Consequently, whether it's the addition of new nodes due to the expansion of the fixed node pool or the exit of nodes due to the shrinking of the elastic node pool, the network resource side can be automatically processed accordingly.
[0148] The above scheme imposes a scheduling restriction on a node indicating network insecurity when its auxiliary network resources are not ready. The scheduling restriction is lifted after the auxiliary network interface card (NIC) is created or confirmed to exist, mounted, routed, and its connectivity verified. This ensures that Celeborn workloads run only on nodes with target data plane communication capabilities, reducing errors in network paths, cross-VPC unreachability, and post-remediation costs. Furthermore, by incorporating the creation, mounting, routed optimization, and dismantling and recycling of auxiliary network resources into a unified control mechanism, combined with state recording and termination control, the risks of external network resource leakage, state mismatch, and blocking during deletion can be reduced.
[0149] Optionally, the fixed node control module is used to control the scaling up and down of fixed nodes. The fixed nodes are used to provide basic stable capacity. Their scaling up and down is not only reflected in the change of the number of workload replicas, but also involves the consistency control between Celeborn cluster operating parameters, node carrying status and data security exit status.
[0150] refer to Figure 5 , Figure 5 This is a schematic diagram of the fixed node expansion process, such as... Figure 5 As shown, when scaling up a fixed number of nodes, the cluster state is switched to an intermediate scaling state. First, the Celeborn cluster updates its understanding of the target capacity of the fixed worker's runtime parameters, i.e., the cluster's custom resources are modified. The Operator compares configuration differences, i.e., tunes the target number of replicas for the fixed worker's workload. Only nodes that have completed network preparation and have had their scheduling restrictions lifted are allowed to carry the cluster workload. After the new worker starts up and joins the cluster, the stable capacity expansion is completed.
[0151] refer to Figure 6 , Figure 6 This is a schematic diagram of the fixed node scaling-down process, such as... Figure 6 As shown, when scaling down fixed nodes, the safe exit of nodes already containing Shuffle data must be prioritized. The Operator first identifies the fixed nodes to be retired and notifies Celeborn to switch them to retirement mode. Retirement mode imposes scheduling restrictions on the nodes to be retired, prohibiting new cluster workloads from being scheduled to them. Simultaneously, the Operator compares configuration differences, updates the target capacity parameters of the fixed nodes, but does not immediately determine the end of the scaling down. The nodes to be retired or the set of Workers to be observed are recorded, and the cluster state is switched to the retirement processing intermediate state. This round of tuning ends, awaiting subsequent checks.
[0152] The container cluster controller (Operator) continuously checks whether the actual number of replicas of the workload corresponding to the fixed node has converged to the target value, and whether there are still fixed worker instances or incomplete data exit tasks on the node to be exited. Only when both the replica status and node status meet the safety conditions, i.e., the cluster workload corresponding to the fixed node to be exited is completed, will the exit mode be restored to the normal operation mode. The effective configuration status will be updated, the set to be exited will be cleared, and the cluster state will be restored to a stable state.
[0153] The above approach allows Celeborn to detect which fixed workers are about to exit at the runtime layer, triggering data flushing, migration, or cessation of new writes, instead of Kubernetes directly deleting the corresponding workloads. This transforms fixed node scaling down from directly reducing replicas to a controlled process of first retiring and then shrinking, making it more suitable for longer data flushing cycles and higher data consistency requirements in cross-VPC data read / write scenarios, reducing the risk of data loss and task anomalies.
[0154] Optionally, the elastic node control module is used to control the expansion and contraction of elastic nodes. Elastic nodes are mainly designed for sudden traffic scenarios. Their control objective is not to maintain a fixed number of replicas, but to ensure that the application-side elastic parameters, the underlying node pool capacity, and the actual number of running replicas are coordinated.
[0155] refer to Figure 7 , Figure 7 A schematic diagram of the elastic node expansion process, such as... Figure 7 As shown, when the Operator detects that the target upper limit of elastic nodes is higher than the currently effective upper limit, it adopts the control order of first increasing the infrastructure capacity limit and then releasing the application-side elastic parameters, specifically including: First, increase the scalable capacity limit of the elastic node pool, then update the target capacity limit parameters of the elastic workers in the Celeborn cluster. This ensures that when Celeborn triggers more elastic workers based on business pressure, the underlying node pool has sufficient capacity to handle the load. If new elastic nodes are created in the node pool as a result, scheduling restrictions are first applied to the new elastic nodes. Only after the new elastic nodes complete network preparation and the scheduling restrictions are lifted can they handle the cluster workload, i.e., the cluster workload is allocated to the new elastic nodes. This avoids situations where only Celeborn elastic parameters are modified without a corresponding increase in infrastructure capacity, resulting in applications having the intention to scale but no available nodes to handle the load.
[0156] refer to Figure 8 , Figure 8 This is a schematic diagram of the elastic node scaling-down process, as shown below. Figure 8As shown, the Operator compares the new target upper limit for elastic nodes with the currently effective upper limit. When the Operator detects that the new target upper limit for elastic nodes is lower than the currently effective upper limit, it cannot immediately synchronize the reduction of the node pool capacity upper limit. Otherwise, it may cause the underlying infrastructure to reclaim nodes first, while the elastic workers on the nodes are still carrying shuffle data. When scaling down elastic nodes, the actual number of replicas or live replicas of the workload corresponding to the elastic node are read. The actual number of replicas is compared with the new target upper limit for elastic nodes. When the actual number of replicas is still higher than the new target upper limit, the reduction of the node pool capacity upper limit is temporarily suspended, and only the scaling down intention is recorded and checked again in subsequent tuning. When the actual number of replicas has dropped to a safe level, that is, the actual number of replicas is less than the new target upper limit for elastic nodes, the node pool capacity upper limit is then synchronously shrunk. The elastic node upper limit configuration in the Celeborn cluster is synchronously updated, the status of the effective configuration is updated, and idempotent convergence is completed.
[0157] By employing the above approach, prioritizing infrastructure capacity enhancement during expansion and checking the actual replica level before scaling down to meet safety requirements, the operational risks arising from a mismatch between application elasticity intent and underlying capacity can be mitigated. Furthermore, the principle of scaling down elastic workers—first reducing business instances and then infrastructure after ensuring a safe level—effectively prevents data loss or task failures caused by premature scaling down of the underlying infrastructure.
[0158] Optionally, when an elastic node is added to scale up, auxiliary network resources must be prepared before it can support cluster workloads. When an elastic node is scaled down and removed, the controller reclaims its auxiliary network resources based on the node's removal event. This combination of scheduling constraints and resource reclamation control avoids mismatches where workloads are scheduled but network resources are not ready, or nodes are deleted but network resources are not released.
[0159] Optionally, elastic nodes can be controlled to scale elastically by directly modifying the replicas of the Advanced StatefulSet by the Operator.
[0160] The resource release and reclamation module controls the release process of external dependent resources, auxiliary network resources, and related finalizers when a cluster is deleted or a node exits, to avoid resource leaks or blockages during the deletion phase. Specifically, it detects whether the CelebornCluster has entered the deletion phase and whether finalizers are still present on the objects. It performs explicit cleanup of external load balancers and their associated resources, rather than relying entirely on Kubernetes' default cascading deletion. After confirming that the relevant network resources have been unoccupied, it further releases the corresponding external network resources. After confirming that critical dependent resources have been safely cleaned up, it removes the finalizers, allowing Kubernetes to continue deleting the remaining associated resources.
[0161] The above approach involves explicitly cleaning up external dependent resources during cluster deletion and removing the termination control flag after the safe release conditions are met, in order to avoid resource release blocking or objects remaining in a terminated state for a long time due to reliance on Kubernetes' default cascading deletion.
[0162] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0163] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0164] Based on the same concept, corresponding to the cluster network scheduling and control method of any of the above embodiments, this application also provides a cluster network scheduling and control device.
[0165] refer to Figure 9 The device includes: The data acquisition module 901 is configured to acquire a first node and impose scheduling restrictions on the first node to prevent the container cluster management system from scheduling cluster workloads to the first node, wherein the first node is not assigned any workload. The status judgment module 902 is configured to obtain the network resource status corresponding to the first node, and in response to the network resource status being ready to complete, determine that the first node has the function of communicating across private networks through the auxiliary network. The restriction removal module 903 is configured to remove the scheduling restriction to allow the container cluster management system to allocate cluster workloads to the first node.
[0166] In one scenario, the state determination module 902 is specifically configured as follows: Obtain the node attributes corresponding to the first node, obtain the auxiliary network resources based on the node attributes, and attach the auxiliary network resources to the first node so that the first node has a communication interface for cross-private network communication. The routing settings for the network resources attached to the first node are configured. In response to the completion of the routing settings, the status of the network resources corresponding to the first node is determined to be ready.
[0167] In one scenario, the state determination module 902 is specifically configured as follows: Check whether the auxiliary network resources are configured in the cluster network; In response to the existence of the auxiliary network resource, the auxiliary network resource is attached to the first node; or, In response to the absence of the auxiliary network resource, the auxiliary network resource is created, and the target auxiliary network resource is attached to the first node.
[0168] In one scenario, the state determination module 902 is specifically configured as follows: Obtain the communication path of the network resource and the type identifier corresponding to the communication path; The first interface is obtained based on the type identifier, and the communication path is matched with the first interface to complete the routing settings.
[0169] In one scenario, the state determination module 902 is specifically configured as follows: In response to the type identifier being a user-side identifier, the auxiliary network interface corresponding to the auxiliary network resource is identified, the address of the auxiliary network interface is configured, and the auxiliary network interface is designated as the first interface; or... In response to the type identifier being a cluster-side identifier, the main network interface is designated as the first interface, wherein the main network interface is the cluster network interface.
[0170] In one embodiment, the device further includes a capacity expansion module, which is specifically configured as follows: Determine the target number of nodes corresponding to the cluster workload, and obtain the number of second nodes, wherein the second node is not assigned any load, and the first node and the second node belong to the same node pool; In response to the target number of nodes being greater than the sum of the number of the first node and the second node, a third node is added to the node pool, wherein the third node is not assigned any load. A scheduling restriction is imposed on the third node, and the scheduling restriction is lifted in response to the network resource status corresponding to the third node being ready.
[0171] In one scenario, the expansion module is specifically configured as follows: Obtain the node pool to which the first node belongs, and find the total number of all nodes in the node pool; In response to the total number of nodes being less than a preset threshold, a third node is added to the node pool; or... In response to the total number of nodes exceeding a preset threshold and receiving a node addition instruction, a third node is added to the node pool.
[0172] In one embodiment, the device further includes a reduction module, which is specifically configured as follows: In response to receiving a scaling-down instruction, the target number corresponding to the scaling-down instruction is obtained, wherein the scaling-down instruction is used to reduce the number of nodes in the node pool; Obtain the number of fourth nodes, and shrink the node pool according to the number of fourth nodes and the target number to obtain a new node pool, wherein the fourth nodes have been assigned load.
[0173] In one scenario, the reduction module is specifically configured as follows: In response to the target number being less than the number of the fourth node, obtain the node to be deleted; A scheduling restriction is imposed on the node to be deleted to prevent the container cluster management system from scheduling new cluster workloads to the node to be deleted; Once the cluster workload corresponding to the node to be deleted is identified as complete, the node to be deleted is deleted, and a new node pool is obtained.
[0174] In one scenario, the reduction module is specifically configured as follows: In response to the target number being greater than the number of the fourth node, obtain the node to be deleted; Delete the node to be deleted to obtain a new node pool.
[0175] In one embodiment, the device further includes a cluster deletion module, which is specifically configured to: A cluster deletion command is received, wherein the cluster deletion command is used to indicate the deletion of the cluster where the first node is located; Obtain all network resources configured with the cluster, and clear all network resources; Once all network resources have been cleared, delete the cluster.
[0176] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0177] The apparatus of the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0178] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the methods described in any of the above embodiments.
[0179] Figure 10 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0180] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0181] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0182] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0183] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0184] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0185] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0186] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0187] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the methods described in any of the above embodiments.
[0188] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0189] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0190] Based on the same concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0191] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0192] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be illustrated in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0193] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., DRAM) may be used with the embodiments discussed.
[0194] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A cluster network scheduling and control method, comprising: Obtain the first node and impose scheduling restrictions on the first node to prevent the container cluster management system from scheduling cluster workloads to the first node, wherein the first node is not assigned any workload; Obtain the network resource status corresponding to the first node, and in response to the network resource status being ready to complete, determine that the first node has the function of communicating across private networks through the auxiliary network. Remove the scheduling restrictions to allow the container cluster management system to allocate cluster workloads to the first node.
2. The method according to claim 1, wherein, The step of obtaining the network resource status corresponding to the first node includes: Obtain the node attributes corresponding to the first node, obtain the auxiliary network resources based on the node attributes, and attach the auxiliary network resources to the first node so that the first node has a communication interface for cross-private network communication. The routing settings for the network resources attached to the first node are configured. In response to the completion of the routing settings, the status of the network resources corresponding to the first node is determined to be ready.
3. The method according to claim 2, wherein, The step of attaching the auxiliary network resources to the first node includes: Check whether the auxiliary network resources are configured in the cluster network; In response to the existence of the auxiliary network resource, the auxiliary network resource is attached to the first node; or, In response to the absence of the auxiliary network resource, the auxiliary network resource is created and attached to the first node.
4. The method according to claim 2, wherein, The routing configuration for the network resources mounted on the first node includes: Obtain the communication path of the network resource and the type identifier corresponding to the communication path; The first interface is obtained based on the type identifier, and the communication path is matched with the first interface to complete the routing settings.
5. The method according to claim 2, wherein, The step of obtaining the first interface based on the type identifier includes: In response to the type identifier being a user-side identifier, the auxiliary network interface corresponding to the auxiliary network resource is identified, the address of the auxiliary network interface is configured, and the auxiliary network interface is designated as the first interface; or... In response to the type identifier being a cluster-side identifier, the main network interface is designated as the first interface, wherein the main network interface is the cluster network interface.
6. The method according to claim 1, wherein, Also includes: Determine the target number of nodes corresponding to the cluster workload, and obtain the number of second nodes, wherein the second node is not assigned any load, and the first node and the second node belong to the same node pool; In response to the target number of nodes being greater than the sum of the number of the first node and the second node, a third node is added to the node pool, wherein the third node is not assigned any load. A scheduling restriction is imposed on the third node, and the scheduling restriction is lifted in response to the network resource status corresponding to the third node being ready.
7. The method according to claim 6, wherein, Adding a third node to the node pool includes: Obtain the node pool to which the first node belongs, and find the total number of all nodes in the node pool; In response to the total number of nodes being less than a preset threshold, a third node is added to the node pool; or... In response to the total number of nodes exceeding a preset threshold and receiving a node addition instruction, a third node is added to the node pool.
8. The method according to claim 1, wherein, Also includes: In response to receiving a scaling-down instruction, the target number corresponding to the scaling-down instruction is obtained, wherein the scaling-down instruction is used to reduce the number of nodes in the node pool; Obtain the number of fourth nodes, and shrink the node pool according to the number of fourth nodes and the target number to obtain a new node pool, wherein the fourth nodes have been assigned load.
9. The method according to claim 8, wherein, The step of scaling down the node pool based on the number of the fourth node and the target number to obtain a new node pool includes: In response to the target number being less than the number of the fourth node, obtain the node to be deleted; A scheduling restriction is imposed on the node to be deleted to prevent the container cluster management system from scheduling new cluster workloads to the node to be deleted; Once the cluster workload corresponding to the node to be deleted is identified as complete, the node to be deleted is deleted, and a new node pool is obtained.
10. The method according to claim 8, wherein, The step of scaling down the node pool based on the number of the fourth node and the target number to obtain a new node pool includes: In response to the target number being greater than the number of the fourth node, obtain the node to be deleted; Delete the node to be deleted to obtain a new node pool.
11. The method according to claim 1, wherein, Also includes: A cluster deletion command is received, wherein the cluster deletion command is used to indicate the deletion of the cluster where the first node is located; Obtain all network resources configured with the cluster, and clear all network resources; Once all network resources have been cleared, delete the cluster.
12. A cluster network scheduling and control device, comprising: The data acquisition module is configured to acquire the first node and impose scheduling restrictions on the first node to prevent the container cluster management system from scheduling cluster workloads to the first node, wherein the first node is not assigned any workload. The status determination module is configured to obtain the network resource status corresponding to the first node, and in response to the network resource status being ready to complete, determine that the first node has the function of communicating across private networks through the auxiliary network. The restriction removal module is configured to remove the scheduling restrictions to allow the container cluster management system to allocate cluster workloads to the first node.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method as described in any one of claims 1 to 11.
15. A computer program product comprising computer program instructions, wherein, When the computer program instructions are executed on a computer, the computer causes the computer to perform the method as described in any one of claims 1 to 11.