Multi-cluster hybrid deployment system of heterogeneous models

By building a multi-cluster hybrid deployment system with heterogeneous machines and utilizing the machine characteristics for grouping and dynamic resource adjustment, the problems of insufficient resource utilization and load balancing in the cluster are solved, and the overall operation efficiency and resource utilization of the cluster are improved.

CN120407213BActive Publication Date: 2025-09-26HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510918619.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-26
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

In existing cluster management, it is difficult for services in the cluster to fully utilize the advantageous resources of specific nodes, and it is difficult to achieve the load balancing goal between online and offline co-location in a single cluster, resulting in reduced overall cluster operation efficiency.

Method used

Build a multi-cluster hybrid deployment system for heterogeneous machines. By constructing a multi-dimensional resource model of cluster-machine group-machine-target service, group machines according to their own characteristics to form machine groups with differentiated capabilities, and dynamically adjust resources within and outside the cluster to achieve load balancing scheduling.

Benefits of technology

It improves the overall machine resource utilization of the cluster, ensures the stability and flexibility of the service, reduces resource waste, and improves the efficiency of task scheduling and system operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407213B_ABST
    Figure CN120407213B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-cluster hybrid deployment system for heterogeneous machines. This relates to the fields of computer technology, particularly cluster management and heterogeneous resource management. A specific implementation scheme includes: multiple clusters built based on heterogeneous machines; each cluster includes multiple machine groups; each machine group corresponds to its own first target service; for each machine group, the machine group is grouped based on the hardware feature requirements of the corresponding first target service; within the same cluster, resources are dynamically adjusted between different machine groups within the cluster based on the resource requirements of the second target service supported by the cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as cluster management and heterogeneous resource management. Background Art

[0002] In the field of computer technology, as services continue to expand and diversify, the demand for computing resources is becoming increasingly complex. Multi-cluster architectures have emerged to address this challenge. By dividing different computing resources into multiple clusters, each responsible for its own workload, they enable categorized resource management and efficient utilization. Summary of the Invention

[0003] The present disclosure provides a multi-cluster hybrid deployment system of heterogeneous machines to solve or alleviate one or more technical problems in the prior art.

[0004] In a first aspect, the present disclosure provides a multi-cluster hybrid deployment system for heterogeneous machines, including:

[0005] Multiple clusters built on heterogeneous models;

[0006] Each cluster includes multiple machine groups; each machine group corresponds to its own first target service; for each machine group, the machine group is grouped based on the hardware feature requirements of the corresponding first target service;

[0007] Within the same cluster, resources are dynamically adjusted between different machine groups within the cluster based on the resource requirements of the second target service supported by the cluster.

[0008] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0010] Figure 1 is an exemplary architecture diagram of a multi-cluster hybrid deployment system of heterogeneous machines according to the first embodiment of the present disclosure;

[0011] Figure 2 is a schematic diagram of a process for dynamically adjusting resources between different machine groups in a cluster according to the second embodiment of the present disclosure;

[0012] Figure 3is a schematic diagram of the structural flow of adjusting machine resources between different machine groups according to the third embodiment of the present disclosure;

[0013] Figure 4 is a schematic diagram of a flow chart of load balancing between clusters according to the fourth embodiment of the present disclosure;

[0014] Figure 5 is a schematic diagram of a process for configuring a cluster for a newly added service according to the fifth embodiment of the present disclosure;

[0015] Figure 6 2. It is a schematic diagram of execution scheduling of a multi-cluster hybrid deployment system of heterogeneous machines according to the sixth embodiment of the present disclosure;

[0016] Figure 7 2 is a schematic diagram of the electronic device structure of a multi-cluster hybrid deployment system of heterogeneous models according to the seventh embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] The present disclosure will be described in further detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0018] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, circuits, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.

[0019] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0020] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order between multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.

[0021] Cluster scheduling systems play a crucial role in computer technology. However, existing cluster colocation solutions have numerous limitations: First, services within the cluster struggle to fully utilize the resources of specific nodes; second, load balancing between online and offline colocation within a single cluster is difficult to achieve. These two factors can reduce overall cluster operational efficiency.

[0022] In light of this, and to address at least one of the aforementioned issues, embodiments of the present disclosure provide a system for the hybrid deployment of heterogeneous machines across multiple clusters. This system constructs a multi-dimensional resource model encompassing clusters, machine groups, machines, and target services. This system effectively groups machines based on their characteristics, forming machine groups with differentiated capabilities. Based on these groupings, load balancing is implemented across clusters and within a single cluster, thereby improving overall cluster resource utilization.

[0023] like Figure 1 The figure shows an exemplary architecture diagram of a multi-cluster hybrid deployment system of heterogeneous machines provided by the present disclosure, including:

[0024] Multiple clusters 101 are built based on heterogeneous machine models. Heterogeneous machine models refer to machines with different hardware characteristics, such as operating systems, graphics card types, disk types, and network card types. Clusters built based on heterogeneous machine models are a collection of computing resources consisting of multiple machines interconnected via a network. These machines work together to provide unified service capabilities.

[0025] Each cluster includes multiple machine groups 102. Each machine group corresponds to a respective first target service. Each machine group is grouped based on the hardware feature requirements of the corresponding first target service. Each cluster contains multiple machine groups, each grouped based on the hardware feature requirements of the corresponding first target service. A first target service can be understood as a specific function or application service. Different first target services may have different requirements for machine hardware features.

[0026] During implementation, each machine can report its own feature information to the monitoring node of the cluster through its own agent. Each cluster can receive the feature information reported by each machine through its own monitoring node, such as the resource manager. The cluster monitoring center divides the machines that meet the hardware features required by the first target service into a machine group based on the requirements of the first target service for the machine hardware features. Figure 1As shown in the figure, in cluster A, there are machine groups A1 and A2. The first target service corresponding to machine group A1 may only be executed in a Linux (a computer system) environment with high-performance graphics cards and fast read / write hard drives. Therefore, machines with these characteristics are assigned to machine group A1. Similarly, the first target service corresponding to machine group A2 may only be executed in a Windows (a computer system) environment with high-performance graphics cards and fast read / write hard drives. Therefore, machines with these characteristics are assigned to machine group A2.

[0027] In this system, within the same cluster, resources are dynamically adjusted between different machine groups within the cluster based on the resource requirements of the second target service supported by the cluster.

[0028] Secondary target services are services that require coordination and management within the same cluster. Their resource requirements may fluctuate over time, depending on factors such as business load. Within a cluster, the services required to support them can also change dynamically. Therefore, based on the resource requirements of secondary target services, resources can be adjusted across different machine groups within the same cluster to adjust the machines within each group, thereby ensuring efficient and stable operation of the entire cluster.

[0029] In the disclosed embodiments, within a cluster architecture built on heterogeneous machine models, multiple clusters can be constructed, each catering to different service requirements. When a new service is needed, the system can quickly find and allocate appropriate machine resources within existing clusters and machine groups based on the resource requirements of the service, thereby supporting flexible service expansion. Each cluster includes multiple machine groups. By grouping machines that meet the hardware requirements of a first target service into a machine group, the system ensures that the first target service receives a stable supply of corresponding machine resources, mitigating issues such as wasted machine resources or insufficient performance due to over- or under-configured hardware. Furthermore, machine groups can fully utilize the advantageous resources of specific nodes. Within the same cluster, resources can be dynamically adjusted across different machine groups based on the resource requirements of a second target service supported by the cluster. When load imbalances occur between different machine groups, machine resources can be real-timely allocated from machine groups with relatively abundant resources to those in need. This mitigates the issue of service performance degradation due to insufficient machine resources in some machine groups within the same cluster, prevents idle machine resources within certain machine groups, and further optimizes the overall utilization efficiency of machine resources in the system.

[0030] In the disclosed embodiments, a cluster typically deploys a mix of online and offline tasks to achieve efficient utilization of cluster machine resources. Online tasks are highly real-time and, in emergencies, may run out of available machine resources. Offline tasks, on the other hand, have relatively loose requirements for machine resource quality and can be run flexibly when machine resources are available.

[0031] Based on this, while the total amount of cluster machine resources is fixed, the allocation strategy can be dynamically adjusted based on the machine resource requirements of online and offline tasks. For example, when the online task load surges, idle machine resources of offline tasks can be temporarily used to ensure the normal operation of the online task.

[0032] To improve the security and stability of machine resource migration and reduce disruption to existing services caused by switching machines between groups, we can set up idle resource zones to achieve smoother resource transitions. This effectively prevents service interruptions and performance fluctuations caused by machines suddenly joining or leaving a group. This allows for dynamic resource adjustments between different machine groups within a cluster.

[0033] During implementation, in the disclosed embodiments, in a multi-cluster hybrid deployment system of heterogeneous machines, a monitoring node may be set up in each cluster, and a general dispatching center may be set up in the entire system. The monitoring nodes are used to manage resource allocation issues within their clusters. The general dispatching center can manage task allocation issues between clusters. It is understood that the monitoring nodes can also be used as a medium for transmitting information within the corresponding cluster, passing relevant information to the general dispatching center, which will manage resource allocation within the cluster and task allocation issues between clusters.

[0034] During implementation, within the same cluster, by maintaining idle resource areas, resources can be dynamically adjusted between different machine groups within the cluster. The specific implementation method can be as follows: Figure 2 As shown, it can be implemented by the monitoring node of the cluster, including:

[0035] S201: Maintain an idle resource area, where the idle resource area includes idle machines released from a first target machine group.

[0036] The free resource zone is an area within a cluster dedicated to storing and managing idle machine resources. In a heterogeneous, multi-cluster hybrid deployment system, it manages temporarily unused machine resources within the cluster so that they can be quickly allocated to services that require them when needed.

[0037] During implementation, the implementation process of maintaining the idle resource area by releasing the idle machines from the first target machine group can be as follows: Figure 2 Shown, including:

[0038] S2011, based on the preset task eviction strategy, filters out the target tasks in the cluster.

[0039] In this disclosed embodiment, the second target service is an online task whose resource demand growth exceeds a threshold, and the target task is an offline task. This threshold is used to indicate a sudden increase in online tasks, allowing the cluster to respond to emergencies. Deploying online and offline tasks within the same cluster allows the system to dynamically adjust resource allocation strategies based on real-time resource usage and business needs, improving the efficiency and flexibility of cluster management.

[0040] A task eviction policy is a pre-defined rule or method. Based on this policy, the system identifies target tasks within the cluster that meet the eviction criteria. For example, it evicts the target tasks with the least impact on the system.

[0041] During implementation, the preset task eviction strategy may determine the eviction score of each candidate task in the cluster and filter out the target task based on the eviction score of each candidate task.

[0042] The eviction score is a quantitative metric used to evaluate the suitability of candidate tasks within a cluster for eviction. It helps the system quickly and objectively identify candidate tasks that can be paused or terminated to free up idle machines for free resources.

[0043] The eviction score is determined based on at least one of the following parameters and has a positive correlation with each of the following parameters: processor utilization of the candidate task, storage resource utilization of the candidate task, and execution time of the candidate task.

[0044] For example, the expulsion score can be described by formula (1):

[0045]

[0046] In formula (1), EvictStore represents the eviction score of each candidate task; CPU_util / 100 represents the processor utilization of each candidate task, that is, the CPU utilization of each candidate task in the corresponding machine group; Mem_util / 100 represents the storage resource utilization of each candidate task, which may include the memory utilization of each candidate task in the corresponding machine group; TaskAge / 24h represents the execution time of each candidate task, in hours, with a maximum value of 24 hours, that is, the running time of the target task does not exceed 24 hours.

[0047] Candidate tasks are sorted by eviction score, with the candidate task with the lowest eviction score selected as the target task for eviction. A lower score indicates lower processor and storage resource utilization within the machine group, and a shorter execution time. Evicting this task effectively frees up idle machine resources with minimal impact on other tasks.

[0048] In the disclosed embodiment, an eviction score is calculated by calculating at least one of the processor utilization, storage utilization, and task execution duration of each candidate task in the cluster. Determining which candidate tasks are selected as target tasks for eviction based on the eviction score allows for the release of machine resources at minimal cost, avoiding service interruptions caused by the eviction of high-load or critical tasks. This enables refined management of on-demand release of machine resources, improving resource utilization and task scheduling efficiency in heterogeneous cluster environments.

[0049] S2012: Evict the target task from the target machine running the target task in the first target machine group to release resources of the target machine and obtain an idle machine.

[0050] After selecting the target tasks based on the eviction scores of each candidate task, the system evicts the target tasks from the target machines running them. The system also changes the status of the target machines from busy to idle, indicating that they are not currently running any tasks and are available as idle resources for the system to call and allocate.

[0051] Among them, evicting the target task may involve pausing the execution of the task, migrating the task to another machine, or directly terminating the task, which depends on the specific strategy and the nature of the task, and is not limited in this embodiment of the present disclosure.

[0052] S2013: Set the idle machines to the idle resource area.

[0053] Idle machines are added to the idle resource zone. In this way, when other services or tasks in the same cluster need to expand resources, they can quickly obtain available machine resources from this idle resource zone.

[0054] In this disclosed embodiment, a preset task eviction policy is used to screen out target tasks with minimal impact within the cluster. These tasks are then evicted from the target machines running them in the first target machine group, freeing up resources on the target machines and converting them into idle machines. Furthermore, the idle machines are consolidated into an idle resource zone. When other services in the cluster (such as the second target service) require more resources, these idle machines can be directly acquired from the idle resource zone, thus preventing resource idleness and improving resource utilization efficiency across the entire cluster.

[0055] S202 : For the second target service, dynamically adjust resources among different machine groups through idle resource areas based on resource expansion capacity required by the second target service.

[0056] During implementation, the system will monitor the resource usage of the second target service in real time and assess its current resource requirements.

[0057] When the second target machine group where the second target service is located needs to be expanded, based on the resource expansion capacity required by the second target service, idle machines that match the hardware feature requirements of the second target service are obtained from the idle resource area; and the idle machines are allocated to the second target machine group.

[0058] Specifically, the system selects eligible idle machines from the free resource pool based on the hardware requirements of the second target service and assigns them to the second target machine group to be expanded. This increases the resource capacity of the second target machine group, enabling it to more efficiently handle the increased load. Furthermore, the system continuously monitors the resource usage of the second target machine group, allowing it to further adjust resource allocation strategies as necessary.

[0059] In the disclosed embodiments, if the second target machine group hosting the second target service requires capacity expansion, based on the resource expansion capacity required by the second target service, idle machines matching the hardware requirements of the second target service are retrieved from the idle resource zone for capacity expansion. This effectively prevents system crashes or severe performance degradation caused by resource overload, further ensuring the continuity of the second target service. By dynamically allocating idle machines to the required machine groups based on the actual load of different machine groups, this achieves on-demand allocation of machine resources and improves overall resource utilization.

[0060] In the embodiment of the present disclosure, when the load of the second target service increases, the system (such as the monitoring node of the cluster) will calculate the required resource expansion capacity. The calculation of the required resource expansion capacity of the second target service can be achieved based on the following steps:

[0061] Step A1: When the second target service is an online task, determine the current amount of idle resources in the total resources allocated for the offline task and an online indicator; the online indicator is determined based on the resource quota allocated for the online task, and the online indicator is less than the resource quota;

[0062] The current idle resources in the total resources allocated to offline tasks refer to the unused portion of the allocated resources for offline tasks. The online indicator is the amount of resources allocated to online tasks to ensure their basic operation.

[0063] During implementation, online metrics can be planned and constrained through a quota (system quota) mechanism. Minimum and maximum quotas are set to limit the resource usage of online tasks. Quota usage can also be monitored, with adjustments made when necessary to optimize resource allocation. This allows for the rational allocation and management of limited resources, preventing overconsumption and ensuring stable and efficient system operation.

[0064] The minimum quota specifies the minimum resources required for online tasks to run normally, preventing them from failing due to insufficient resources. When configuring the minimum quota for online tasks, ensure that the total minimum quota for each task does not exceed the total physical resources available in the cluster. This ensures that each service receives the minimum resource requirements it needs without exceeding the overall cluster capacity.

[0065] The maximum quota is the upper limit of resources that can be used by online tasks, preventing online tasks from excessively occupying resources and affecting other services.

[0066] In the disclosed embodiment, the maximum quota is selected for online metrics. This maximum quota allows online tasks to borrow more idle resources from offline tasks. This allows online tasks to apply for more resources in a short period of time, thereby better coping with sudden high loads.

[0067] Step A2: Determine the resource expansion capacity required for the second target service based on the current idle resource amount and the online indicator.

[0068] During implementation, the resource expansion capacity required for the second target service can be expressed by formula (2):

[0069] Resource expansion capacity = min (current idle resource amount, online index × γ) (2)

[0070] Among them, γ is the maximum borrowing ratio and can be set to 50%.

[0071] In the disclosed embodiment, by determining the resource expansion capacity required by the second target service, the system can determine the maximum amount of offline resources that can be borrowed by the online task. This ensures that the resources of the offline task are fully utilized without affecting the basic operation of the offline task, thereby improving the resource utilization of the entire system.

[0072] In an embodiment of the present disclosure, in order to reduce the problem of offline task resources being unable to operate normally for a long time due to being borrowed for a long time, when the second target service is an online service and an idle machine is obtained by expelling the offline service, the borrowing time of the idle machine by the second target service can be set as a time threshold; and the second target service can be controlled to return a preset proportion of the idle machines borrowed from the offline task at every specified period.

[0073] When an online task borrows machines from an offline task, the system sets a time threshold for the borrowed machines. This threshold specifies the maximum preset duration that the online service can borrow these idle machines. Furthermore, the system sets a specified time period, requiring the online service to return the idle machines borrowed from the offline service at a preset ratio at the end of each period. This ensures that the offline service can gradually resume resource use. For example, if the borrowing period is ≤ 2 hours, 20% of the borrowed resources must be returned every hour.

[0074] In the disclosed embodiments, when idle machines are obtained by evicting offline tasks, the resource requirements of online and offline tasks within a single cluster can be balanced by setting a threshold for the duration of time that the idle machines are borrowed by a second target service. This threshold controls the second target service's return of a preset proportion of the idle machines borrowed from offline tasks at specified intervals. By setting the borrowing duration and controlling the return period and proportion, the system ensures that online services receive sufficient resources when needed while avoiding the long-term occupation of offline service resources, thereby improving the resource utilization efficiency and flexibility of the entire system.

[0075] In summary, by maintaining idle resource areas, resources can be dynamically adjusted between different machine groups within the cluster. Figure 3 For example, when the task load of machine group A surges, idle machine 1 resources that can meet the hardware feature requirements of its tasks can be obtained from the idle resource area to ensure the normal operation of the tasks of machine group A.

[0076] During implementation, the task on machine 1 in machine group B can be evicted using a pre-set task eviction policy. This releases the resources of machine 1 and places them in the idle resource zone. Machine group A then acquires machine 1 from the idle resource zone to meet the surge in task load.

[0077] In the disclosed embodiments, by maintaining an idle resource zone and using it as a resource buffer pool, the required machine resources can be quickly provided to machine groups requiring capacity expansion, promptly responding to changes in the load of the secondary target service and meeting sudden demands for the secondary target service. For the secondary target service, dynamic resource adjustments are made between different machine groups using the idle resource zone based on the resource expansion capacity required by the secondary target service. This ensures that each machine group has sufficient machine resources to respond to load changes, reducing the risk of secondary target service interruptions due to insufficient machine resources and improving resource utilization efficiency.

[0078] In the embodiment of the present disclosure, in addition to the dynamic resource adjustment between the machine groups in the same cluster as described above, in the case of load imbalance between the clusters, the following strategies can also be used to achieve resource optimization scheduling, specifically: Figure 4 As shown:

[0079] S401 : Determine a task migration amount in the first cluster when it is determined that the difference between the load of the first cluster and the load of the second cluster meets a cross-cluster task adjustment condition; the first cluster and the second cluster support the same task type.

[0080] During implementation, for example, cluster 1 is a high-load cluster and cluster 2 is a low-load cluster. The system (e.g., the central dispatch center) sets a threshold to determine whether the load difference between the two clusters meets the conditions for cross-cluster task adjustment. Once the conditions are determined to be met, the task migration volume for cluster 1 is calculated, specifically the number of tasks to be migrated from the high-loaded first cluster to the low-loaded second cluster.

[0081] In the embodiment of the present disclosure, determining the task migration amount in the first cluster can be achieved based on the following steps:

[0082] Step A1: determining a first candidate migration amount based on the load of the first cluster; and determining a second candidate migration amount based on the load of the second cluster;

[0083] Step A2: Filtering out the task migration amount from the first candidate migration amount and the second candidate migration amount.

[0084] During implementation, the first candidate migration amount is determined based on the load of the first cluster, which can be obtained by subtracting the number of services in the first cluster that are critical to business operations and should not be easily migrated from the number of all services currently running in the first cluster; similarly, the second candidate migration amount is determined based on the load of the second cluster, which can also be obtained by subtracting the number of services in the second cluster that are critical to business operations and should not be easily migrated from the number of all services currently running in the second cluster.

[0085] Then, the minimum value can be selected from the first candidate migration amount and the second candidate migration amount as the task migration amount in the first cluster. Of course, a migration amount can be randomly selected from the range represented by the first candidate migration amount and the second candidate migration amount to obtain the task migration amount.

[0086] In the disclosed embodiment, by calculating the candidate migration amounts of the first cluster and the second cluster and selecting the minimum value as the final task migration amount, the excess load in the first cluster is transferred to the second cluster, so that the loads of the two clusters tend to be balanced, while avoiding overloading of the second cluster due to migrating too many tasks.

[0087] It should be understood that task migration between clusters also needs to ensure that the machine resource requirements of the tasks in the first cluster match the hardware characteristics of the machines in the second cluster.

[0088] S402: Filter out tasks to be migrated with a large task migration amount from the first cluster.

[0089] The selected tasks for migration are those in the first cluster that do not have the core task flag. By migrating these tasks, we ensure that these core tasks continue to run stably in the first cluster without being disrupted by the migration process. This prevents service interruptions or performance degradation caused by migrating core tasks, ensuring service continuity and reliability.

[0090] S403: Migrate the task to be migrated to a machine group in the second cluster that supports the task to be migrated for execution.

[0091] In the embodiment of the present disclosure, by migrating tasks in the first cluster to the second cluster with a lighter load, computing resources can be distributed more evenly, avoiding overloading of some clusters while idling resources in other clusters, thereby improving overall resource utilization.

[0092] In the disclosed embodiments, as business scale and complexity increase, resource management strategies must be continuously optimized to meet diverse needs. When faced with new service deployment requirements, new service deployment can be implemented based on the strategy shown in Figure 5. Specifically, the following operations can be performed by the central dispatch center:

[0093] S501: Determine the allocation weight of each machine based on the load of each cluster.

[0094] The system monitors the load of each cluster in real time, including resource utilization, task execution status, etc. For example, a cluster may have multiple machine groups running various tasks, and the system will count indicators such as CPU utilization and memory usage of each cluster to determine its load situation.

[0095] The weight assigned to each cluster is determined based on its load. A cluster with a lighter load may be assigned a higher weight, indicating that it can accept more tasks; while a cluster with a heavier load may be assigned a lower weight, indicating that fewer tasks need to be assigned to avoid overload.

[0096] During implementation, for each cluster, the allocation weight of the cluster is determined based on at least one of the following parameters of the cluster: idle processor resources of the cluster, idle storage resources of the cluster, idle processor resources of all clusters, and idle storage resources of all clusters.

[0097] Among them, the allocation weights are positively correlated with the amount of idle processor resources and idle storage resources of the cluster respectively;

[0098] The allocation weights are inversely correlated with the amount of idle processor resources of all clusters and the amount of idle storage resources of all clusters, respectively.

[0099] In the embodiment of the present disclosure, the allocation weight of the cluster can be expressed by formula (3):

[0100]

[0101] In formula (3), represents the allocation weight of cluster i; CPU_free represents the amount of idle processor resources in cluster i, is the weight coefficient of the idle processor resources of cluster i; Mem_free represents the amount of free storage resources in cluster i, is the weight coefficient of the free storage resources of cluster i; It is the sum of the idle processor resources of all clusters and the idle storage resources of all clusters.

[0102] The higher the weight ratio, the greater the probability that the cluster will be selected and assigned to a new instance when a new task needs to be assigned, thereby obtaining more resources to run the task.

[0103] In the embodiment of the present disclosure, by calculating the ratio of the idle processor resources and the idle storage resources of the cluster to the idle processor resources and the idle storage resources of all clusters, the allocation weight of the cluster is determined to optimize resource matching, so that the allocation weight of the cluster can simultaneously reflect the available status of multiple hardware resources, avoid scheduling deviations caused by a single dimension, and improve the resource utilization and service stability of the overall system.

[0104] S502 : When a third target service is newly added, a cluster that meets the hardware feature requirements of the third target service is selected as a target cluster based on the allocation weight.

[0105] When adding a third-party service, the system first matches the hardware characteristics of the machines required by the third-party service. Then, based on the allocation weights of each cluster, it determines the target cluster that can host the service. For example, if the hardware resources required for the new third-party service are a Linux system and a hard disk drive (HDD), the cluster with the highest allocation weight is selected as the target cluster to ensure reasonable resource allocation and efficient utilization.

[0106] S503: Allocate the third target service to the target cluster, and create a corresponding machine group for the third target service.

[0107] Deploy the third target service to the selected target cluster. In the target cluster, create a dedicated machine group for the third target service.

[0108] During implementation, based on the aforementioned information, machines within the cluster that meet the hardware requirements of the third target service can be migrated to the idle resource zone. The required machine resources are then obtained from the idle resource zone and combined to form a machine group corresponding to the third target service.

[0109] In the embodiment of the present disclosure, the allocation weight of each machine is determined based on the load situation of each cluster. The weight can reflect the current resource idleness or carrying capacity of the cluster. Based on the allocation weight, the cluster that adapts to the hardware feature requirements is selected as the target cluster to prevent the new service from being allocated to the already highly loaded cluster, thereby avoiding service performance degradation or failure due to insufficient resources.

[0110] In summary, the present disclosure implements the overall scheduling process of providing a multi-cluster hybrid deployment system of heterogeneous models as follows: Figure 6 As shown:

[0111] During implementation, a general dispatching center can be set up in a multi-cluster hybrid deployment system of heterogeneous models to manage load balancing issues between clusters by monitoring the load conditions of each cluster in real time.

[0112] A monitoring node can also be set up in each cluster to receive hardware characteristics reported by each machine through its own agent (not shown in the figure). The machines are then divided into different machine groups based on the hardware characteristics required by each target service. Furthermore, the monitoring node in each cluster can monitor the load balancing between machine groups within the cluster in real time to manage load balancing issues within the cluster.

[0113] The monitoring node can also be used as a medium for transmitting information within the corresponding cluster, and transmit relevant information to the general dispatching center, which will manage the resource allocation within the cluster and the task allocation between clusters.

[0114] It is understandable that in a multi-cluster hybrid deployment system of heterogeneous models, the overall dispatching center and the monitoring nodes in each cluster can be flexibly deployed on any electronic device in the system.

[0115] Figure 7 FIG. 1 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 7 As shown, the electronic device includes: a memory 710 and a processor 720. The memory 710 stores a computer program that can be executed on the processor 720. The number of memory 710 and processor 720 can be one or more. The memory 710 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device performs the method provided by the above method embodiment. The electronic device may also include: a communication interface 730 for communicating with external devices and performing data exchange.

[0116] If the memory 710, processor 720, and communication interface 730 are implemented independently, the memory 710, processor 720, and communication interface 730 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0117] Optionally, in a specific implementation, if the memory 710, the processor 720 and the communication interface 730 are integrated on a chip, the memory 710, the processor 720 and the communication interface 730 can communicate with each other through an internal interface.

[0118] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0119] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).

[0120] In the description of the embodiments of the present disclosure, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0121] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means or. For example, A / B can mean A or B. "And / or" in this document is only a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0122] In the description of the embodiments of the present disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.

[0123] The above description is merely an exemplary embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.

Claims

1. A multi-cluster hybrid deployment system with heterogeneous models, including: Multiple clusters built on heterogeneous models; Each cluster includes multiple machine groups; Each machine group corresponds to its own first target service; For each machine group, the machine group is grouped based on hardware feature requirements of the corresponding first target service; Within the same cluster, dynamically adjusting resources between different machine groups within the cluster based on the resource requirements of the second target service supported by the cluster, including: Maintaining an idle resource area, wherein the idle resource area includes idle machines released from the first target machine group; For the second target service, dynamically adjust resources among different machine groups through the idle resource area based on resource expansion capacity required by the second target service; The maintaining of the idle resource area includes: Based on a preset task eviction strategy, target tasks in the cluster are screened out, including determining an eviction score for each candidate task in the cluster; the eviction score is positively correlated with a processor utilization rate and a storage resource utilization rate of the candidate task, and negatively correlated with an execution time of the candidate task; and the target tasks are screened out in ascending order of the eviction scores; evicting the target task from the target machine running the target task in the first target machine group to release resources of the target machine and obtain the idle machine; placing the idle machine into the idle resource area; The second target service is an online task whose resource demand growth exceeds a threshold, and the target task is an offline task; Also includes: If the second target service is an online service and the idle machine is obtained by expelling an offline service, setting a time period for which the idle machine is borrowed by the second target service as a time period threshold; The second target service is controlled to return a preset proportion of idle machines borrowed from the offline task at specified intervals, so as to ensure that the offline task gradually resumes resource usage.

2. The system according to claim 1, wherein: The dynamic resource adjustment between different machine groups using the idle resource area includes: When the second target machine group where the second target service is located needs to be expanded, based on the resource expansion capacity required by the second target service, an idle machine that matches the hardware feature requirements of the second target service is obtained from the idle resource area; The idle machines are divided into the second target machine group.

3. The system according to claim 1, wherein: Determine the resource expansion capacity required for the second target service based on the following method: In a case where the second target service is an online task, determining a current amount of idle resources in the total resources allocated for the offline task and an online index; the online index is determined based on a resource quota allocated for the online task, and the online index is less than the resource quota; Based on the current idle resource amount and the online indicator, the resource expansion capacity required for the second target service is determined.

4. The system according to claim 1, further comprising: Determine the allocation weight of each machine based on the load of each cluster; In the case of adding a third target service, a cluster that meets the hardware feature requirements of the third target service is selected as a target cluster based on the allocation weight; Allocate the third target service to the target cluster, and create a corresponding machine group for the third target service.

5. The system according to claim 4, wherein: The allocation weight of each machine is determined based on the load of each cluster, including: For each cluster, determining an allocation weight of the cluster based on at least one of the following parameters of the cluster: an amount of idle processor resources of the cluster, an amount of idle memory resources of the cluster, an amount of idle processor resources of all clusters, and an amount of idle memory resources of all clusters; The allocation weights are respectively positively correlated with the amount of idle processor resources of the cluster and the amount of idle storage resources of the cluster; The allocation weights are respectively inversely correlated with the idle processor resources of all clusters and the idle storage resources of all clusters.

6. The system of claim 1 , further comprising: determining a task migration amount in the first cluster when it is determined that a difference between the load of the first cluster and the load of the second cluster meets a cross-cluster task adjustment condition; the first cluster and the second cluster support the same task type; Filtering tasks to be migrated with the task migration amount from the first cluster; Migrate the task to be migrated to a machine group in the second cluster that supports the task to be migrated for execution.

7. The system according to claim 6, wherein: The determining the task migration amount in the first cluster includes: determining a first candidate migration amount based on the load of the first cluster; and determining a second candidate migration amount based on the load of the second cluster; The task migration amount is selected from the first candidate migration amounts and the second candidate migration amounts.

8. The system according to claim 6, wherein: The selected tasks to be migrated are tasks in the first cluster that do not have a core task mark.

Citation Information

Patent Citations

  • Application characteristic-based isomeric group operation self-adapting dispatching method and system

    CN101739292A

  • Rendering task scheduling method, device and equipment in heterogeneous multi-cluster

    CN118535345A

  • Computing node distribution method and device, storage medium and electronic equipment

    CN118819834A