Heterogeneous task scheduling method based on multi-dimensional resource optimization
By introducing multi-dimensional resource optimization and AHP hierarchical analysis method into the task scheduling method, combining rate factor and load factor, the problem of resource allocation imbalance in traditional scheduling methods is solved, resource utilization and load balancing allocation are achieved, and the overall performance of the system is improved.
Patent Information
- Application Number
- CN202510001133.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional task scheduling methods have limitations in the fields of multi-dimensional resource optimization and task scheduling. They fail to fully consider the dynamic evolution of multi-dimensional resource requirements and node load states, resulting in unbalanced resource allocation, and some nodes may face resource overload or idleness, weakening the system's resource utilization and task processing efficiency.
The heterogeneous task scheduling method based on multi-dimensional resource optimization is adopted. By building a node cluster and a real-time data acquisition environment, real-time resource information of nodes is recorded, parameter weights of task types are determined using AHP hierarchical analysis method, a multi-dimensional node resource utilization matrix is constructed, and a rate factor and load factor are introduced to calculate the scheduling score of nodes to achieve maximum resource utilization and load balancing allocation.
The maximum utilization rate of resources is achieved, the waste of node resources is reduced, the overall resource utilization efficiency of the cluster is improved, the resource overload or idleness is avoided, and the task processing efficiency and user experience are improved.
Smart Images

Figure CN120011010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-dimensional resource and task scheduling, and in particular to a heterogeneous task scheduling method based on multi-dimensional resource optimization. Background Art
[0002] With the rapid development of cloud computing, edge computing, and big data technologies, task scheduling plays an increasingly important role in high-concurrency heterogeneous computing environments. Traditional task scheduling methods are usually limited to a single or limited resource dimension as the optimization target. For example, they focus on the utilization of CPU and memory, but ignore the heterogeneity of task types in actual applications, that is, the differences and dynamics of resource requirements of different types of tasks. Given the heterogeneity of the tasks themselves, there are obvious differences in the requirements of various tasks in terms of computing, storage, network and other resources, and traditional methods often fail to fully consider these differences during the scheduling process, resulting in a mismatch between resource allocation and actual demand. In the face of complex cluster environments, such scheduling strategies are difficult to meet actual needs.
[0003] In addition, the traditional resource scheduling model usually fails to fully evaluate the impact of the correlation of resource requirements between tasks on node resource allocation during the resource allocation process, making it difficult to achieve batch scheduling of resources. When scheduling in sequence according to models such as task queues, the resource status of all scheduling nodes is obtained in real time. However, when the scheduled task executes the corresponding scheduling strategy, the resource parameters of the scheduled node cannot be updated immediately. If the overall idleness of the node is evaluated and tasks are assigned based solely on the resource status of the current node, the node may eventually face resource overload or idleness, thereby weakening the system's resource utilization and task processing efficiency. These defects not only limit the effect of scheduling optimization, but may also cause resource waste and reduce user experience.
[0004] The main problems facing current task scheduling technology include: traditional task scheduling methods are limited to optimizing a single or limited resource dimension, ignoring the significant differences in multi-dimensional resource requirements and various node load states; at the same time, most existing technologies adopt resource scheduling methods based on task queue models, which are more likely to encounter performance bottlenecks, especially when multiple tasks are running in parallel and resource requirements fluctuate significantly. Summary of the invention
[0005] In view of the above-mentioned defects in the prior art, the core technical problem to be solved by the present invention lies in the limitations of the existing heterogeneous task scheduling methods in the field of multi-dimensional resource optimization and task scheduling. Most traditional scheduling methods only focus on the optimization of single or a few dimensional resources, and fail to fully consider the dynamic evolution of multi-dimensional resource requirements and node load status. At the same time, they ignore the batch processing characteristics of multiple tasks and the reasonable allocation of dynamic resource requirements. Such resource scheduling strategies are difficult to achieve the expected effect of intelligent scheduling in multi-task scenarios, which can easily lead to imbalanced resource allocation. Some nodes may face the dilemma of resource overload or idleness, thereby weakening the resource utilization efficiency and task processing speed of the overall system, and ultimately causing resource waste and reducing user experience. Therefore, the present invention provides a heterogeneous task scheduling method based on multi-dimensional resource optimization, which realizes the maximum utilization of resources and load balancing distribution, provides key technical support for modern computing cluster environments, and has broad application prospects.
[0006] To achieve the above object, the present invention provides a heterogeneous task scheduling method based on multi-dimensional resource optimization, comprising the following steps:
[0007] Build and configure the node cluster and the corresponding node real-time data collection environment, record the real-time resource information of each node in the cluster, including the utilization of CPU, memory, bandwidth, and disk IO resources, and obtain the type information of all tasks to be scheduled in advance through the historical scheduling statistics of the tasks;
[0008] The parameter weights of different types of tasks are determined through the AHP hierarchical analysis method, which facilitates the subsequent use of the weights to calculate the scores of the nodes;
[0009] Construct a multi-dimensional node resource utilization matrix A using the data collected at the current moment;
[0010] Set the rate factor and load factor of the nodes in the cluster for score adjustment in subsequent simulation iterations;
[0011] Perform positive processing on each type of non-maximum value resource in the cluster. The larger the maximum value, the better, to ensure that the trends of all subsequent indicator values are consistent.
[0012] Standardize the resource matrix after forwarding to eliminate the influence of different indicator dimensions;
[0013] According to the resource weight of the task and the Euclidean distance formula, the optimal and worst solutions of the node distance are calculated, and then the Euclidean distance score is calculated;
[0014] Introduce rate factor and load factor, calculate the scheduling score of the current task on each node, and select the node with the highest scheduling score to schedule the task;
[0015] When there are tasks to be scheduled in this round, the system will continuously and virtually update the node resource utilization matrix, load factor and adjustment rate factor to complete the complete allocation strategy for all tasks to be scheduled and issue it, thereby realizing batch scheduling of multiple tasks.
[0016] Furthermore, a cluster is built and a node data collection environment is configured to record node resource data in the cluster, including:
[0017] Record the real-time resource information of each node in the cluster, mainly obtaining CPU resource utilization, memory resource utilization, bandwidth resource utilization and disk resource utilization;
[0018] The tasks to be scheduled are divided into CPU type, memory type and network type. The importance of resource indicators is compared based on the calculation of historical scheduling statistics of each task and the percentage of node resources occupied by the scheduled tasks to obtain the type of tasks to be scheduled.
[0019] Furthermore, the parameter weights of different types of tasks are determined by the AHP hierarchical analysis method, which is convenient for the subsequent use of the weights to calculate the scores of nodes, specifically including the use of the AHP hierarchical analysis method to determine the parameter weights of different types of tasks. AHP ranks each custom parameter according to its importance, compares the custom parameters pairwise, evaluates the importance of each resource and derives a judgment matrix. The comparison usually uses a scale of 1-9, where 1 means that two factors are equally important and 9 means that one factor is extremely more important than another. After comparing the custom parameters pairwise, the judgment matrix is calculated, and the judgment matrix is further normalized to determine the weights W of the CPU, memory, bandwidth and disk resources. cpu , W mem , W net , W disk .
[0020] Furthermore, the data collected at the current moment is used to construct a multidimensional node resource utilization matrix, including querying the resource utilization information of all nodes. The data collected at the current moment is used to construct a multidimensional node resource utilization matrix A, which is the n*m node resource utilization matrix A at the current moment. The resource utilization rate of the first node is {A 11 , A 11 , A 11 …A 1m}, where subscripts 1 to m correspond to the resources of each dimension of the node, namely CPU, memory, bandwidth, and disk IO.
[0021] Furthermore, a rate factor V and a load factor L are preset for the node, wherein the initial value of the rate factor is set to be custom and does not exceed 1, and is subsequently set to gradually increase with the increase in the number of iterations according to actual conditions, with an upper limit of 1.
[0022] Furthermore, each type of resource in the cluster that is not a maximum value type (the larger the value, the better) is processed positively:
[0023] X ij =max-A ij
[0024] max is the maximum value of the same type of resources in all nodes of the cluster, X ij is the availability of resource j on node i, A ij It is an element in the node resource utilization matrix A. The forward matrix composed of n nodes and m evaluation indicators (after forward transformation) is as follows:
[0025]
[0026] 7. Further, the normalized resource matrix is normalized to eliminate the influence of different indicator dimensions. Specifically, the normalized matrix is recorded as Z. The normalized result of the j (j = 1, 2, 3, ... m) resource of the i-th (i = 1, 2, 3, ... n) node in Z is calculated by the following formula:
[0027]
[0028] The normalized matrix Z is as follows:
[0029]
[0030] Furthermore, the optimal node distance solution d is calculated by combining the resource weight of the task with the Euclidean distance formula. i + , the worst solution d i - , then calculate the Euclidean distance score, select the maximum and minimum values from each column of the standardized matrix (each column corresponds to each type of resource) and record the maximum value of all resources in the cluster as Z + , the minimum value is recorded as Z - , the weights obtained by AHP are combined with the Euclidean distance formula to calculate the optimal solution for the distance of the i-th (i=1,2,3,…n) node and the size of the worst solution d i + , d i - , then the Euclidean distance score D is obtained i .
[0031]
[0032] Furthermore, the rate factor and load factor are introduced to calculate the scheduling score of the current task on each node. Under the influence of the adjustment of the rate factor V and the load factor L, the Euclidean distance score D i Calculate the scheduling score S of the current task at each node i ; Therefore, the scheduling scores of each node are sorted, and the node with the highest score is selected for scheduling. The scheduling score calculation formula is as follows:
[0033] S i =L i ·V i ·D i
[0034] L i is the load factor obtained based on the node resource situation; V i is the rate factor of gradual increase; D i is the node Euclidean distance score.
[0035] Furthermore, the iteration is performed by judging whether there are still tasks to be scheduled in the current round of scheduling. During the iteration process, the resource utilization matrix and load factor are continuously updated virtually and the rate factor is adjusted. After the iteration process is completed, the complete allocation strategy of the task queue to be scheduled can be calculated, and the strategy will eventually be sent to the cluster to realize the scheduling of heterogeneous tasks in the system.
[0036] Technical Effects
[0037] The present invention provides a heterogeneous task scheduling method based on multi-dimensional resource optimization, which uses a node resource collection system to collect the real-time resource utilization information of nodes and pre-determines the types of all tasks to be scheduled through historical data statistics; then the parameter weights of different types of tasks are determined by the AHP hierarchical analysis method, and then a multi-dimensional node resource matrix is constructed and the load factor and rate factor are set; after the resource matrix is forward-oriented and standardized, the optimal solution of the node distance and the worst solution size are calculated according to the resource weight of the task and the Euclidean distance formula, and then the Euclidean distance score is calculated; finally, under the adjustment of the rate factor and the load factor, the scheduling score of the current task at each node is calculated and the node with the highest score is pre-selected for scheduling. The present invention provides a heterogeneous task scheduling method based on multi-dimensional resource optimization, which realizes the maximum utilization of resources, reduces the waste of node resources, improves the overall resource utilization efficiency of the cluster, has significant advantages in load balancing, avoiding task concentration, efficient resource utilization and flexible scalability, maximizes the resource utilization of the system cluster, and implements the load balancing distribution of resources more reasonably, which is particularly suitable for a cluster environment with dynamic changes in multiple tasks and multiple nodes.
[0038] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flow chart of a heterogeneous task scheduling method based on multi-dimensional resource optimization according to a preferred embodiment of the present invention;
[0040] Figure 2 It is a schematic diagram of an iterative update algorithm process of a heterogeneous task scheduling method based on multi-dimensional resource optimization in a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0042] In the following description, specific details such as specific internal procedures and techniques are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention can be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.
[0043] like Figure 1 As shown, an embodiment of the present invention provides a heterogeneous task scheduling method based on multi-dimensional resource optimization, which specifically includes the following steps:
[0044] Step S101, build and configure the node cluster and the corresponding node real-time data collection environment, record the real-time resource information of each node in the cluster, and obtain the type information of all tasks to be scheduled through the historical scheduling statistics of each task in the past, specifically including:
[0045] Obtain CPU resource utilization, memory resource utilization, bandwidth resource utilization, and disk resource utilization;
[0046] The tasks to be scheduled can be generally divided into three types: CPU type, memory type and network type. The importance of resource indicators is compared based on past data calculation and statistics of the percentage of node resources occupied after task scheduling to determine the type of task to be scheduled. The statistical calculation formula for determining the type of task to be scheduled is as follows:
[0047]
[0048] Where R represents the total resource load, N represents the required resource application of the Pod, T represents the maximum resource, i is the resource number, representing various types of resources such as CPU, memory, etc., j is the node label (0 < j < n, n is the number of nodes), and S i represents the resource occupancy rate of the task scheduled to the node in the statistical data, is the average percentage of this type of resource scheduled to all nodes, and type is the displayed resource type after comparison. According to the comparison of values, the type type can be obtained.
[0049] Step S102, determine the parameter weights of different types of tasks through the AHP (Analytic Hierarchy Process) method, which is convenient for calculating the scores of nodes using these weights later. AHP ranks according to the importance of each custom parameter, compares the custom parameters in pairs, evaluates the importance of each resource, and obtains a judgment matrix. (The comparison usually uses a scale of 1 - 9, where 1 means the two elements are equally important, and 9 means one element is extremely more important than the other element)
[0050] The AHP defines the parameter importance criteria and values as shown in the following table:
[0051]
[0052]
[0053] According to the criteria in the above table, compare the above-mentioned custom influencing factors, namely CPU usage rate, memory usage rate, bandwidth usage rate, and disk I / O usage rate, to obtain the values of the judgment matrix. In this embodiment, taking CPU-type tasks as an example, according to the resource demand data of the task for the cluster resources during task deployment, combined with the importance criteria values in the above table, evaluate the importance of the custom scheduling parameters, determine the importance of each resource, calculate the judgment matrix, and perform the consistency test of the judgment matrix. If not passed, readjust a reasonable judgment matrix. After passing the test, determine the values in the judgment matrix. The judgment matrix is determined as the following example:
[0054] Parameter variables CPU MEM NET DISK CPU 1 4 3 3 MEM 1 / 4 1 1 / 2 1 / 3 NET 1 / 3 2 1 1 DISK 1 / 3 3 1 1
[0055] The algorithm for calculating the weights is as follows:
[0056] i and j represent the rows and columns of the input judgment matrix S. For a judgment matrix S, first traverse each column of the matrix S, then traverse each row of S, find the sum of the j-th column of the matrix S and denote it as sum, and then traverse each row of the matrix S to perform column normalization on each element of this column:
[0057] S ij = S ij / sum
[0058] After normalizing the columns, you need to normalize the rows. Traverse each row of the matrix S, then traverse each column, and find the sum of each row of the matrix S. The i-th row is recorded as sum. Then you will get a one-dimensional array consisting of each row {sum1, sum2, sum3, sum m}. Then normalize the rows of this one-dimensional array to get the W matrix. Calculate the eigenvalues based on the obtained W and the input matrix S. Then calculate the CI, where n is the matrix dimension:
[0059] CI = (λ-n) / (n-1)
[0060] Obtain RI by looking up the official standard RI table, and finally calculate CR:
[0061] CR=CI / RI
[0062] If CR is less than 0.1, the consistency check is passed and the weight is calculated, otherwise the judgment matrix is readjusted.
[0063] According to the judgment matrix, the weight W of CPU, memory, bandwidth and disk resource dimensions can be calculated through the weight calculation algorithm. cpu , W mem , W net , W disk .
[0064] Step S103, using the data collected at the current moment to construct a multi-dimensional node resource utilization matrix A, which is the n*m node resource utilization matrix at the current moment; the resource utilization rate of the first node is {A 11 , A 11 , A 11... A 1m}, where subscripts 1 to m correspond to the resources of each dimension of the node, namely CPU, memory, bandwidth, and disk IO.
[0065] Step S104, preset the rate factor V and load factor L for the node. The initial value of the rate factor is customized and does not exceed 1. It is subsequently set to increase gradually with the increase of the number of iterations according to the actual situation, and the upper limit is 1; and the load factor of each node is calculated as follows:
[0066]
[0067] m is the number of resource dimensions, and A is the resource utilization matrix.
[0068] For the rate factor, set its change strategy, for example, set the initial value of V to 0.2, and then as the number of iterations increases, when the number of nodes scheduled increases, the rate factor increases and gradually approaches 1.
[0069] Step S105: Perform positive processing on each type of resource in the cluster that is not a maximum value type (the larger the value, the better), because we hope that the utilization index is as low as possible, but in order to unify the subsequent trends, positive processing is performed here:
[0070] X ij =max-A ij
[0071] max is the maximum value of the same type of resources in all nodes of the cluster, X ij is the availability of resource j on node i, A ij It is an element in the node resource utilization matrix A. The forward matrix composed of n nodes and m evaluation indicators (after forward transformation) is as follows:
[0072]
[0073] Step S106, normalize the resource matrix after the forward transformation. The normalized matrix is denoted as Z, where the elements can be obtained from the elements of the forward transformation matrix:
[0074]
[0075] The normalized matrix Z is as follows:
[0076]
[0077] Step S107, calculate the optimal node distance solution d according to the resource weight of the task and the Euclidean distance formula i + , the worst solution d i -, and then calculate the Euclidean distance score. Select the maximum and minimum values from each column of the standardized matrix (each column corresponds to each type of resource) and record the maximum value of all resources in the cluster as Z + , the minimum value is recorded as Z - . Z + It can be simplified to Z + ={Z1 + ,Z2 + ,…,Z m +}, indicating the maximum value of all cluster resources; Z - Simplified as Z - ={Z1 - ,Z2 - ,…,Z m -}, which represents the minimum value of all cluster resources. The weights obtained by AHP are combined with the Euclidean distance formula to calculate the optimal solution and the worst solution of the i-th (i=1,2,3,…n) node distance.
[0078]
[0079]
[0080] where d i + , d i - is the distance between the current candidate node and the optimal solution and the worst solution of the cluster. i + , d i - Get the Euclidean distance score D i .
[0081]
[0082] Step S108, introduce the rate factor and load factor to calculate the scheduling score of the current task at each node. Under the influence of the adjustment of the rate factor V and the load factor L, based on the Euclidean distance score D i Calculate the scheduling score S of the current task at each node i ; Therefore, the scheduling scores of each node are sorted, and the node with the highest score is selected for scheduling. The scheduling score calculation formula is as follows:
[0083] S i =L i ·V i ·D i
[0084] L i is the load factor obtained based on the node resource situation; V i is the rate factor of gradual increase; D i is the node Euclidean distance score.
[0085] Step S109, iterates by judging whether there are still tasks to be scheduled in this round of scheduling. During the iteration process, the resource utilization matrix and load factor are continuously updated virtually and the rate factor is adjusted. After the iteration process is completed, the complete allocation strategy of the task queue to be scheduled can be calculated, and the strategy will eventually be sent to the cluster to realize the scheduling of heterogeneous tasks in the system.
[0086] The specific iterative update algorithm process is as follows Figure 2 As shown, the specific process is as follows:
[0087] First, calculate the pre-scheduling node for the current task, and select the first one according to the scheduling score. At this time, determine whether there are still tasks that need to be scheduled to determine whether the iterative algorithm needs to continue. When all tasks to be scheduled have been assigned pre-scheduling nodes, the iterative process will end, and the system cluster will execute the assigned strategy. When there are still tasks that need to select scheduling nodes, since the current task already has a corresponding pre-scheduling node, the resource information of the pre-scheduling node needs to be virtually updated to cope with the next iteration process. For the nodes that are assigned tasks in this round, their resource utilization information will be updated. Assuming that the requirements for the utilization of CPU, memory, bandwidth, and disk IO resources in this round of tasks are a, b, c, and d respectively, the resource utilization information of the i-th node assigned the task will be updated:
[0088] A′ i1 =A i1 +a
[0089] A′ i2 =A i2 +b
[0090] A′ i3 =A i3 +c
[0091] A′ i4 =A i4 +d
[0092] The scales of 1 to 4 correspond to the utilization of CPU, memory, bandwidth, and disk IO resources.
[0093] Then the load factor of the node corresponding to the virtual update will also change accordingly, according to the calculation formula of the load factor:
[0094]
[0095] You can get the load factor value after virtual update:
[0096] L′ i =(1-A′ i1 )(1-A′ i2 )(1-A′ i3 )(1-A′ i4 )
[0097] The introduction of rate factors enables the system to gradually improve the scheduling response speed on idle nodes, ensuring the rational use of resources. It also avoids the situation where when a new node appears in the cluster, the resource utilization matrix information value of the node is generally too small, and the value after forward and normalization is generally too large, which will cause its scheduling score to be significantly higher than other nodes in the cluster. Therefore, the introduction of rate factors limits the scheduling score of new nodes to a certain extent, thereby controlling the number of tasks scheduled by the new node, avoiding the dilemma of a large number of tasks being scheduled to the same node, causing the problem of insufficient resource supply.
[0098] The update of the rate factor is related to the selected strategy, which depends on the technician's requirements for the wake-up speed of the new node. It is assumed here that the initial value of the rate factor is 0.2, and for any node, every time it is not selected by the pre-scheduling, its number of unscheduled times increases by 1. When the times value is 2, that is, the node has not been assigned a task in two rounds of iterations, at this time, its rate factor is adjusted to increase by 0.2, and its upper limit is 1. The update process of the rate factor reflects the idle state of the corresponding node in task scheduling; as the rate factor increases, the node gradually increases the scheduling priority of new tasks. When the rate factor does not reach the upper limit of 1, it will continue to increase as unassigned tasks until it reaches 1 and remains unchanged. This design ensures that the system can gradually increase the response rate of the node under low load conditions to improve the scheduling capability before the high load state arrives.
[0099] After the rate factor and load factor are virtually updated, the resource utilization of the scheduling node selected for the current round of tasks will be updated again, and the resource matrix will be re-normalized and standardized. Then the system will recalculate the distance between each node and the optimal solution and the worst solution and derive the Euclidean distance, thereby obtaining the new scheduling score for each node, and then enter a new round of iteration.
[0100] The embodiment of the present invention provides a heterogeneous task scheduling method based on multi-dimensional resource optimization. The node resource collection system collects the real-time resource utilization information of the node and determines the types of all tasks to be scheduled in advance through historical data statistics. Then, the parameter weights of different types of tasks are determined by applying the AHP hierarchical analysis method, a multi-dimensional node resource matrix is constructed, and the load factor and rate factor are set. After the resource matrix is forward-oriented and standardized, the optimal solution and the worst solution of the node distance are calculated according to the resource parameter weight of the task and the Euclidean distance formula, and then the Euclidean distance score is calculated; finally, under the adjustment of the rate factor and the load factor, the scheduling score of the current task at each node is calculated, and the node with the highest score is pre-selected for scheduling. Compared with the traditional method, this method more significantly improves the resource utilization of the system cluster and more reasonably realizes the balanced distribution of resource load.
[0101] The specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes to the concept of the present invention without creative work. Therefore, any technical solution obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A heterogeneous task scheduling method based on multi-dimensional resource optimization, characterized in that: The following steps are involved: Build and configure the node cluster and the corresponding node real-time data collection environment, record the real-time resource information of each node in the cluster, and obtain the type information of all tasks to be scheduled through the historical scheduling statistics of each task in the past; The parameter weights of different types of tasks are determined through the AHP hierarchical analysis method, which facilitates the subsequent use of the weights to calculate the scores of the nodes; Use the data collected at the current moment to build a multi-dimensional node resource utilization matrix; Preset rate factors and load factors for nodes for score adjustment in subsequent simulation iterations; Perform positive processing on each type of non-maximum value resource in the cluster to ensure that the trends of all subsequent indicator values are consistent; Standardize the resource matrix after forwarding to eliminate the influence of different indicator dimensions; The optimal and the worst node distance solutions are calculated based on the resource weight of the task and the Euclidean distance formula, and then the Euclidean distance score is calculated; Under the adjustment of rate factor and load factor, the scheduling score of the current task on each node is calculated, and the node with the highest scheduling score is selected for scheduling; Iteration is performed by determining whether there are tasks that need to be scheduled in this round, and the resource utilization matrix, load factor and adjustment rate factor are continuously virtually updated during the iteration process to complete the complete allocation strategy for all tasks to be scheduled and issue them.
2. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 1, characterized in that: Build and configure the node cluster and the corresponding node real-time data collection environment, record the real-time resource information of each node in the cluster, and obtain the type information of all tasks to be scheduled through the historical scheduling statistics of each task in the past, including: Obtain CPU resource utilization, memory resource utilization, bandwidth resource utilization, and disk resource utilization; The tasks to be scheduled include CPU type, memory type and network type. The importance of resource indicators is compared based on the historical scheduling statistical data of each task and the percentage of node resources occupied by the scheduled tasks, so as to obtain the type of tasks to be scheduled.
3. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 1, characterized in that: The parameter weights of different types of tasks are determined by the AHP hierarchical analysis method, which is convenient for the subsequent use of the weight to calculate the node score. Specifically, the weights are calculated by the AHP hierarchical analysis method. AHP ranks each custom parameter according to its importance, and the judgment matrix is calculated after comparing the custom parameters pairwise. The judgment matrix is further normalized to determine the weights W of CPU, memory, bandwidth and disk resources. cpu , W mem , W net , W disk .
4. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 1, characterized in that: The multi-dimensional node resource utilization matrix is constructed using the data collected at the current moment, including querying the resource utilization information of all nodes to construct the n*m node resource utilization matrix A at the current moment. The resource utilization rate of the first node is {A 11 , A 11 , A 11… A 1m }, where subscripts 1 to m correspond to the resources of each dimension of the node, namely CPU, memory, bandwidth, and disk IO.
5. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 4, characterized in that: The rate factor V and load factor L are preset for the node, where the initial value of the rate factor is set to custom and is set not to exceed 1. The subsequent settings are gradually increased with the increase in the number of iterations, and its upper limit is 1.
6. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 5, characterized in that: Each type of non-maximum resource in the cluster is processed positively, specifically: X ij =max-A ij Among them, max is the maximum value of the same type of resources of all nodes in the cluster, X ij is the availability of resource j on node i, A ij It is an element in the node resource utilization matrix A; the forward matrix composed of n nodes and m forward-oriented evaluation indicators is as follows:
7. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 6, characterized in that: The resource matrix after forwarding is standardized to eliminate the influence of different indicator dimensions. Specifically, the standardized matrix is recorded as Z. The standardized result of the j (j = 1, 2, 3, ... m) resource of the i-th (i = 1, 2, 3, ... n) node in Z is calculated by the following formula: The normalized matrix Z is as follows:
8. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 7, characterized in that: The optimal and the worst node distance solutions are calculated by combining the task resource weight with the Euclidean distance formula, and then the Euclidean distance score is calculated. Specifically, the maximum and minimum values are selected from each column of the standardized matrix and the maximum value of all cluster resources is recorded as Z. + , the minimum value is recorded as Z - , where each column in the matrix corresponds to each type of resource; the weights obtained by AHP are combined with the Euclidean distance formula to calculate the optimal solution for the distance of the i-th (i=1,2,3,…n) node, and the Euclidean distance score D is obtained. i : Among them, d i + , d i - is the size of the worst solution.
9. A heterogeneous task scheduling method based on multi-dimensional resource optimization as claimed in claim 8, characterized in that: Under the influence of the adjustment of rate factor V and load factor L, based on the Euclidean distance score D i Calculate the scheduling score S of the current task at each node i ; Therefore, the scheduling scores of each node are sorted, and the node with the highest score is selected for scheduling. The scheduling score calculation formula is as follows: S i =L i ·V i ·D i Among them, L i is the load factor obtained based on the node resource situation; V i is the rate factor of gradual increase; D i is the node Euclidean distance score.
10. The heterogeneous task scheduling method based on multi-dimensional resource optimization according to claim 1, characterized in that: Iteration is performed by judging whether there are still tasks that need to be scheduled in this round of scheduling. During the iteration process, the resource utilization matrix and load factor are continuously updated virtually and the rate factor is adjusted. After the iteration process is completed, the complete allocation strategy of the task queue to be scheduled is calculated. Finally, the strategy is sent to the cluster to realize the scheduling of heterogeneous tasks in the system.
Citation Information
Cited By
Industrial equipment intelligent operation optimization method and system
CN120256138A
Multi-task processing method and system based on large model
CN120973497A