Intelligent scheduling method and device for distributed idle computing power cluster

By using intelligent scheduling methods in distributed idle computing power clusters, using node feature data to form spatial point clusters and match task requirements, the problems of low resource utilization and frequent task interruptions in resource scheduling are solved, and more efficient and stable task execution is achieved.

CN120029777AInactive Publication Date: 2025-05-23BEIJING GONGJI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510144507.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Resource scheduling in distributed idle computing power clusters faces challenges such as network latency, hardware heterogeneity, and dynamic task load changes, resulting in low resource utilization and frequent task interruptions.

Method used

By obtaining the real-time feature data of the node, the spatial points in the vector space are determined, and a clustering algorithm is used to form a spatial point cluster. Match the task with the spatial point cluster, select the target space point that meets the task needs, and monitor the task progress and node resource usage in real time, and dynamically adjust task allocation.

Benefits of technology

It improves the resource utilization efficiency of the cluster and the stability of task execution, avoids task interruptions caused by node state changes, and optimizes resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029777A_ABST
    Figure CN120029777A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent scheduling method and device for a distributed idle computing power cluster, and the method comprises the steps: determining space points corresponding to a plurality of nodes in a vector space based on the feature data of the plurality of nodes; clustering the plurality of spatial points to obtain at least one spatial point cluster; matching the plurality of tasks with each spatial point cluster to obtain the spatial point cluster matched with each task; for each task, selecting a target space point from the space point cluster matched with the task; triggering a target node corresponding to the target space point to execute the task, and monitoring the progress of the task and the resource use condition of the target node; and under the condition that the progress of the task is execution interruption and / or the resource use condition of the target node is resource shortage, selecting a new target space point from the space point cluster matched with the task, and triggering a new target node corresponding to the new target space point to execute the task. The method can effectively improve the resource utilization efficiency of the cluster and the stability of task execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of distributed clusters, and in particular to an intelligent scheduling method and device for distributed idle computing power clusters. Background Art

[0002] In a distributed idle computing cluster, computing resources come from personal computers and small and medium-sized IDC (Internet Data Center) nodes on the public Internet. These nodes are not standardized intranet devices. Each node may have different computing power and hardware configuration, and its online and offline time is uncertain. This makes resource scheduling face many challenges, including network latency, hardware heterogeneity, and dynamic load changes of tasks.

[0003] At present, the computing power scheduling methods in the industry generally adopt relatively simple task allocation methods, such as random allocation or allocation based on fixed rules, which cannot adapt to the random online and offline behaviors of distributed idle nodes, resulting in low resource utilization and frequent task interruptions. In addition, the hardware configurations of nodes vary greatly. The existing simple allocation methods cannot intelligently match the most suitable nodes according to the specific requirements of the tasks (such as GPU performance, network latency, etc.). In addition, since the nodes are deployed on the public network, network latency and volatility are inevitable. The existing methods fail to effectively consider the network conditions, resulting in low efficiency in task transmission and execution.

[0004] Therefore, how to improve the resource utilization efficiency of the cluster and the stability of task execution has become an urgent problem to be solved in this field. Summary of the invention

[0005] The present application provides an intelligent scheduling method and device for a distributed idle computing power cluster, aiming to improve the resource utilization efficiency of the cluster and the stability of task execution.

[0006] In order to achieve the above objectives, this application provides the following technical solutions:

[0007] An intelligent scheduling method for distributed idle computing power clusters, comprising:

[0008] Obtain feature data uploaded by multiple nodes in real time;

[0009] Based on the feature data, determining spatial points corresponding to a plurality of the nodes in a vector space;

[0010] Clustering the plurality of spatial points using a preset clustering algorithm to obtain at least one spatial point cluster; the spatial point cluster includes a plurality of spatial points having similarities;

[0011] Matching multiple tasks with each of the spatial point clusters to obtain a spatial point cluster matched by each of the tasks;

[0012] For each of the tasks, selecting a target spatial point that meets the running requirements of the task from a spatial point cluster matching the task;

[0013] Triggering the target node corresponding to the target spatial point to execute the task, and monitoring the progress of the task and the resource usage of the target node in real time;

[0014] When the progress of the task is execution interruption, and / or the resource usage of the target node is resource-constrained, a new target spatial point that meets the running requirements of the task is selected from the spatial point cluster matching the task, and the new target node corresponding to the new target spatial point is triggered to execute the task.

[0015] Optionally, determining, based on the feature data, space points corresponding to the plurality of nodes in the vector space includes:

[0016] Based on the feature data of each of the nodes, determining the feature vector corresponding to each of the nodes in the high-dimensional vector space; the feature data at least includes: GPU performance, CPU performance, memory performance, network latency, network bandwidth, memory occupancy, and video memory occupancy;

[0017] Normalizing all feature vectors in the high-dimensional vector space to obtain a normalized vector of each node;

[0018] Performing dimensionality reduction on each of the normalized vectors to obtain a low-dimensional vector corresponding to each of the nodes in the low-dimensional vector space;

[0019] Based on the low-dimensional vector corresponding to each of the nodes, a spatial point corresponding to each of the nodes in the vector space is determined.

[0020] Optionally, matching multiple tasks with each of the spatial point clusters to obtain a spatial point cluster matched by each of the tasks includes:

[0021] Obtain demand feature data for multiple tasks;

[0022] Based on the demand characteristic data, determining target points corresponding to a plurality of the tasks in the vector space;

[0023] Matching the plurality of target points with each of the spatial point clusters to determine the spatial point cluster to which each of the target points belongs;

[0024] Based on the spatial point cluster to which each of the target points belongs, a spatial point cluster matched by each of the tasks is determined.

[0025] Optionally, matching the plurality of target points with the respective spatial point clusters to determine the spatial point cluster to which each target point belongs includes:

[0026] Determining the centroid of each of the spatial point clusters;

[0027] For each of the target points, respectively calculating the distance between the target point and the centroid of each of the spatial point clusters;

[0028] The spatial point cluster to which the target point belongs is determined based on the spatial point cluster whose centroid is the smallest and the target point.

[0029] Optionally, selecting a target spatial point that meets the running requirements of the task from a spatial point cluster matching the task includes:

[0030] If the running requirement of the task is high-performance AI computing, a target spatial point whose GPU performance of the corresponding node meets the preset GPU performance standard and whose network delay meets the preset delay standard is selected from the spatial point cluster matching the task.

[0031] Optionally, a target spatial point that meets the running requirements of the task is selected from a spatial point cluster that matches the task.

[0032] If the task requires long-term continuous operation, a target spatial point whose online time of the corresponding node meets the preset online time standard and whose network bandwidth meets the preset bandwidth standard is selected from the spatial point cluster matching the task.

[0033] Optionally, the method further includes:

[0034] If the computing power provided by the spatial point cluster matching the task does not meet the computing power requirements of the task, a spare spatial point cluster with a matching degree second only to the spatial point cluster matching the task is selected from the remaining spatial point clusters to provide additional computing power support for the task; the matching degree is determined based on the distance between the target point corresponding to the task and the center of mass of the spare spatial point cluster.

[0035] An intelligent scheduling device for distributed idle computing power clusters, comprising:

[0036] A data collection unit, used to obtain feature data uploaded by multiple nodes in real time;

[0037] A spatial point updating unit, used for determining spatial points corresponding to a plurality of the nodes in the vector space based on the feature data;

[0038] A spatial point clustering unit, used to cluster the plurality of spatial points using a preset clustering algorithm to obtain at least one spatial point cluster; the spatial point cluster includes a plurality of spatial points having similarities;

[0039] A task matching unit, used for matching a plurality of tasks with each of the spatial point clusters to obtain a spatial point cluster matched by each of the tasks;

[0040] A spatial point screening unit, configured to select, for each of the tasks, a target spatial point that meets the running requirements of the task from a spatial point cluster that matches the task;

[0041] A task monitoring unit, used to trigger the target node corresponding to the target spatial point to execute the task, and to monitor the progress of the task and the resource usage of the target node in real time;

[0042] The task monitoring unit is also used to select a new target spatial point that meets the running requirements of the task from the spatial point cluster matching the task when the progress of the task is interrupted and / or the resource usage of the target node is resource-constrained, and trigger the new target node corresponding to the new target spatial point to execute the task.

[0043] A storage medium includes a stored program, wherein the program is executed by a processor when it is run to execute the intelligent scheduling method for distributed idle computing power clusters.

[0044] An electronic device, comprising: a processor, a memory and a bus; the processor and the memory are connected via the bus;

[0045] The memory is used to store programs, and the processor is used to run programs, wherein the program is executed by the processor when it is run to execute the intelligent scheduling method for distributed idle computing power clusters.

[0046] The technical solution provided by the present application determines the spatial points corresponding to multiple nodes in the vector space based on the feature data of multiple nodes. Cluster multiple spatial points to obtain at least one spatial point cluster. Match multiple tasks with each spatial point cluster to obtain a spatial point cluster matched by each task. For each task, select a target spatial point from the spatial point cluster that matches the task. Trigger the target node corresponding to the target spatial point to execute the task, and monitor the progress of the task and the resource usage of the target node. When the progress of the task is an execution interruption, and / or the resource usage of the target node is resource-constrained, select a new target spatial point from the spatial point cluster that matches the task, and trigger the new target node corresponding to the new target spatial point to execute the task. The present application determines the spatial points corresponding to multiple nodes in the vector space based on the characteristic data of multiple nodes, obtains at least one spatial point cluster based on the multiple spatial points, and determines the target spatial points that meet the running requirements of the task by matching the task with each spatial point cluster, thereby directly solving the resource waste and task terminal caused by problems such as node hardware differences and network delays. In addition, during the task execution process, the task allocation is adjusted (i.e., the new target node is triggered to execute the task) according to the progress of the task and the resource usage of the target node, to ensure the stable execution of the task and avoid the problem of task interruption due to changes in node status. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 A flow chart of an intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application;

[0049] Figure 2 A flowchart of another intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application;

[0050] Figure 3 A flowchart of another intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application;

[0051] Figure 4 A flowchart of another intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application;

[0052] Figure 5A flowchart of another intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application;

[0053] Figure 6 A schematic diagram of the architecture of an intelligent scheduling device for a distributed idle computing power cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0055] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0056] like Figure 1 As shown, it is a flow chart of an intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application, which can be applied to the server side and includes the steps shown below.

[0057] S101: Obtain feature data uploaded by multiple nodes in real time.

[0058] Among them, multiple nodes can be working nodes included in the distributed idle computing power cluster.

[0059] In some examples, the working node can be deployed in a computer device, which is pre-installed with a client. The client can collect characteristic data of each working node in the computer device in real time (for example, every 3 minutes). The characteristic data includes but is not limited to: GPU performance, CPU performance, memory performance, network latency, network bandwidth, memory occupancy, and video memory occupancy.

[0060] In a possible implementation, the client can use a professional GPU parameter collection tool (such as the nvidia-smi tool) to collect GPU performance and video memory usage, can use a network monitoring tool (such as iperf) to collect network latency and network bandwidth, and can access the task manager of the computer device to collect CPU performance, memory performance, and memory usage.

[0061] It should be noted that after obtaining the characteristic data of multiple nodes, the server will also archive the characteristic data of the multiple nodes and add the latest data records of the multiple nodes to the preset node list so as to update the status of the multiple nodes in real time.

[0062] S102: Determine spatial points corresponding to multiple nodes in the vector space based on the feature data.

[0063] Among them, after obtaining the feature data of multiple nodes, the feature data can be represented as spatial points in the vector space, so that in the subsequent task scheduling process, the nodes matching the task can be selected through geometric distance and feature matching, to ensure that task allocation can be based on the close match between node characteristics and task requirements, solving the problem of uneven node resource utilization in the cluster.

[0064] Optionally, based on feature data, determine the spatial points corresponding to multiple nodes in the vector space. See Figure 2 The steps are shown and the corresponding explanations.

[0065] S103: Clustering the multiple spatial points using a preset clustering algorithm to obtain at least one spatial point cluster.

[0066] The spatial point cluster includes multiple spatial points with similarities.

[0067] In some examples, the preset clustering algorithms include but are not limited to K-Means algorithm, DBSCAN algorithm, etc.

[0068] It should be noted that the multiple spatial points with similarity shown in the spatial point cluster can be understood as clustering multiple spatial points with similar feature vectors of corresponding nodes into the same spatial point cluster. In a possible implementation, each spatial point cluster can be regarded as a node group (or node cluster), which includes multiple nodes with similar feature vectors.

[0069] S104: Match the multiple tasks with each spatial point cluster to obtain a spatial point cluster matched by each task.

[0070] Among them, multiple tasks include computing tasks whose computing power requirements are not met, and computing tasks that have not yet been assigned nodes. Multiple tasks are matched with each spatial point cluster to obtain the spatial point cluster matched by each task. The main purpose is to assign a suitable node group to each task, that is, to assign a suitable computing power support to each task (the nodes in the node group can be regarded as computing power support, which is used to provide corresponding services for the task).

[0071] Optionally, multiple tasks are matched with each spatial point cluster to obtain the implementation process of the spatial point cluster matched by each task. For details, see Figure 3 The steps are shown and the corresponding explanations.

[0072] It should be noted that the centroid and spread of each spatial point cluster can also be compared with the requirements of each task to obtain the most suitable spatial point cluster sequence for each task. The spatial point clusters in the spatial point cluster sequence are arranged in descending order according to the matching degree.

[0073] In some examples, based on the requirements of the task, a target point cluster in the vector space can be obtained, and the centroid of the target point cluster is compared with the centroid of each spatial point cluster to obtain each centroid comparison result, and the scatter of the target point cluster is compared with the scatter of each spatial point cluster to obtain each scatter comparison result, and based on each centroid comparison result and each scatter comparison result, the comparison result between the target point cluster and each spatial point cluster is determined, and based on the comparison result between the target point cluster and each spatial point cluster, the similarity between the target point cluster and each spatial point cluster is determined, and finally, based on the similarity between the target point cluster and each spatial point cluster, the matching degree between the task and each spatial point cluster is determined.

[0074] In a possible implementation, the higher the similarity between the target point cluster and the space point cluster, the higher the matching degree between the task and the space point cluster; the lower the similarity between the target point cluster and the space point cluster, the lower the matching degree between the task and the space point cluster.

[0075] S105: For each task, a target spatial point that meets the running requirements of the task is selected from the spatial point cluster that matches the task.

[0076] The operation requirement may be a requirement of the task on the number of nodes. Specifically, after determining a spatial point cluster matching the task, N target spatial points may be selected from the spatial point cluster according to the requirement N of the task on the number of nodes.

[0077] Optionally, if the task requires high-performance AI computing, select a target spatial point from a spatial point cluster that matches the task, whose corresponding node's GPU performance meets a preset GPU performance standard and whose network delay meets a preset delay standard.

[0078] Optionally, if the task requires long-term continuous operation, a target spatial point whose online time of the corresponding node meets the preset online time standard and whose network bandwidth meets the preset bandwidth standard is selected from the spatial point cluster matching the task.

[0079] S106: Trigger the target node corresponding to the target spatial point to execute the task, and monitor the progress of the task and the resource usage of the target node in real time.

[0080] The server distributes the task to the corresponding target node so that the target node executes the task and monitors the progress of the task and the resource usage of the target node in real time.

[0081] S107: When the progress of the task is execution interruption, and / or the resource usage of the target node is resource-constrained, a new target spatial point that meets the running requirements of the task is selected from the spatial point cluster that matches the task, and the new target node corresponding to the new target spatial point is triggered to execute the task.

[0082] Among them, the reason for the interruption of task execution may be that the target node is offline or faulty. Therefore, a new target node needs to be used to replace the task to ensure that the task is completed. In addition, if the resource usage of the target node is tight, the execution efficiency of the task will be reduced. Therefore, a new target node needs to be used to replace the task to ensure that the execution efficiency of the task is not affected.

[0083] It is understandable that during the task execution process, if a node is interrupted due to offline or other resource competition, the server will automatically reschedule the task and select a new target node to take over to ensure that the task can continue to be executed.

[0084] Optionally, if the computing power provided by the spatial point cluster matching the task does not meet the computing power requirements of the task, a spare spatial point cluster with a matching degree second only to the spatial point cluster matching the task is selected from the remaining spatial point clusters to provide additional computing power support for the task, and the matching degree is determined based on the distance between the target point corresponding to the task and the center of mass of the spare spatial point cluster.

[0085] In some examples, the calculation method of the matching degree between the task and the spatial point cluster can also refer to the explanation of S104. After the matching degree between the task and each spatial point cluster is calculated, each spatial point cluster can be sorted in order from high to low matching degree to obtain a point cluster sequence, so as to determine the backup spatial point cluster directly through the point cluster sequence.

[0086] Combination Figure 2-Figure 4 The method shown in the figure can also be summarized as follows: Figure 5The process shown includes: step 1, data collection: collecting working node features (i.e., feature data); step 2, coordinate update: updating the coordinates of working nodes in high-dimensional vector space; step 3, dimension reduction: normalizing the high-dimensional vector space and reducing the dimension to obtain a low-dimensional vector space; step 4, clustering: clustering the low-dimensional vector space to obtain several node clusters; step 5, matching: comparing the centroid and dispersion of each node cluster with the requirements of each task to obtain the most suitable node cluster order for each task; step 6, scheduling: for each computing task with computing power requirements to be met, randomly (or according to custom rules) find a node from the best node cluster to provide service. If the nodes in the node cluster are not enough (i.e., the computing power does not meet the computing power requirements of the computing task), it will be postponed to the second-best node cluster. Based on steps 1 to 6, the server can realize intelligent scheduling and efficient utilization of distributed idle computing power resources, solving multiple problems in existing scheduling schemes (low resource utilization efficiency of the cluster and low stability of task execution).

[0087] Compared with the prior art, the embodiments of the present application have the following innovations: (1) High-dimensional vector space modeling: The high-dimensional features of each node (such as GPU performance, network latency, CPU usage, etc.) are represented as vectors, so that the scheduling algorithm can select the optimal node through geometric matching in the high-dimensional vector space; (2) Dynamic task scheduling and adjustment: The system can dynamically adjust the allocation of tasks according to the real-time status of the node (such as online time, network fluctuations, etc.), ensuring that the task can be smoothly transitioned when the node is online or offline or the status changes, avoiding task interruption; (3) Combination of clustering and dimensionality reduction technology: PCA dimensionality reduction technology and preset clustering algorithm are used to efficiently group heterogeneous nodes, improve scheduling efficiency and reduce computational complexity; (4) Node cluster strategy optimization: In task allocation, an optimization strategy based on node clusters in high-dimensional vector space is applied to prioritize the allocation of the best node clusters, maximize system resource utilization and task execution efficiency; (5) Accurate matching of task requirements and node characteristics: The system accurately matches the specific requirements of each task (such as GPU memory requirements and computing power requirements), improving the stability and success rate of task execution.

[0088] The process shown in S101-S107 above determines the spatial points corresponding to the multiple nodes in the vector space based on the feature data of the multiple nodes, obtains at least one spatial point cluster based on the multiple spatial points, and determines the target spatial points that meet the running requirements of the task by matching the task with the various spatial point clusters, thereby directly solving the resource waste and task terminal caused by problems such as node hardware differences and network delays. In addition, during the task execution process, the task allocation is adjusted (i.e., the new target node is triggered to execute the task) according to the progress of the task and the resource usage of the target node, to ensure the stable execution of the task and avoid the problem of task interruption due to changes in node status.

[0089] like Figure 2 As shown, a flow chart of an intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application includes the following steps.

[0090] S201: Determine a feature vector corresponding to each node in a high-dimensional vector space based on feature data of each node.

[0091] The characteristic data includes at least: GPU performance, CPU performance, memory performance, network latency, network bandwidth, memory occupancy, and video memory occupancy.

[0092] In some examples, the dimension of the high-dimensional vector space is determined based on the total categories of parameter items in the feature data. Specifically, if the feature data contains 7 categories of parameter items, the dimension of the high-dimensional vector space can be considered as 7 dimensions.

[0093] It should be noted that constructing the feature vector of a node in a high-dimensional vector space based on the feature data of the node is a common use of mathematical knowledge and will not be elaborated here.

[0094] S202: Normalize all feature vectors in the high-dimensional vector space to obtain a normalized vector of each node.

[0095] Among them, the measurements of various parameter items in the feature data are not uniform, and the format of the feature data of each node is not necessarily the same. Therefore, all feature vectors in the high-dimensional vector space are normalized to obtain the normalized vector of each node to improve data reliability.

[0096] S203: Perform dimensionality reduction on each normalized vector to obtain a low-dimensional vector corresponding to each node in the low-dimensional vector space.

[0097] Among them, a dimensionality reduction algorithm (such as the PCA algorithm and the T-SNE algorithm) can be used to reduce the dimensionality of each normalized vector to obtain a low-dimensional vector corresponding to each node in the low-dimensional vector space, so as to reduce the computational overhead of the subsequent clustering process while maintaining the effective expression of node features, which is convenient for subsequent spatial point clustering and task scheduling.

[0098] In some examples, the dimension of the low-dimensional vector space is less than that of the high-dimensional vector space. Specifically, the dimension of the low-dimensional vector space may be less than 5.

[0099] S204: Based on the low-dimensional vector corresponding to each node, determine the space point corresponding to each node in the vector space.

[0100] Among them, a vector refers to a quantity with size and direction, which can be visualized as a line segment with an arrow. The line segment includes a starting point and an end point. Generally speaking, the starting point of a low-dimensional vector is the origin in the vector space, and the end point can be regarded as the spatial point corresponding to the node.

[0101] The process shown in S201-S204 above uses the feature data of the node to determine the spatial point of the node in the vector space, thereby providing effective reference data for subsequent spatial point clustering and task scheduling.

[0102] like Figure 3 As shown, a flow chart of an intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application includes the following steps.

[0103] S301: Obtaining demand feature data of multiple tasks.

[0104] The demand characteristic data of the task include but are not limited to GPU performance requirements, network latency requirements, etc.

[0105] In a possible implementation manner, the task may be analyzed to obtain the task requirement characteristic data.

[0106] S302: Determine target points corresponding to multiple tasks in the vector space based on the demand feature data.

[0107] Among them, if the demand feature data contains more parameter items, the process of determining the target points corresponding to multiple tasks in the vector space based on the demand feature data of multiple tasks can be: based on the demand feature data of each task, determine the feature vector corresponding to each task in the high-dimensional vector space; normalize the feature vector corresponding to each task in the high-dimensional vector space to obtain the normalized vector of each task; reduce the dimension of the normalized vector of each task to obtain the low-dimensional vector corresponding to each task in the low-dimensional vector space (that is, the vector space to which the spatial point corresponding to the node belongs); based on the low-dimensional vector corresponding to each task, determine the target point corresponding to each task in the vector space.

[0108] In some examples, if the requirement feature data contains fewer parameter items and the total category of the parameter items conforms to the dimension of the low-dimensional vector space, then the process of determining the target points corresponding to multiple tasks in the vector space based on the requirement feature data of multiple tasks can be: based on the requirement feature data of each task, determining the feature vector corresponding to each task in the low-dimensional vector space; based on the feature vector corresponding to each task in the low-dimensional vector space, determining the target point corresponding to each task in the vector space.

[0109] S303: Match the multiple target points with the respective spatial point clusters to determine the spatial point cluster to which each target point belongs.

[0110] Among them, the target point and each spatial point cluster can be matched according to the distance between the target point and the centroid of each spatial point cluster to determine the spatial point cluster to which the target point belongs. Generally speaking, the spatial point cluster to which the target point belongs, the node corresponding to each spatial point in the spatial point cluster, is the node most suitable for the task corresponding to the target point.

[0111] Optionally, multiple target points are matched with each spatial point cluster to determine the spatial point cluster to which each target point belongs. For details, see Figure 4 The steps are shown and the corresponding explanations.

[0112] S304: Based on the spatial point cluster to which each target point belongs, determine the spatial point cluster matched by each task.

[0113] Among them, based on the spatial point cluster to which each target point belongs, the spatial point cluster that matches each task is determined. In fact, each task is bound to the spatial point cluster to which the corresponding target point belongs. In the subsequent task scheduling process, tasks are preferentially distributed to the nodes corresponding to the spatial points in the bound spatial point cluster (which can be regarded as the best spatial point cluster).

[0114] In the above process S301 - S304 , multiple tasks are matched with each spatial point cluster according to the demand feature data of multiple tasks, so as to determine the spatial point cluster (ie, node grouping) matched with each task.

[0115] like Figure 4 As shown, a flow chart of an intelligent scheduling method for a distributed idle computing power cluster provided in an embodiment of the present application includes the following steps.

[0116] S401: Determine the centroid of each spatial point cluster.

[0117] Among them, after clustering each spatial point using a preset clustering algorithm, each spatial point cluster and the corresponding centroid can be obtained.

[0118] S402: For each target point, respectively calculate the distance between the target point and the centroid of each spatial point cluster.

[0119] Among them, the farther the distance between the target point and the centroid of the spatial point cluster, the lower the matching degree between the target point and the spatial point cluster, and the closer the distance between the target point and the centroid of the spatial point cluster, the higher the matching degree between the target point and the spatial point cluster.

[0120] S403: Determine the spatial point cluster to which the target point belongs based on the spatial point cluster whose centroid has the smallest distance with the target point.

[0121] Among them, after determining the distance between the target point and the centroid of each spatial point cluster, the spatial point cluster with the smallest distance between the centroid and the target point can be regarded as the spatial point cluster with the highest matching degree with the target point.

[0122] In the above process S401-S403, the distance between the target point corresponding to the task and the centroid of the spatial point cluster is used to determine the matching degree between the task and the spatial point cluster, so as to achieve matching between multiple tasks and each spatial point cluster.

[0123] like Figure 6 As shown, it is a schematic diagram of the architecture of an intelligent scheduling device for a distributed idle computing power cluster provided in an embodiment of the present application, including the units shown below.

[0124] The data collection unit 100 is used to obtain feature data uploaded by multiple nodes in real time.

[0125] The spatial point updating unit 200 is used to determine the spatial points corresponding to the multiple nodes in the vector space based on the feature data.

[0126] Optionally, the spatial point update unit 200 is specifically used to: determine the feature vector corresponding to each node in the high-dimensional vector space based on the feature data of each node; the feature data includes at least: GPU performance, CPU performance, memory performance, network latency, network bandwidth, memory occupancy, and video memory occupancy; normalize all feature vectors in the high-dimensional vector space to obtain a normalized vector for each node; reduce the dimension of each normalized vector to obtain a low-dimensional vector corresponding to each node in the low-dimensional vector space; and determine the spatial point corresponding to each node in the vector space based on the low-dimensional vector corresponding to each node.

[0127] The spatial point clustering unit 300 is used to cluster multiple spatial points using a preset clustering algorithm to obtain at least one spatial point cluster; the spatial point cluster includes multiple spatial points with similarities.

[0128] The task matching unit 400 is used to match multiple tasks with each spatial point cluster to obtain a spatial point cluster matched by each task.

[0129] Optionally, the task matching unit 400 is specifically used to: obtain requirement feature data for multiple tasks; determine target points corresponding to multiple tasks in the vector space based on the requirement feature data; match multiple target points with each spatial point cluster to determine the spatial point cluster to which each target point belongs; and determine the spatial point cluster matched by each task based on the spatial point cluster to which each target point belongs.

[0130] Optionally, the task matching unit 400 is specifically used to: determine the center of mass of each spatial point cluster; for each target point, calculate the distance between the target point and the center of mass of each spatial point cluster; and determine the spatial point cluster to which the target point belongs based on the spatial point cluster with the smallest distance between the center of mass and the target point.

[0131] The spatial point screening unit 500 is used to select, for each task, a target spatial point that meets the running requirements of the task from the spatial point cluster that matches the task.

[0132] Optionally, the spatial point screening unit 500 is specifically used for: if the running requirement of the task is high-performance AI computing, selecting a target spatial point whose corresponding node's GPU performance meets the preset GPU performance standard and whose network delay meets the preset delay standard from the spatial point cluster matching the task.

[0133] Optionally, the spatial point screening unit 500 is specifically used for: if the task requires long-term continuous operation, selecting a target spatial point from the spatial point cluster matching the task, the target spatial point whose corresponding node's online time meets the preset online time standard and whose network bandwidth meets the preset bandwidth standard.

[0134] The task monitoring unit 600 is used to trigger the target node corresponding to the target space point to execute the task, and to monitor the progress of the task and the resource usage of the target node in real time.

[0135] The task monitoring unit 600 is also used to select a new target spatial point that meets the running requirements of the task from the spatial point cluster matching the task when the progress of the task is interrupted and / or the resource usage of the target node is resource-constrained, and trigger the new target node corresponding to the new target spatial point to execute the task.

[0136] Optionally, the task monitoring unit 600 is also used to: if the computing power provided by the spatial point cluster matching the task does not meet the computing power requirements of the task, then select a spare spatial point cluster from the remaining spatial point clusters whose matching degree is second only to the spatial point cluster matching the task, to provide additional computing power support for the task; the matching degree is determined based on the distance between the target point corresponding to the task and the center of mass of the spare spatial point cluster.

[0137] Each unit shown above determines the spatial points corresponding to multiple nodes in the vector space based on the characteristic data of multiple nodes, obtains at least one spatial point cluster based on the multiple spatial points, and determines the target spatial point that meets the running requirements of the task by matching the task with each spatial point cluster, thereby directly solving the resource waste and task terminal caused by problems such as node hardware differences and network delays. In addition, during the task execution process, the task allocation is adjusted (i.e., the new target node is triggered to execute the task) according to the progress of the task and the resource usage of the target node, to ensure the stable execution of the task and avoid the problem of task interruption due to node status changes.

[0138] The present application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the intelligent scheduling method for distributed idle computing power clusters provided by the present application.

[0139] The present application also provides an electronic device, including: a processor, a memory and a bus. The processor and the memory are connected via a bus, the memory is used to store programs, and the processor is used to run programs, wherein when the program is running, the intelligent scheduling method for distributed idle computing power clusters provided by the present application is executed.

[0140] In addition, the functions described above in the embodiments of the present application may be at least partially performed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0141] Although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present application. Certain features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination.

[0142] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.

Claims

1. An intelligent scheduling method for distributed idle computing power clusters, characterized in that: include: Obtain feature data uploaded by multiple nodes in real time; Based on the feature data, determining spatial points corresponding to a plurality of the nodes in a vector space; Clustering the plurality of spatial points using a preset clustering algorithm to obtain at least one spatial point cluster; the spatial point cluster includes a plurality of spatial points having similarities; Matching multiple tasks with each of the spatial point clusters to obtain a spatial point cluster matched by each of the tasks; For each of the tasks, selecting a target spatial point that meets the running requirements of the task from a spatial point cluster matching the task; Triggering the target node corresponding to the target spatial point to execute the task, and monitoring the progress of the task and the resource usage of the target node in real time; When the progress of the task is execution interruption, and / or the resource usage of the target node is resource-constrained, a new target spatial point that meets the running requirements of the task is selected from the spatial point cluster matching the task, and the new target node corresponding to the new target spatial point is triggered to execute the task.

2. The method according to claim 1, characterized in that Determining, based on the feature data, space points corresponding to the plurality of nodes in the vector space, comprising: Based on the feature data of each of the nodes, determining the feature vector corresponding to each of the nodes in the high-dimensional vector space; the feature data at least includes: GPU performance, CPU performance, memory performance, network latency, network bandwidth, memory occupancy, and video memory occupancy; Normalizing all feature vectors in the high-dimensional vector space to obtain a normalized vector of each node; Performing dimensionality reduction on each of the normalized vectors to obtain a low-dimensional vector corresponding to each of the nodes in the low-dimensional vector space; Based on the low-dimensional vector corresponding to each of the nodes, a spatial point corresponding to each of the nodes in the vector space is determined.

3. The method according to claim 1, characterized in that: Matching the plurality of tasks with each of the spatial point clusters to obtain a spatial point cluster matched by each of the tasks includes: Obtain demand feature data for multiple tasks; Based on the demand characteristic data, determining target points corresponding to a plurality of the tasks in the vector space; Matching the plurality of target points with each of the spatial point clusters to determine the spatial point cluster to which each of the target points belongs; Based on the spatial point cluster to which each of the target points belongs, a spatial point cluster matched by each of the tasks is determined.

4. The method according to claim 3, characterized in that Matching the plurality of target points with the respective spatial point clusters to determine the spatial point cluster to which each target point belongs includes: Determining the centroid of each of the spatial point clusters; For each of the target points, respectively calculating the distance between the target point and the centroid of each of the spatial point clusters; The spatial point cluster to which the target point belongs is determined based on the spatial point cluster whose centroid is the smallest and the target point.

5. The method according to claim 1, characterized in that Selecting a target spatial point that meets the operation requirements of the task from a spatial point cluster that matches the task, including: If the running requirement of the task is high-performance AI computing, a target spatial point whose GPU performance of the corresponding node meets the preset GPU performance standard and whose network delay meets the preset delay standard is selected from the spatial point cluster matching the task.

6. The method according to claim 1, characterized in that From the spatial point clusters matching the task, select the target spatial point that meets the running requirements of the task, If the task requires long-term continuous operation, a target spatial point whose online time of the corresponding node meets the preset online time standard and whose network bandwidth meets the preset bandwidth standard is selected from the spatial point cluster matching the task.

7. The method according to claim 1, characterized in that The method further comprises: If the computing power provided by the spatial point cluster matching the task does not meet the computing power requirements of the task, a spare spatial point cluster with a matching degree second only to the spatial point cluster matching the task is selected from the remaining spatial point clusters to provide additional computing power support for the task; the matching degree is determined based on the distance between the target point corresponding to the task and the center of mass of the spare spatial point cluster.

8. An intelligent scheduling device for distributed idle computing power clusters, characterized in that: include: A data collection unit, used to obtain feature data uploaded by multiple nodes in real time; A spatial point updating unit, used for determining spatial points corresponding to a plurality of the nodes in the vector space based on the feature data; A spatial point clustering unit, used to cluster the plurality of spatial points using a preset clustering algorithm to obtain at least one spatial point cluster; the spatial point cluster includes a plurality of spatial points having similarities; A task matching unit, used for matching a plurality of tasks with each of the spatial point clusters to obtain a spatial point cluster matched by each of the tasks; A spatial point screening unit, configured to select, for each of the tasks, a target spatial point that meets the running requirements of the task from a spatial point cluster that matches the task; A task monitoring unit, used to trigger the target node corresponding to the target spatial point to execute the task, and to monitor the progress of the task and the resource usage of the target node in real time; The task monitoring unit is also used to select a new target spatial point that meets the running requirements of the task from the spatial point cluster matching the task when the progress of the task is interrupted and / or the resource usage of the target node is resource-constrained, and trigger the new target node corresponding to the new target spatial point to execute the task.

9. A storage medium, characterized in that: The storage medium includes a stored program, wherein the program, when executed by a processor, executes the intelligent scheduling method for a distributed idle computing power cluster as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: include: processor, memory, and bus; The processor is connected to the memory via the bus; The memory is used to store programs, and the processor is used to run programs, wherein the program, when run by the processor, executes the intelligent scheduling method for distributed idle computing power clusters described in any one of claims 1-7.