Heterogeneous resource intelligent scheduling method and system for large-scale model training and push-pull integrated machine

Through the intelligent scheduling method of heterogeneous resources of the large-scale model training and pushing all-in-one machine, the reinforcement learning model is used to dynamically adjust resource allocation, which solves the problem of task delay and failure of traditional scheduling algorithms under dynamic loads, and realizes efficient resource utilization and stable task execution.

CN119781991BActive Publication Date: 2025-09-26BEIJING WANGZHI TIANYUAN BIG DATA TECH CO LTD +1

Patent Information

Application Number
CN202510278768.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-09-26
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Traditional resource scheduling algorithms are unable to adjust resource allocation in a timely manner when faced with dynamic load changes, resulting in task delays or failures. Especially in a multi-task parallel execution environment, task allocation conflicts occur frequently.

Method used

A heterogeneous resource intelligent scheduling method using a large-scale model training and pushing machine is adopted. By collecting task lists, the maintenance task list of resource nodes is initialized, and the pre-trained reinforcement learning model is used to screen the minimum cost features and slack variables, dynamically adjust resource allocation, optimize scheduling load costs, and realize resource redistribution.

Benefits of technology

Effectively prevent task delays or failures, ensure timely task completion, improve resource utilization efficiency and task execution reliability, reduce allocation conflicts, and dynamically respond to load changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781991B_ABST
    Figure CN119781991B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for intelligent scheduling of heterogeneous resources of a large-model training and pushing all-in-one machine, which initializes the maintenance task list of each heterogeneous resource node in the training and pushing all-in-one machine, and generates a cost matrix for resource scheduling of maintenance tasks based on all maintenance task lists; the minimum cost characteristics of each heterogeneous resource node are screened out from the scheduling cost of the resource scheduling cost matrix, and then the slack variables of each heterogeneous resource node are determined when the target task has an allocation conflict; the scheduling load cost of the target task for each heterogeneous resource node is determined based on each minimum cost characteristic and the load characteristic of the target task; the target task is heterogeneously allocated resources using each scheduling load cost, and each slack variable is used to reallocate resources to the target task each time a task allocation conflict occurs in the heterogeneous resource allocation. Based on the above scheme, resource reallocation when a task allocation conflict occurs in resource scheduling can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of resource scheduling, and more specifically, to a method and system for intelligently scheduling heterogeneous resources of a large-scale model training and promotion integrated machine. Background Art

[0002] With the widespread application of large models, the demand for training and inference has increased dramatically, and the training and inference all-in-one machine has come into being. The training and inference all-in-one machine integrates training and inference functions on the same hardware platform, and realizes efficient task allocation and resource management through intelligent scheduling and resource optimization. The integrated design not only improves computing efficiency, but also reduces energy consumption and operation and maintenance costs.

[0003] Traditional resource scheduling algorithms usually schedule tasks based on static load predictions and pre-set resource allocation strategies. They have limitations when dealing with dynamic load changes. The computing requirements of tasks and the availability of resources are constantly changing. Especially in an environment where multiple tasks are executed in parallel, the task load may suddenly increase or decrease, and the resource consumption of nodes also changes at any time. Traditional scheduling algorithms usually do not take these dynamic factors into consideration and rely on pre-calculated scheduling strategies. They cannot make real-time adjustments based on the actual resource load of the current node. When node resources are insufficient or load conflicts occur, tasks will be delayed or unable to execute due to the lack of timely adjustment of resource allocation. Tasks may not be completed within the scheduled time or may even fail due to resource conflicts. Therefore, how to achieve resource reallocation when task allocation conflicts occur in resource scheduling has become a difficult problem faced by the industry. Summary of the Invention

[0004] The present application provides a method and system for intelligent scheduling of heterogeneous resources of a large-scale model training and pushing integrated machine, which can realize resource reallocation when task allocation conflicts occur in resource scheduling.

[0005] In a first aspect, the present application provides a method for intelligently scheduling heterogeneous resources of a large-scale model training and promotion machine, comprising:

[0006] Collecting a list of tasks to be assigned in the training and pushing integrated machine, and selecting a task to be assigned from the list of tasks to be assigned as a target task;

[0007] Initialize the maintenance task list for each heterogeneous resource node in the training and push machine, and generate the cost matrix for resource scheduling when the training and push machine performs maintenance tasks based on all the maintenance task lists;

[0008] Based on the pre-trained reinforcement learning model, the minimum cost feature of each heterogeneous resource node is screened from the cost matrix of the resource scheduling, and then the slack variables of each heterogeneous resource node are determined when the allocation conflict of the target task occurs;

[0009] Determine the scheduling load cost of the target task on each heterogeneous resource node based on each minimum cost feature and the load feature of the target task;

[0010] The training-pushing machine is used to perform heterogeneous resource allocation for the target task in combination with the various scheduling load costs. When a task allocation conflict occurs in each heterogeneous resource allocation, the various slack variables are used to reallocate resources to the target task until the resource allocation of each task to be allocated in the task list is completed.

[0011] In some embodiments, generating a cost matrix for resource scheduling when the training and push-pushing machine performs maintenance tasks based on all maintenance task lists specifically includes:

[0012] For each heterogeneous resource node, obtain the scheduling cost of each maintenance task of the heterogeneous resource node from the maintenance task list of the heterogeneous resource node;

[0013] Determine the scheduling cost sequence of heterogeneous resource nodes for maintenance tasks through all scheduling costs, and then obtain the scheduling cost sequence of each heterogeneous resource node for maintenance tasks;

[0014] Determine the cost matrix of resource scheduling when the training and push machine performs maintenance tasks based on all scheduling cost sequences.

[0015] In some embodiments, screening out the minimum cost features of each heterogeneous resource node from the cost matrix of the resource scheduling based on the pre-trained reinforcement learning model specifically includes:

[0016] For each heterogeneous resource node, screening out a scheduling cost sequence of the heterogeneous resource node from the cost matrix of resource scheduling;

[0017] Extracting cost markers for each maintenance task in the scheduling cost sequence based on a pre-trained reinforcement learning model;

[0018] The minimum value of all cost flags is taken as the minimum cost feature of the heterogeneous resource node, and then the minimum cost feature of each heterogeneous resource node is obtained.

[0019] In some embodiments, determining the slack variables of each heterogeneous resource node when a target task allocation conflict occurs specifically includes:

[0020] For each heterogeneous resource node, obtain the minimum cost feature of the heterogeneous resource node;

[0021] Determine the execution cost of the target task on heterogeneous resource nodes;

[0022] The slack variables of the heterogeneous resource nodes when the target task allocation conflict occurs are determined according to the execution cost and the minimum cost feature, and the slack variables of each heterogeneous resource node when the target task allocation conflict occurs are further obtained.

[0023] In some embodiments, determining the scheduling load cost of the target task for each heterogeneous resource node based on each minimum cost feature and the load feature of the target task specifically includes:

[0024] For each heterogeneous resource node, obtain the idle computing power of the heterogeneous resource node;

[0025] Determining the scheduling load of the target task on the heterogeneous resource nodes according to the idle computing power and the load characteristics of the target task;

[0026] The scheduling load is optimized according to the minimum cost characteristics of the heterogeneous resource nodes to obtain the scheduling load cost of the target task for the heterogeneous resource nodes, and then the scheduling load cost of the target task for each heterogeneous resource node is obtained.

[0027] In some embodiments, using the integrated training and pushing machine in combination with various scheduling load costs to perform heterogeneous resource allocation for target tasks specifically includes:

[0028] For each heterogeneous resource node, the scheduling load cost of the target task on the heterogeneous resource node is used as the allocated load of the training and pushing machine;

[0029] Determine the load ratio of the target task in the resource heterogeneous nodes based on the distributed load in the training and push machine, and then obtain the load ratio of the target task in each resource heterogeneous node;

[0030] The execution nodes of the target task are selected from all resource heterogeneous nodes through all load ratios.

[0031] In some embodiments, using each slack variable to reallocate resources to the target task specifically includes:

[0032] A heterogeneous resource node with the largest slack variable is selected from the heterogeneous resource nodes where task allocation conflicts occur as a reallocated resource node, and the reallocated resource node is used to execute the target task.

[0033] In a second aspect, the present application provides a heterogeneous resource intelligent scheduling system for a large-scale model training and promotion all-in-one machine, comprising:

[0034] A collection module is used to collect a list of tasks to be assigned in the training and pushing integrated machine, and select a task to be assigned from the list of tasks to be assigned as a target task;

[0035] The processing module is used to initialize the maintenance task list of each heterogeneous resource node in the training and push integration machine, and generate the cost matrix of resource scheduling when the training and push integration machine performs maintenance tasks based on all the maintenance task lists;

[0036] The processing module is further configured to filter out the minimum cost feature of each heterogeneous resource node from the cost matrix of the resource scheduling based on a pre-trained reinforcement learning model, and then determine the slack variables of each heterogeneous resource node when an allocation conflict occurs in the target task;

[0037] The processing module is further configured to determine the scheduling load cost of the target task for each heterogeneous resource node based on each minimum cost feature and the load feature of the target task;

[0038] The execution module is used to use the training and pushing machine in combination with each scheduling load cost to perform heterogeneous resource allocation for the target task. When a task allocation conflict occurs in each heterogeneous resource allocation, each slack variable is used to reallocate resources to the target task until the resource allocation of each task to be allocated in the task list is completed.

[0039] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-mentioned heterogeneous resource intelligent scheduling method of the large-model training and push all-in-one machine.

[0040] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions or codes. When the instructions or codes are run on a computer, the computer implements the above-mentioned intelligent scheduling method for heterogeneous resources of the large-model training and push integrated machine.

[0041] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0042] The present application provides a heterogeneous resource intelligent scheduling method and system for a large-model training and pushing integrated machine, which collects a list of tasks to be assigned in the training and pushing integrated machine, selects a task to be assigned from the list of tasks to be assigned as a target task; initializes a maintenance task list for each heterogeneous resource node in the training and pushing integrated machine, and generates a cost matrix for resource scheduling when the training and pushing integrated machine performs maintenance tasks based on all maintenance task lists; based on a pre-trained reinforcement learning model, the minimum cost feature of each heterogeneous resource node is screened from the cost matrix of resource scheduling, and then determines the slack variables of each heterogeneous resource node when an allocation conflict occurs in the target task; determines the scheduling load cost of the target task for each heterogeneous resource node based on each minimum cost feature and the load feature of the target task; uses the training and pushing integrated machine in combination with each scheduling load cost to perform heterogeneous resource allocation on the target task, and when a task allocation conflict occurs in each heterogeneous resource allocation, uses each slack variable to reallocate resources to the target task until the resource allocation of each task to be assigned in the list of tasks to be assigned is completed.

[0043] It can be seen that in this application, the training and pushing integrated machine is used to perform heterogeneous resource allocation for the target task in combination with each scheduling load cost. When a task allocation conflict occurs in each heterogeneous resource allocation, each slack variable is used to reallocate resources for the target task until the resource allocation of each task to be allocated in the task list to be allocated is completed; first, the slack variable reflects the adjustable ability of the heterogeneous resource node when a task allocation conflict occurs, that is, the node can appropriately adjust the task load or scheduling strategy without exceeding its carrying capacity. Through the pre-trained reinforcement learning model, the slack variable can accurately quantify the resource adjustment potential of each node, and provide a basis for task allocation decision-making. When a task conflict occurs, the slack variable can be used as an adjustment factor to dynamically adjust The resource allocation of target tasks mitigates the risk of overloading heterogeneous resource nodes or uneven resource allocation. By continuously optimizing the resource allocation process for conflicting tasks, task execution delays or failures are effectively prevented, ensuring the stability of the training and push machine and the timely completion of tasks. Slack variables provide the training and push machine with flexible scheduling space, enabling flexible adjustments based on real-time load conditions to maximize resource utilization efficiency and task execution reliability. The scheduling load cost then helps assess the load-bearing capacity of each resource node, ensuring that tasks are allocated to nodes with low resource consumption and optimized loads. Real-time calculation of the scheduling load cost enables the training and push machine to dynamically respond to changing demand for different tasks, avoiding node overload or resource waste. By continuously optimizing the scheduling load cost, the execution burden of the target task on the resource node can be precisely adjusted, thereby reducing the occurrence of allocation conflicts. In summary, the above scheme enables resource reallocation when task allocation conflicts occur during resource scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0045] Figure 1 This is an exemplary flow chart of a method for intelligently scheduling heterogeneous resources of a large-scale training and push-pull machine according to some embodiments of the present application;

[0046] Figure 2 This is an execution logic diagram of the training and pushing all-in-one machine shown in some embodiments of the present application;

[0047] Figure 3 is a schematic diagram of a process for determining a dispatch load cost according to some embodiments of the present application;

[0048] Figure 4This is a schematic diagram of the structure of a heterogeneous resource intelligent scheduling system for a large-scale training and promotion all-in-one machine according to some embodiments of the present application;

[0049] Figure 5 It is a structural diagram of a computer device for implementing a method for intelligent scheduling of heterogeneous resources of a large-scale model training and promotion all-in-one machine as shown in some embodiments of the present application. DETAILED DESCRIPTION

[0050] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0051] refer to Figure 1 This figure is an exemplary flow chart of a method for intelligently scheduling heterogeneous resources of a large-scale training and push-pull machine according to some embodiments of the present application. The method for intelligently scheduling heterogeneous resources of a large-scale training and push-pull machine mainly includes the following steps:

[0052] In step 101, a list of tasks to be assigned in the training and pushing integrated machine is collected, and a task to be assigned is selected from the list of tasks to be assigned as a target task.

[0053] It should be noted that, in this application, the list of tasks to be assigned represents a set of tasks that have not yet been assigned to a specific heterogeneous resource node for execution, and the target task represents the task currently selected as the assignment object in the list of tasks to be assigned; in specific implementation, all the tasks to be assigned of the training and pushing integrated machine at the current moment are collected from the task management module of the training and pushing integrated machine, and then the collection of all the tasks to be assigned is used as a list of assigned tasks, and a task to be assigned is randomly selected from the list of tasks to be assigned as the target task.

[0054] In some embodiments, reference Figure 2 The figure shows the execution logic of the integrated training and push system shown in some embodiments of this application. The integrated training and push system is divided into three main parts: the command and control group, the participating group, and the information support platform. The command and control group is composed of decision-makers and a supervisory group. The decision-makers provide guidance and supervision to the entire system through the supervisory group. The participating group includes the control group, the participating group leader, the game staff, and the participating experts. They conduct confrontation simulations through the game center and the nervous system to improve the efficiency and effectiveness of the participation.

[0055] The information support platform is composed of a situation system, a support system, and a communication system, providing the necessary information and technical support for the entire system. The control group connects with multiple units through the nervous system to achieve control and coordination of the participating activities. The various roles within the participating group exchange information and collaborate through the nervous system to ensure the smooth progress of the participating activities. The entire system achieves efficient command, control, and participation simulation through the combination of man and machine.

[0056] In step 102, a maintenance task list of each heterogeneous resource node in the training and pushing machine is initialized, and a cost matrix for resource scheduling when the training and pushing machine performs maintenance tasks is generated based on all the maintenance task lists.

[0057] In some embodiments, initializing the maintenance task list of each heterogeneous resource node in the training and pushing integrated machine can be implemented in the following manner, namely: identifying all available heterogeneous resource nodes in the training and pushing integrated machine, and for each heterogeneous resource node, obtaining all tasks that need to be maintained from the task module of the heterogeneous resource node as maintenance tasks, which include task serial number, resource occupancy and task progress, and then arranging all maintenance tasks according to the task serial number as a maintenance task list of the heterogeneous resource node; it should be noted that in this application, the maintenance task list represents a set of tasks that need to be executed in the heterogeneous resource node to ensure the normal operation of the system, maintain equipment, repair faults, etc., which usually includes the task priority, resource requirements and scheduled execution time.

[0058] In some embodiments, generating a cost matrix for resource scheduling when the training and push-pull machine performs maintenance tasks based on a list of all maintenance tasks can be achieved by using the following steps:

[0059] For each heterogeneous resource node, obtain the scheduling cost of each maintenance task of the heterogeneous resource node from the maintenance task list of the heterogeneous resource node;

[0060] Determine the scheduling cost sequence of heterogeneous resource nodes for maintenance tasks through all scheduling costs, and then obtain the scheduling cost sequence of each heterogeneous resource node for maintenance tasks;

[0061] Determine the cost matrix of resource scheduling when the training and push machine performs maintenance tasks based on all scheduling cost sequences.

[0062] It should be noted that, in this application, the cost matrix of resource scheduling represents the scheduling cost relationship between multiple heterogeneous resource nodes and multiple tasks; the scheduling cost represents the total resource consumption overhead of computing, storage and communication generated when a task is assigned to a heterogeneous resource node; the scheduling cost sequence represents the record sequence of the scheduling cost of maintaining the task on each heterogeneous resource node.

[0063] In the specific implementation, first, for each heterogeneous resource node, the scheduling cost of the heterogeneous resource node for each maintenance task is obtained from the maintenance task list of the heterogeneous resource node. This can be achieved in the following way, namely: for each heterogeneous resource node, the resource occupancy of the heterogeneous resource node for each maintenance task is obtained from the maintenance task list of the heterogeneous resource node as the scheduling cost; then, the scheduling cost sequence of the heterogeneous resource node for the maintenance task is determined through all the scheduling costs, and then the scheduling cost sequence of each heterogeneous resource node for the maintenance task is obtained. This can be achieved in the following way, namely: all scheduling costs are arranged according to the execution order of the maintenance task in the training-push integrated machine to obtain the scheduling cost of the heterogeneous resource node for the maintenance task. Sequence, the scheduling cost sequence of each heterogeneous resource node for the maintenance task can be obtained by the above method; finally, the cost matrix of resource scheduling when the training-pushing integrated machine performs maintenance tasks is determined according to all the scheduling cost sequences, which can be implemented in the following way, namely: the cost matrix of resource scheduling is a matrix corresponding to all maintenance task lists, and each element in the matrix represents the resource occupancy of the heterogeneous resource node for the maintenance task as the scheduling cost. If there is no scheduling relationship between the heterogeneous resource node and the maintenance task, the value of this element in the cost matrix of resource scheduling is 0, and each scheduling cost sequence is arranged according to the arrangement order of the heterogeneous resource nodes in the training-pushing integrated machine as the cost matrix of resource scheduling when the training-pushing integrated machine performs maintenance tasks.

[0064] In step 103, the minimum cost feature of each heterogeneous resource node is screened from the cost matrix of the resource scheduling based on the pre-trained reinforcement learning model, and then the slack variables of each heterogeneous resource node when the target task allocation conflict occurs are determined.

[0065] In some embodiments, screening out the minimum cost feature of each heterogeneous resource node from the cost matrix of the resource scheduling based on the pre-trained reinforcement learning model can be achieved by the following steps:

[0066] For each heterogeneous resource node, screening out a scheduling cost sequence of the heterogeneous resource node from the cost matrix of resource scheduling;

[0067] Extracting cost markers for each maintenance task in the scheduling cost sequence based on a pre-trained reinforcement learning model;

[0068] The minimum value of all cost flags is taken as the minimum cost feature of the heterogeneous resource node, and then the minimum cost feature of each heterogeneous resource node is obtained.

[0069] In a specific implementation, first, for each heterogeneous resource node, the scheduling cost sequence of the heterogeneous resource node is screened out from the cost matrix of the resource scheduling, which can be implemented in the following manner: for each heterogeneous resource node, the scheduling cost sequence of the heterogeneous resource node in the cost matrix of the resource scheduling is obtained; then, based on the pre-trained reinforcement learning model, the cost mark of each maintenance task in the scheduling cost sequence is extracted, which can be implemented in the following manner: initializing a scheduling optimization model based on reinforcement learning, using the scheduling cost sequence as the state space in the scheduling optimization model, designing a reward function to optimize the cost mark of each maintenance task, the reward function can be set to the negative value of the scheduling cost, using the scheduling optimization model to infer the cost mark of each maintenance task in the scheduling cost sequence, and thus using the quantized value of the inference result as the cost mark of each maintenance task; finally, using the minimum value of all cost marks as the minimum cost feature of the heterogeneous resource node, thereby obtaining the minimum cost feature of each heterogeneous resource node, which can be implemented in the following manner: using the minimum value of all cost marks as the minimum cost feature of the heterogeneous resource node, and the minimum cost feature of each heterogeneous resource node can be obtained by the above method.

[0070] It should be noted that in this application, the minimum cost feature represents the lowest scheduling cost of heterogeneous resource nodes when performing specific tasks during the task scheduling process; the cost indicator is a comprehensive indicator reflecting node resource consumption and task execution overhead; the scheduling optimization model is a mathematical model based on reinforcement learning. This scheduling optimization model uses the scheduling cost sequence as a state space and uses the intelligent agent to learn scheduling strategies to minimize the task scheduling cost. The core technical principle of this scheduling optimization model is the deep Q-network (DQN) in reinforcement learning. The intelligent agent continuously optimizes its decision-making strategy by interacting with the environment. In this scheduling optimization model, the scheduling cost sequence is used as the state space to reflect the various resource consumption and task allocation situations in task scheduling. The reward function is designed to be the negative value of the scheduling cost to incentivize the intelligent agent to choose a scheduling solution that reduces cost. Through the training process, the intelligent agent can gain experience from environmental feedback and use reasoning capabilities to optimize future task scheduling and predict the cost indicator of each maintenance task. The quantified value of the reasoning result reflects the scheduling cost of each maintenance task. In this way, scheduling decisions can be effectively optimized, the overall resource allocation efficiency can be improved, the task execution cost can be reduced, and the resource utilization of the system can be improved.

[0071] In some embodiments, determining the slack variables of each heterogeneous resource node when a target task allocation conflict occurs may be achieved by using the following steps:

[0072] For each heterogeneous resource node, obtain the minimum cost feature of the heterogeneous resource node;

[0073] Determine the execution cost of the target task on heterogeneous resource nodes;

[0074] The slack variables of the heterogeneous resource nodes when the target task allocation conflict occurs are determined according to the execution cost and the minimum cost feature, and the slack variables of each heterogeneous resource node when the target task allocation conflict occurs are further obtained.

[0075] It should be noted that, in this application, the slack variable is an indicator to measure the adjustable ability of each heterogeneous resource node in the case of allocation conflict; in specific implementation, first, for each heterogeneous resource node, the minimum cost characteristics of the heterogeneous resource node are obtained; then, the execution cost of the target task on the heterogeneous resource node can be determined in the following way, namely: the task execution time, computing resource consumption and data transmission delay of the target task on the heterogeneous resource node can be evaluated by using a simulation algorithm, so as to calculate the weighted sum of the task execution time, computing resource consumption and data transmission delay as the execution cost of the target task on the heterogeneous resource node, wherein the weight parameters of task execution time and computing resource consumption are 1 and 2. The number can be adjusted through reinforcement learning or historical task data, and the execution cost represents the total amount of computing and storage resource consumption required to complete the execution of the target task on a specific resource node; finally, the slack variable of the heterogeneous resource node when the target task has an allocation conflict is determined based on the execution cost and the minimum cost feature, and then the slack variable of each heterogeneous resource node when the target task has an allocation conflict is obtained. This can be achieved in the following way, namely: the absolute value of the difference between the execution cost and the minimum cost feature is used as the slack variable of the heterogeneous resource node when the target task has an allocation conflict. The slack variable of each heterogeneous resource node when the target task has an allocation conflict can be obtained by the above method.

[0076] In step 104, the scheduling load cost of the target task on each heterogeneous resource node is determined according to each minimum cost feature and the load feature of the target task.

[0077] In some embodiments, the scheduling load cost of the target task for each heterogeneous resource node is determined based on each minimum cost feature and the load feature of the target task, referring to Figure 3 The figure is a schematic diagram of a process for determining the dispatching load cost in some embodiments of the present application. In this embodiment, the dispatching load cost can be determined by the following steps:

[0078] In step 1041, for each heterogeneous resource node, the idle computing power of the heterogeneous resource node is obtained;

[0079] In step 1042, the scheduling load of the target task on the heterogeneous resource nodes is determined based on the idle computing power and the load characteristics of the target task;

[0080] In step 1043, the scheduling load is optimized according to the minimum cost characteristics of the heterogeneous resource nodes to obtain the scheduling load cost of the target task for the heterogeneous resource nodes, and then obtain the scheduling load cost of the target task for each heterogeneous resource node.

[0081] In the specific implementation, first, for each heterogeneous resource node, obtaining the idle computing power of the heterogeneous resource node can be achieved in the following manner, namely: for each heterogeneous resource node, obtaining the total computing power and used computing power of the heterogeneous resource node from the resource module of the training and pushing integrated machine, and taking the difference between the total computing power and the used computing power as the idle computing power of the heterogeneous resource node; then, determining the scheduling load of the target task on the heterogeneous resource node according to the idle computing power and the load characteristics of the target task can be achieved in the following manner, namely: using a load prediction algorithm based on a long-short-term memory network to evaluate the load characteristics of the target task, and taking the ratio of the load characteristics of the target task to the idle computing power as the scheduling load of the target task on the heterogeneous resource node, wherein the load prediction algorithm based on the long-short-term memory network can predict the load characteristics of the target task through time series analysis, and use its memory capacity to capture long-term dependencies, Accurately estimate the task load change, so as to provide accurate input data for the load optimization model to optimize the scheduling cost; finally, perform load optimization on the scheduling load according to the minimum cost characteristics of the heterogeneous resource nodes to obtain the scheduling load cost of the target task for the heterogeneous resource nodes, and then obtain the scheduling load cost of the target task for each heterogeneous resource node. This can be achieved in the following way, namely: initialize a load optimization model based on a support vector machine, use the minimum cost characteristics of the heterogeneous resource nodes as the feature vector in the load optimization model, use the scheduling load as the initial optimization value in the load optimization model, use the load optimization model to perform regression analysis on the scheduling load cost function, and use the result of the regression analysis as the scheduling load cost of the target task for the heterogeneous resource nodes. The scheduling load cost of the target task for each heterogeneous resource node can be obtained in the above way.

[0082] It should be noted that, in this application, the scheduling load cost refers to the scheduling resource consumption caused by the computing load distribution when the target task is executed on a specific resource node; the idle computing power refers to the computing power of the heterogeneous resource node that can be used to execute new tasks at the current moment; the scheduling load refers to the degree of occupancy of the computing, storage and bandwidth resources of a specific resource node when the target task is executed on the node; the load optimization model is a mathematical model based on support vector machine (SVM), which uses the minimum cost feature of the heterogeneous resource node as the feature vector, takes the scheduling load as the initial optimization value, and predicts the scheduling load of the target task on each node through regression analysis. Cost, the SVM regression model captures the nonlinear relationship between task load and resource node performance by mapping input features (such as minimum cost features and scheduling load) to a high-dimensional space. The load optimization model trains a best-fit hyperplane by minimizing the regression error to predict the cost of unknown task scheduling. The result of the regression analysis is used as the scheduling load cost of the target task for each node, which can effectively reflect the task's demand for resources and the node's load-bearing capacity. Support vector machines can provide an accurate and efficient way to optimize task scheduling, balance resource load, reduce computing overhead, and improve the overall system's resource utilization and task execution efficiency.

[0083] In step 105, the training-pushing machine is used in combination with each scheduling load cost to perform heterogeneous resource allocation for the target task. When a task allocation conflict occurs in each heterogeneous resource allocation, each slack variable is used to reallocate resources for the target task until the resource allocation of each task to be allocated in the task list to be allocated is completed.

[0084] In some embodiments, using the training-pushing machine in combination with various scheduling load costs to perform heterogeneous resource allocation for target tasks can be achieved using the following steps:

[0085] For each heterogeneous resource node, the scheduling load cost of the target task on the heterogeneous resource node is used as the allocated load of the training and pushing machine;

[0086] Determine the load ratio of the target task in the resource heterogeneous nodes based on the distributed load in the training and push machine, and then obtain the load ratio of the target task in each resource heterogeneous node;

[0087] The execution nodes of the target task are selected from all resource heterogeneous nodes through all load ratios.

[0088] It should be noted that in this application, the execution node refers to the heterogeneous resource node that actually undertakes the execution of the target task; the allocated load refers to the consumption of computing, storage and network resources allocated to the target task on the heterogeneous resource node; the load ratio refers to the proportion of computing resources occupied by the target task on each heterogeneous resource node.

[0089] In the specific implementation, first, for each resource heterogeneous node, the scheduling load cost of the target task on the heterogeneous resource node is used as the allocated load of the training-pushing machine; then, the load ratio of the target task in the resource heterogeneous node is determined according to the allocated load in the training-pushing machine, and then the load ratio of the target task in each resource heterogeneous node is obtained. This can be achieved in the following way, namely: the ratio of the allocated load in the training-pushing machine to the idle computing power of the heterogeneous resource node is used as the load ratio of the target task in the resource heterogeneous node. The load ratio of the target task in each resource heterogeneous node can be obtained by the above method; finally, the execution node of the target task is screened out from all resource heterogeneous nodes through all load ratios. This can be achieved in the following way, namely: the resource heterogeneous nodes with load ratios lower than the preset load threshold in all are used as the execution nodes of the target task. If there are multiple execution nodes, it means that there is a task allocation conflict, and then the various slack variables are continued to be used to reallocate resources for the target task.

[0090] In some embodiments, resource reallocation to target tasks using slack variables may be achieved by the following steps:

[0091] A heterogeneous resource node with the largest slack variable is selected from the heterogeneous resource nodes where task allocation conflicts occur as a reallocated resource node, and the reallocated resource node is used to execute the target task.

[0092] It should be noted that by screening the heterogeneous resource nodes with the largest slack variables for resource reallocation, the flexibility of task scheduling and the optimal utilization of computing resources are achieved. The slack variables measure the adjustable ability of each node in the case of allocation conflicts. Prioritizing the nodes with the largest slack variables helps to balance resource loads, reduce scheduling conflicts and task failure rates, and dynamically adapt to changes in resource status, ensuring the stability and continuity of task execution, thereby improving the overall computing efficiency and throughput of the system.

[0093] In addition, in another aspect of the present application, in some embodiments, the present application provides a heterogeneous resource intelligent scheduling system for a large model training and pushing integrated machine, referring to Figure 4 This figure is a schematic diagram of the structure of a heterogeneous resource intelligent scheduling system for a large-scale model training and push-integrated machine according to some embodiments of the present application. The heterogeneous resource intelligent scheduling system for a large-scale model training and push-integrated machine includes: a collection module 201, a processing module 202, and an execution module 203, which are described as follows:

[0094] The collection module 201 in this application is mainly used to collect a list of tasks to be assigned in the training and pushing integrated machine, and select a task to be assigned from the list of tasks to be assigned as a target task;

[0095] Processing module 202, in this application, is used to initialize a maintenance task list for each heterogeneous resource node in the training and pushing machine, and generate a cost matrix for resource scheduling when the training and pushing machine performs maintenance tasks based on all the maintenance task lists;

[0096] It should be noted that the processing module 202 is further configured to filter out the minimum cost feature of each heterogeneous resource node from the cost matrix of the resource scheduling based on a pre-trained reinforcement learning model, and further determine the slack variables of each heterogeneous resource node when an allocation conflict occurs in the target task;

[0097] In addition, the processing module 202 is further configured to determine the scheduling load cost of the target task on each heterogeneous resource node according to each minimum cost feature and the load feature of the target task;

[0098] Execution module 203. In this application, execution module 203 is mainly used to use the training and pushing integrated machine in combination with each scheduling load cost to perform heterogeneous resource allocation on the target task. When a task allocation conflict occurs in each heterogeneous resource allocation, each slack variable is used to reallocate resources to the target task until the resource allocation of each task to be allocated in the task list to be allocated is completed.

[0099] The above describes in detail an example of a method and system for intelligent scheduling of heterogeneous resources of a large-scale training and pushing integrated machine provided in an embodiment of the present application. It can be understood that in order to realize the above functions, the corresponding device includes a hardware structure and / or software module corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware-driven or computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0100] In some embodiments, the present application also provides a computer device, which includes a memory and a processor, the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that the computer device executes the above-mentioned heterogeneous resource intelligent scheduling method of the large-model training and push all-in-one machine.

[0101] In some embodiments, reference Figure 5The dotted line in the figure indicates that the unit or module is optional. The figure is a structural diagram of a computer device for implementing a method for intelligently scheduling heterogeneous resources of a large-scale model training and pushing integrated machine according to an embodiment of the present application. The method for intelligently scheduling heterogeneous resources of a large-scale model training and pushing integrated machine described in the above embodiment can be Figure 5 The computer device shown in the figure is implemented, and the computer device includes at least one processor 301, a memory 302 and at least one communication unit 305. The computer device can be a terminal device, a server or a chip.

[0102] The processor 301 may be a general-purpose processor or a dedicated processor. For example, the processor 301 may be a central processing unit (CPU), which may be used to control the computer device, execute software programs, and process data from the software programs. The computer device may also include a communication unit 305 for inputting (receiving) and outputting (transmitting) signals.

[0103] For example, the computer device may be a chip, the communication unit 305 may be an input and / or output circuit of the chip, or the communication unit 305 may be a communication interface of the chip, and the chip may be a component of a terminal device, a network device, or other device.

[0104] For another example, the computer device may be a terminal device or a server, and the communication unit 305 may be a transceiver of the terminal device or the server, or the communication unit 305 may be a transceiver circuit of the terminal device or the server.

[0105] The computer device may include one or more memories 302, on which a program 304 is stored. The program 304 can be executed by the processor 301 to generate instructions 303, so that the processor 301 executes the method described in the above method embodiment according to the instructions 303. Optionally, data (such as a target audit model) can also be stored in the memory 302. Optionally, the processor 301 can also read data stored in the memory 302. The data can be stored at the same storage address as the program 304, or at a different storage address from the program 304.

[0106] The processor 301 and the memory 302 may be provided separately or integrated together, for example, integrated on a system on chip (SOC) of a terminal device.

[0107] It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or software-based instructions in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.

[0108] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] For example, in some embodiments, the present application also provides a computer-readable storage medium, which stores instructions or codes. When the instructions or codes are run on a computer, the computer implements the above-mentioned intelligent scheduling method for heterogeneous resources of the large-model training and push integrated machine.

[0110] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0111] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for intelligently scheduling heterogeneous resources of a large-scale training and pushing machine, characterized in that: The steps include: Collecting a list of tasks to be assigned in the training and pushing integrated machine, and selecting a task to be assigned from the list of tasks to be assigned as a target task; Initialize the maintenance task list for each heterogeneous resource node in the training and push machine, and generate a cost matrix for resource scheduling when the training and push machine performs maintenance tasks based on all the maintenance task lists. The cost matrix represents the scheduling cost relationship between multiple heterogeneous resource nodes and multiple tasks. Based on the pre-trained reinforcement learning model, the minimum cost feature of each heterogeneous resource node is screened from the cost matrix of the resource scheduling, and then the slack variable of each heterogeneous resource node is determined when the allocation conflict of the target task occurs. The slack variable is an indicator of the adjustable ability of each heterogeneous resource node in the case of allocation conflict; Determine the scheduling load cost of the target task on each heterogeneous resource node based on each minimum cost feature and the load feature of the target task; Use the training and push machine to combine the various scheduling load costs to perform heterogeneous resource allocation for the target task. When a task allocation conflict occurs during each heterogeneous resource allocation, use various slack variables to reallocate resources to the target task until all resources in the task list are allocated. The slack variables for determining each heterogeneous resource node when a target task allocation conflict occurs specifically include: For each heterogeneous resource node, obtain the minimum cost feature of the heterogeneous resource node; Determine the execution cost of the target task on heterogeneous resource nodes; The slack variables of the heterogeneous resource nodes when the target task allocation conflict occurs are determined according to the execution cost and the minimum cost feature, and the slack variables of each heterogeneous resource node when the target task allocation conflict occurs are further obtained.

2. The method according to claim 1, wherein The cost matrix for resource scheduling when the training and push machine performs maintenance tasks is generated based on the maintenance task list. Specifically, it includes: For each heterogeneous resource node, obtain the scheduling cost of each maintenance task of the heterogeneous resource node from the maintenance task list of the heterogeneous resource node; Determine the scheduling cost sequence of heterogeneous resource nodes for maintenance tasks through all scheduling costs, and then obtain the scheduling cost sequence of each heterogeneous resource node for maintenance tasks; Determine the cost matrix of resource scheduling when the training and push machine performs maintenance tasks based on all scheduling cost sequences.

3. The method according to claim 1, wherein The minimum cost features of each heterogeneous resource node are screened from the cost matrix of the resource scheduling based on the pre-trained reinforcement learning model, specifically including: For each heterogeneous resource node, screening out a scheduling cost sequence of the heterogeneous resource node from the cost matrix of resource scheduling; Extracting cost markers for each maintenance task in the scheduling cost sequence based on a pre-trained reinforcement learning model; The minimum value of all cost flags is taken as the minimum cost feature of the heterogeneous resource node, and then the minimum cost feature of each heterogeneous resource node is obtained.

4. The method according to claim 1, wherein The scheduling load cost of the target task on each heterogeneous resource node is determined based on the minimum cost characteristics and the load characteristics of the target task, specifically including: For each heterogeneous resource node, obtain the idle computing power of the heterogeneous resource node; Determining the scheduling load of the target task on the heterogeneous resource nodes according to the idle computing power and the load characteristics of the target task; The scheduling load is optimized according to the minimum cost characteristics of the heterogeneous resource nodes to obtain the scheduling load cost of the target task for the heterogeneous resource nodes, and then the scheduling load cost of the target task for each heterogeneous resource node is obtained.

5. The method according to claim 1, wherein Using the training and pushing machine to combine the various scheduling load costs to allocate heterogeneous resources to target tasks specifically includes: For each heterogeneous resource node, the scheduling load cost of the target task on the heterogeneous resource node is used as the allocated load of the training and pushing machine; Determine the load ratio of the target task in the resource heterogeneous nodes based on the distributed load in the training and push machine, and then obtain the load ratio of the target task in each resource heterogeneous node; The execution nodes of the target task are selected from all resource heterogeneous nodes through all load ratios.

6. The method according to claim 1, wherein Using various slack variables to reallocate resources to target tasks specifically includes: A heterogeneous resource node with the largest slack variable is selected from the heterogeneous resource nodes where task allocation conflicts occur as a reallocated resource node, and the reallocated resource node is used to execute the target task.

7. A heterogeneous resource intelligent scheduling system for a large-scale model training and pushing machine, which uses the method described in any one of claims 1 to 6 to perform heterogeneous resource intelligent scheduling, characterized in that: The system includes: A collection module is used to collect a list of tasks to be assigned in the training and pushing integrated machine, and select a task to be assigned from the list of tasks to be assigned as a target task; The processing module is used to initialize the maintenance task list of each heterogeneous resource node in the training and push integration machine, and generate the cost matrix of resource scheduling when the training and push integration machine performs maintenance tasks based on all the maintenance task lists; The processing module is further configured to filter out the minimum cost feature of each heterogeneous resource node from the cost matrix of the resource scheduling based on a pre-trained reinforcement learning model, and then determine the slack variables of each heterogeneous resource node when an allocation conflict occurs in the target task; The processing module is further configured to determine the scheduling load cost of the target task for each heterogeneous resource node based on each minimum cost feature and the load feature of the target task; The execution module is used to use the training and pushing machine in combination with each scheduling load cost to perform heterogeneous resource allocation for the target task. When a task allocation conflict occurs in each heterogeneous resource allocation, each slack variable is used to reallocate resources to the target task until the resource allocation of each task to be allocated in the task list is completed.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that the computer device executes the heterogeneous resource intelligent scheduling method of the large model training and promotion integrated machine described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions or codes, and when the instructions or codes are executed on a computer, the computer implements the heterogeneous resource intelligent scheduling method of the large-model training and promotion integrated machine as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Heterogeneous resource-oriented adaptive task scheduling method based on deep reinforcement learning

    CN118916132A

Cited By

  • An AGI-oriented global cost optimization phased closed-loop decision method, system and storage medium

    CN122655854A