An intelligent allocation and scheduling method and system for heterogeneous resources based on deep learning
By building an intelligent allocation scheduling model based on DeepFM and LSTM, the computing power allocation of hybrid heterogeneous server clusters is optimized, and the problem of low execution efficiency of computing tasks is solved, and efficient utilization of heterogeneous resources and rapid completion of task queues is achieved.
Patent Information
- Application Number
- CN202510355450.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In a hybrid heterogeneous server cluster, how to optimize the computing power allocation strategy, improve the execution efficiency and order of computing tasks, and solve the problems of uneven computing power allocation, complex computing task scheduling, and uneven execution efficiency.
By building an intelligent allocation scheduling model based on DeepFM and LSTM, combining feature data sets, analyzing the resource occupancy of the calculation task queue and the remaining resources of the platform, and dynamically adjusting the resource scheduling strategy is adopted to achieve efficient execution of computing tasks on heterogeneous resource platforms such as CPU, GPU, and NPU.
It realizes the rapid completion of batch computing tasks, greatly improves overall computing efficiency and resource utilization, and ensures the timely and effective handling of key computing tasks.
Smart Images

Figure CN119883648B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent software development for heterogeneous resources, and specifically relates to a method and system for intelligent allocation and scheduling of heterogeneous resources based on deep learning. Background Art
[0002] With the continuous progress and profound transformation of artificial intelligence technology, a series of cutting-edge intelligent computing tasks have emerged, covering multiple key fields such as speech recognition, image recognition, intelligent control systems, intelligent monitoring, autonomous driving technology, and natural language processing. These complex and diverse computing tasks all rely on the powerful support of the hardware platform to perform efficient and accurate data processing and analysis.
[0003] In the current technology ecosystem, the hardware cornerstones for executing deep learning computing tasks mainly include CPUs, GPUs, and emerging NPUs. As a general-purpose processor, the CPU can handle various types of computing tasks, but its efficiency is relatively low when dealing with large-scale parallel data operations; the GPU, with its powerful parallel processing ability, shows significant advantages in graphics rendering and specific types of computing acceleration; while the NPU is designed specifically for neural network operations such as deep learning and can provide more efficient and professional computing capabilities.
[0004] In the actual deployment process, due to cost control factors and the inherent dependence of some specific computing tasks on GPUs, server clusters often exhibit a hybrid heterogeneous characteristic, that is, they are equipped with both GPUs and NPUs. Although this hybrid architecture improves the flexibility and compatibility of the system to a certain extent, when faced with large-scale and diverse computing tasks, it also brings problems such as uneven computing power allocation, complex computing task scheduling, and uneven execution efficiency. Therefore, how to optimize the computing power allocation strategy of the hybrid heterogeneous server cluster and improve the execution efficiency and orderliness of computing tasks has become a key technical problem to be solved urgently. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for intelligent allocation and scheduling of heterogeneous resources based on deep learning, which realizes the efficient execution of multiple concurrent computing tasks on heterogeneous computing resource platforms such as CPUs, GPUs, and NPUs.
[0006] A method for intelligent allocation and scheduling of heterogeneous resources based on deep learning provided by the present invention specifically includes the following steps:
[0007] Step 1: Obtain the historical computing task queue, collect the task key features, importance values, and computing task serial numbers, and construct a feature dataset by calculating the total execution duration of the computing task queue and the remaining resource amount of the platform; the task key features include the priority of executing the calculation, the task execution time, the execution computing platform, and the computing resource occupancy.
[0008] Step 2: Build an intelligent allocation and scheduling model based on DeepFM and LSTM. The input of the intelligent allocation and scheduling model is the task key features, total execution duration, and remaining platform resources of the computing task queue, and the output is the computing task sequence. Use the feature dataset constructed in Step 1 to complete the training of the intelligent allocation and scheduling model.
[0009] In actual use, in Step 3: Obtain the task key features of the computing tasks to be processed, divide the computing tasks to be processed into multiple computing task queues according to the execution computing platform, and calculate the total resource occupancy of the computing task queues that need to be executed on each computing platform. If the total resource occupancy is not greater than the remaining platform resources, then schedule and execute the computing task queue; otherwise, input the task key features, total execution duration, and remaining platform resources into the intelligent allocation and scheduling model obtained in Step 2 to obtain the computing task sequence, and then complete the execution of the computing task queue on the computing platform according to the computing task sequence.
[0010] Further, the intelligent allocation and scheduling model is formed by cascading a DeepFM layer, an LSTM layer, and an output layer. Among them, the input of the DeepFM layer is the task key features of the computing tasks in the computing task queue, and the output is the importance value of each computing task; the input of the LSTM layer is the importance value and key features of the computing tasks in the computing task queue, the total execution duration of the computing task queue, and the remaining platform resources, and the output is the timing features of the computing tasks; the input of the output layer is the timing features of the computing tasks, and the output is the computing task sequence.
[0011] Further, use the GAUC evaluation method to evaluate the training results of the intelligent allocation and scheduling model.
[0012] Further, while following the computing task sequence, adopt a dynamic resource scheduling strategy to complete the execution of the computing task queue on the computing platform.
[0013] Further, the dynamic resource scheduling strategy is to use the immediate preemption priority scheduling algorithm and the fast switching scheduling algorithm to dynamically schedule the computing tasks.
[0014] Further, the method of scheduling and executing the computing task queue in Step 3 is: Schedule and execute according to the priority of the computing tasks.
[0015] An intelligent heterogeneous resource allocation and scheduling system based on deep learning provided by the present invention includes a model training module, an input module, a parsing module, a heterogeneous computing module, a log module, a data management module, and a resource monitoring module;
[0016] The model training module is responsible for training the intelligent allocation and scheduling model, including a training sub-module and an evaluation sub-module; the input module is responsible for receiving the basic information of the computing tasks and organizing it into a computing task queue configuration file, and sending the computing task queue configuration file to the parsing module; the parsing module is used to interpret the computing task queue configuration file to obtain the task key features, and send the task key features to the heterogeneous computing module; the heterogeneous computing module is used to analyze and obtain the computing task sequence based on the task key features using the intelligent allocation and scheduling model; the log module is used to record the log information and abnormal events during the execution of the computing tasks in real time; the data management module is used to record the task key features of the computing tasks, and provide functions for data modification and query; the resource monitoring module is used by the user to obtain the computing resource occupancy of the computing tasks, the remaining resources of the platform, and the execution progress of the computing tasks.
[0017] Furthermore, it is deployed and run in a Docker containerized manner.
[0018] Furthermore, the input module uses the command-line tool Curl to send a POST input request to the parsing module to complete the transfer of the computing task queue configuration file.
[0019] Furthermore, the parsing module performs standardization processing on the obtained task key features. Beneficial effects
[0020] The present invention establishes a feature dataset by analyzing the optimal allocation scheme of historical computing task resources, constructs an intelligent allocation and scheduling model based on DeepFM and LSTM and uses the feature dataset to complete the training. During actual use, the task key features of the computing tasks to be processed are obtained, and the computing tasks to be processed are divided into computing task queues. The total resource occupancy of the computing task queues required by each computing platform is obtained, and the scheduling execution method of the computing task queues is determined according to the relationship between the total resource occupancy and the remaining resources of the platform, making full use of heterogeneous computing resources such as CPUs, GPUs, and NPUs, thereby realizing the rapid completion of batch computing tasks and greatly improving the overall computing efficiency and resource utilization rate. Description of the drawings
[0021] Figure 1 It is a schematic flowchart of a method for intelligent allocation and scheduling of heterogeneous resources based on deep learning provided by the present invention.
[0022] Figure 2 It is a schematic structural diagram of an intelligent allocation and scheduling model constructed in a method for intelligent allocation and scheduling of heterogeneous resources based on deep learning provided by the present invention.
[0023] Figure 3Flowchart of the immediate preemption priority scheduling algorithm adopted in a heterogeneous resource intelligent allocation and scheduling method based on deep learning provided by the present invention.
[0024] Figure 4 Schematic structural diagram of a heterogeneous resource intelligent allocation and scheduling system based on deep learning provided by the present invention. Detailed implementation manners
[0025] The following are specific embodiments listed with reference to the accompanying drawings to describe the present invention in detail.
[0026] A heterogeneous resource intelligent allocation and scheduling method and system based on deep learning provided by the present invention, the core idea of which is: to establish a feature data set by analyzing the optimal allocation scheme of historical computing task resources, construct an intelligent allocation and scheduling model based on DeepFM and LSTM and complete the training using the feature data set. During actual use, obtain the task key features of the computing task to be processed, divide the computing task to be processed into a computing task queue, obtain the total resource occupancy of the computing task queue that each computing platform needs to execute, and determine the scheduling execution method of the computing task queue according to the relationship between the total resource occupancy and the remaining resources of the platform.
[0027] A heterogeneous resource intelligent allocation and scheduling method based on deep learning provided by the present invention, the specific process is as Figure 1 shown, specifically including the following steps:
[0028] Step 1: Define the task key features of the computing tasks in the optimal allocation scheme of historical computing task resources. During the process of users using heterogeneous resource CPU, GPU, NPU hardware platforms to execute computing tasks such as training, deployment, or debugging, collect the task key features, importance values, and computing task serial numbers of each computing task, as well as the total execution duration of all computing tasks in the computing task queue and the remaining resources of the computing platform before the computing task queue is executed, form the computing task queue features, and establish a feature data set.
[0029] The computing task queue features in the present invention include the task key features of the computing tasks, the total execution duration of all computing tasks in the computing task queue, and the remaining resources of the platform before the computing task queue is executed. Among them, the task key features of the computing tasks include the computing task ID, computing task type, execution priority of the computing, computing task execution time, execution computing platform, and computing resource occupancy.
[0030] The computing task ID is automatically assigned by the system and is the unique identifier of the computing task. For example, it increments automatically starting from 1.
[0031] The computing task types include training, deployment, and debugging, etc.
[0032] The priority of executing a calculation represents the urgency of the calculation task, indicating the priority level when multiple calculation tasks are executed. The calculation task with a higher priority obtains computing resources first, and the calculation task with a lower priority needs to wait until there are idle computing resources before execution, including: P1 (urgent), P2 (high), P3 (medium), and P4 (low), etc.
[0033] The execution time of a calculation task is the time from the start of execution to the end of the calculation task, measured in milliseconds.
[0034] The execution computing platform represents the computing hardware platform used for the calculation task, including: CPU, NPU, GPU.
[0035] The computing resource occupancy represents the size of the maximum computing resources required on different heterogeneous hardware platforms when executing a calculation task, measured in MB.
[0036] The importance value represents the degree of importance of the execution of a calculation task, which is a weight value, and the value range is from 0 to 1. The closer it is to 1, the more important the calculation task is.
[0037] The calculation task serial number represents the sorting serial number of the calculation task in the optimal allocation scheme of historical calculation task resources.
[0038] The total execution duration of all calculation tasks in the calculation task queue, measured in seconds.
[0039] The remaining platform resources before the execution of the calculation task queue is the remaining amount of heterogeneous resources before the execution of the calculation task queue, measured in MB.
[0040] Taking a calculation task queue containing two calculation tasks as an example, the characteristics of the calculation task queue are shown in Table 1:
[0041] Table 1 Characteristics Table of Calculation Task Queue
[0042]
[0043] Step 2: Perform data preprocessing on the characteristics of the calculation task queue in the feature dataset to form a feature dataset.
[0044] In the data preprocessing process of the calculation task queue characteristics in the present invention, the following steps are included:
[0045] Step 2.1: Perform data identification processing on the calculation task type, the priority of executing the calculation, and the execution computing platform, and assign a unique ID value.
[0046] The calculation task type is assigned an ID value. The value 1 is used for training, the value 2 is used for deployment, and the value 3 is used for debugging.
[0047] The priority of executing calculations is assigned an ID value. P1 (urgent) is represented by the numerical value 1; P2 (high) is represented by the numerical value 2; P3 (medium) is represented by the numerical value 3; P4 (low) is represented by the numerical value 4.
[0048] The ID value is assigned to the execution calculation platform. The CPU is represented by the numerical value 1, the NPU is represented by the numerical value 2, and the GPU is represented by the numerical value 3.
[0049] Taking the calculation task queue containing two calculation tasks as an example, the characteristics of the calculation task queue after data identification processing are shown in Table 2:
[0050] Table 2 Characteristics Table of the Calculation Task Queue with Data Identification
[0051]
[0052] Step 2.2: Use the StandardScaler() function of the machine learning tool Sklearn (Scikit-learn) in the Python language to standardize the characteristics of the calculation task queue after data identification processing.
[0053] Step 2.3: Use the data analysis tool Pandas in the Python language to write the feature set after the standardization process in Step 2.2 into the Comma-Separated Values file format (CSV) to form a feature data set.
[0054] Step 3: Split the feature data set formed in Step 2 into two CSV files, a training set and an evaluation set.
[0055] First, use the open file function in Python, that is, the open() function, to read the CSV file of the feature data set; then, use the train_test_split() function of the machine learning tool Sklearn (Scikit-learn) based on the Python language to split the read data. For example, randomly sample 70% of the data to construct the training set, 15% of the data to construct the training set, and the remaining data as the evaluation set to form the training set of the calculation task queue characteristics.
[0056] Step 4: Build an intelligent allocation and scheduling model based on the hybrid model of DeepFM and LSTM, and use the training set of the calculation task queue characteristics formed in Step 3 to complete the training of the intelligent allocation and scheduling model.
[0057] To achieve intelligent allocation and scheduling of heterogeneous resources, that is, to intelligently allocate heterogeneous resources and schedule them to maximize resource utilization, the present invention uses DeepFM and LSTM in machine learning to establish a hybrid model. Among them, DeepFM is good at capturing high-order interactions and low-order linear relationships between features and is suitable for processing structured data, such as computing task features, resource features, etc.; LSTM is good at capturing temporal dependence relationships and is suitable for processing dynamically changing computing tasks and resource states. The present invention establishes an intelligent allocation and scheduling model based on the DeepFM and LSTM hybrid model. This hybrid model can simultaneously model static feature interactions and dynamic temporal dependencies, thereby more comprehensively optimizing resource allocation.
[0058] The intelligent allocation and scheduling model constructed by the present invention has a model structure as Figure 2 shown, which is formed by cascading a DeepFM layer, an LSTM layer, and an output layer. Among them, the input of DeepFM is the computing task ID, computing task type, execution priority of the computation, computing task execution time, execution computing platform, and computing resource occupancy in the computing task queue feature, and the output is the importance value of each computing task in the computing task queue; the input of LSTM is the importance value of each computing task output by DeepFM and the computing task ID, computing task type, execution priority of the computation, computing task execution time, execution computing platform, computing resource occupancy, total execution duration of all computing tasks in the computing task queue, and remaining platform resources before the computing task queue execution in the computing task queue feature, and the output is the temporal feature of the computing task; the input of the output layer is the temporal feature of the computing task, and the output is a sorted computing task sequence, and this computing task sequence is the optimal resource allocation combination scheme.
[0059] The optimal resource allocation combination scheme refers to a queue formed by load balancing and optimal resource allocation combination of all computing tasks on different heterogeneous hardware platforms. It is mainly a new computing task sequence formed after recombination and sorting according to its execution platform, execution time, computing power requirements, and importance level. The computing tasks are scheduled according to this computing task execution sequence to achieve load balancing and reasonable scheduling of heterogeneous resources. In the new computing task queue, the resource allocation of the computing tasks follows the order from front to back.
[0060] In practical applications, computing power requirements, computing task execution time, and importance are often interrelated. A computing task may simultaneously have high computing power requirements, long execution time, and high importance. In such a case, the system needs to comprehensively consider these three factors, and the present invention uses an intelligent allocation and scheduling model for trade-offs. By sorting according to the three aspects of time, computing power, and importance, the order and priority of computing task execution can be adjusted more flexibly, thereby ensuring the balance between resource requirements and computing task execution. Specifically, the present invention preferentially arranges the execution of computing tasks with short execution time, moderate computing power requirements, and high importance to ensure that critical computing tasks can be processed in a timely and effective manner. Taking GPU resource allocation as an example, assume that the remaining resource amount of the current GPU platform is 24GB, and in the computing task queue that needs to be executed on the GPU, the resource requirements of each computing task are different, and the total demand is 36GB. In this case, the computing task queue will first undergo analysis and calculation by the intelligent allocation and scheduling model to obtain an optimized computing task sequence. Then, resource allocation is performed according to the computing task sequence, and the computing tasks are executed until the remaining resource amount of the GPU platform is completely occupied.
[0061] Step 5: Use the evaluation set formed in Step 3 to evaluate the intelligent allocation and scheduling model trained in Step 4. The process of using the evaluation set to evaluate the trained intelligent allocation and scheduling model is the process of evaluating whether the model fits the evaluation set.
[0062] The present invention uses GAUC to calculate the score of the intelligent allocation and scheduling model, evaluates the training effect of the intelligent allocation and scheduling model according to the score, and outputs the evaluation result. The specific evaluation process is as follows:
[0063] Step 5.1: Input the data in the evaluation set into the trained intelligent allocation and scheduling model to obtain the prediction result.
[0064] Step 5.2: Calculate the True Positive Rate (TPR) and False Positive Rate (FPR) according to the prediction result and the actual result in the evaluation set. The calculation process is as follows:
[0065]
[0066] Among them, TP (True Positive) means that the actual is a positive example and the prediction is also a positive example. FP (False Positive) means that the actual is a negative example, but the prediction is a positive example. FN (False Negative) means that the actual is a positive example, but the prediction is a negative example. TN (True Negative) means that the actual is a negative example and the prediction is also a negative example.
[0067] Step 5.3: Draw an ROC curve based on the TPR value and the FPR value, with the TPR as the vertical axis and the FPR as the horizontal axis.
[0068] Step 5.4: Calculate the AUC value according to the ROC curve, that is, the area under the ROC curve. The calculation process is as follows:
[0069]
[0070] Step 5.5: Further calculate the GAUC according to the AUC value. GAUC actually calculates the AUC of each calculation task queue, then performs a weighted average, and finally obtains the group AUC, that is, GAUC. The calculation process is as follows:
[0071]
[0072] Step 5.6: Determine whether the intelligent allocation and scheduling model meets the usage requirements by setting a threshold for GAUC. For example, set the threshold to 0.7, that is, when GAUC is greater than or equal to 0.7, it is determined that the current intelligent allocation and scheduling model can be used for heterogeneous resource intelligent allocation and scheduling; otherwise, adjust the parameters in the training to retrain the intelligent allocation and scheduling model until GAUC is greater than 0.7 to obtain the optimal intelligent allocation and scheduling model. Among them, adjusting the parameters in the training includes adjusting the parameters in DeepFM and LSTM and adjusting the training process parameters.
[0073] Among them, DeepFM includes the following parameters: k, the hidden vector dimension of the factorization machine, that is, the dimension of the hidden vector corresponding to each feature. This parameter determines the complexity of the feature interaction learned by the intelligent allocation and scheduling model; embedding_size, the embedding dimension of the feature in the embedding layer. Before each feature is input into the DNN, it is first converted into a high-dimensional vector through an embedding layer, and this dimension is the embedding_size; hidden_units, the number of hidden layer units in the DNN is a list that specifies the number of neurons in each hidden layer. For example, [128, 64] means there are two hidden layers, the first layer has 128 neurons, and the second layer has 64 neurons. The LSTM algorithm includes the following parameters: hidden_size, the dimension of the hidden state, that is, the output dimension of the LSTM cell; num_layers, the number of layers of the LSTM; time_steps, the time step.
[0074] Adjust the parameters of the training process, including the following parameters: learning_rate, the learning rate, which controls the step size of weight updates; epochs, the number of training cycles, that is, the number of times the entire training set is used for training; batch_size, the size of each batch, that is, the number of samples used for each gradient update; optimizer, the optimization algorithm, such as Adam, SGD, etc., which is used to update the weights of the intelligent allocation and scheduling model.
[0075] Step 6. In actual use, obtain the basic information of multiple deep learning computing tasks to form a computing task queue. The basic information of the computing task includes the computing task name, computing task path, Python interpreter path, parameter details, input and output formats of the computing task, execution priority, execution method, computing task execution time, execution computing platform, computing resource occupancy, etc., and check the correctness of the input according to the requirements.
[0076] Among them, the computing task name is a custom name for the computing task. The computing task path represents the absolute path address of the executable program or py program of the computing task. The Python interpreter path represents the interpreter path address specified by the user when the computing task path is a py program to execute the py program. The parameter details represent the parameter details that need to be passed in when the executable program or py program of the computing task is executed. The input and output formats of the computing task represent the data formats that the computing task needs to input, including the format of passing parameters, and the data format of the output after the computing is completed, so that the data is saved to the specified location according to the specified data format after the computing is completed. The execution priority of the computing represents the priority level when multiple computing tasks are executed. The execution method represents the type of execution method of the computing task, including: one-time computing task, loop computing task, and timed computing task. The computing task execution time is the time from the start to the end of the execution of the computing task, in milliseconds. The execution computing platform represents the computing platform device for the execution of the computing task. The system judges the computing hardware supported by the computing task and automatically allocates the execution computing platform for the computing task, so that the computing task is pushed to the allocated hardware platform during resource scheduling and the computing is executed. The computing resource occupancy represents the size of the maximum computing resources required on different heterogeneous hardware platforms when the computing task is executed.
[0077] The computing task queue configuration file contains the basic information of all computing tasks and is stored in the Json file format. For example, the Json format of the configuration file of a certain computing task is as follows:
[0078] {
[0079] "tasks":
[0080] {
[0081] "task_name":"readconfig",
[0082] "task_path":" / home / xx / project / readconfig.bin",
[0083] "python":null,
[0084] "args":"-c / home / xx / project / config.json",
[0085] "remote_output":null,
[0086] "priority":3,
[0087] "trigger_type":1
[0088] },
[0090] }
[0091] Step 7: Obtain the usage data of all computing resources (heterogeneous computing resources such as CPUs, GPUs, NPUs, etc.) and the remaining computing resources.
[0092] For example, the user can use the input interface to input the computing task queue configuration file. Among them, the input interface is implemented by Curl. A POST input request is sent through Curl, and the request parameter structure is in Json format. The data content is the absolute path address of the computing task queue configuration file. For example: curl -H "Content-Type: application / json" -X POST -d '{"config_path":" / root / config.json"}' http: / / 192.168.1.198:19891 / postTaskQueue. If a success message is returned, it means that the computing task queue configuration file has been received. Otherwise, try to resend the request.
[0093] Step 8: Parse the computing task queue configuration file, obtain the basic information of each computing task in the computing task queue configuration file, parse the computing task to obtain the task key features of the computing task, and record the basic information of the computing task in the database for persistent storage.
[0094] Step 9: Initially divide all computing tasks according to heterogeneous resources according to the execution computing platform in the basic information of the computing tasks. For example, NPU computing tasks are divided into the same computing task queue, and GPU computing tasks are divided into another same computing task queue.
[0095] Step 10: Calculate the total resource occupancy of the computing tasks to be scheduled and executed on each computing platform. If the total resource occupancy of the computing tasks does not exceed the remaining platform resources of the computing platform, then schedule and execute these computing tasks according to the priorities of the computing tasks; otherwise, analyze and calculate the task key features of the computing task queue through the intelligent allocation scheduling model. The specific analysis and calculation are as follows:
[0096] Step 10.1: Standardize the task key features of the computing tasks in the computing task queue;
[0097] Step 10.2: Initialize and load the intelligent allocation scheduling model obtained from the training in Step 4;
[0098] Step 10.3: Then input the task key features of the computing tasks processed in Step 10.1 and the remaining platform resources of the computing platform into the intelligent allocation scheduling model for analysis and calculation, and output the optimal resource allocation combination plan.
[0099] Step 11: Adopt a dynamic resource scheduling strategy and schedule the computing tasks onto the heterogeneous computing hardware platform according to the optimal resource allocation combination plan. In the dynamic resource scheduling strategy, the immediate preemption priority scheduling algorithm and the fast switching scheduling algorithm are used to dynamically schedule the computing tasks.
[0100] For the immediate preemption priority scheduling algorithm, the processing flow is as Figure 3 shown, and the scheduling process includes:
[0101] Specify a time slice in advance, use a priority queue to store all computing task processes, and sort them in descending order according to the priorities of the computing tasks. Whenever a computing task process finishes execution or is preempted, the next computing task process with the highest priority is taken out from the queue for execution. If there are multiple computing task processes with the same priority, the round-robin method is used to execute these processes in turn.
[0102] Among them, the round-robin algorithm divides the time on the computing resources into several time slices. Each process executes for a certain period of time within a time slice, and then is suspended and put back into the ready queue to wait for the next scheduling. If a process does not finish execution within a time slice, it will be suspended and put back to the end of the ready queue to wait for the next scheduling. When implementing the round-robin algorithm:
[0103] The length of the time slice should be set according to the actual situation, usually between 10ms and 100ms;
[0104] The ready queue should be implemented using data structures such as queues or linked lists to facilitate the quick insertion and deletion of processes;
[0105] Each process needs to record its own execution time and remaining time to determine whether it needs to continue execution at the end of the time slice;
[0106] At the end of each time slice, the current process needs to be put back to the end of the ready queue, and the next process is selected from the queue for execution.
[0107] The scheduling process of the fast-switching scheduling algorithm is as follows: Under the condition of real-time monitoring of computing resources, when the previous computing task in execution is in the completion critical section or an external interruption occurs, the system prepares the next computing task process in advance. Once the completion or interruption occurs, the system immediately switches the computing tasks, reducing the waiting time of real-time computing tasks.
[0108] Step 12: Record the running log of the computing task during the execution of the computing task to form a computing task log file.
[0109] Step 13: After the execution of the computing task is completed, organize the output result of the computing task into the target format specified in the computing task queue configuration file and save it in the specified directory location.
[0110] Step 14: After all computing tasks in the computing task queue are executed, collect and record all information of the computing task queue into the feature dataset. After the feature dataset is updated, repeat Steps 2, 3, 4, and 5 to iteratively train the intelligent allocation and scheduling model.
[0111] An intelligent allocation and scheduling system for heterogeneous resources based on deep learning provided by the present invention has a structure as Figure 4 shown, including a model training module, an input module, a parsing module, a heterogeneous computing module, a log module, a data management module, and a resource monitoring module. The system is deployed and run in a Docker containerized manner, which can improve the system deployment efficiency to adapt to various different server cluster environment scenarios, and can also provide a completely isolated running environment for the operation of each module. In the complex environment of the server cluster, it can reduce the interference of other libraries or packages and improve the reliability of the system.
[0112] Among them, the model training module is responsible for training the intelligent allocation and scheduling model, and this module includes a training sub-module and an evaluation sub-module. The training sub-module is used to build an intelligent allocation and scheduling model based on machine learning, and extract feature data from the feature dataset for preprocessing to form a dataset to train the intelligent allocation and scheduling model. The evaluation sub-module is used to evaluate the intelligent allocation and scheduling model using the GAUC metric. If the evaluation passes, the optimal model is obtained; otherwise, the parameters in the training are adjusted and retrained until the requirements are met.
[0113] An input module, responsible for receiving the basic information of the computing tasks provided by the user. These information are organized into a computing task queue configuration file, and a POST input request is sent to the system via Curl to input the computing task queue configuration file into the system, and the computing task queue configuration file is passed to the parsing module.
[0114] A parsing module, used to interpret the computing task queue configuration file and the content of each computing task, and extract the key task features from it, including computing task ID, computing task type, computing priority, execution time, execution computing platform, computing resource occupancy, etc.; then perform standardization processing on these key task features, and the processed data will be input into the heterogeneous computing module and synchronized to the data management module to achieve data management.
[0115] A heterogeneous computing module, according to the standardized data passed by the parsing module, uses an intelligent allocation and scheduling model to analyze and obtain the optimal resource allocation combination plan. The module is equipped with an analysis sub-module and a resource scheduling sub-module. The analysis sub-module is mainly responsible for calculating the resource requirements of the computing task queue and the remaining platform resources of the heterogeneous system, and calculating the best resource combination plan through an intelligent model. The resource scheduling sub-module, based on this plan, adopts a dynamic adjustment resource scheduling strategy to reasonably allocate computing tasks to the heterogeneous hardware platform for execution, so as to achieve the optimal configuration of resources and the rapid completion of the computing task queue. At the same time, the resource scheduling sub-module is also connected to the resource monitoring module to monitor the execution status of computing tasks.
[0116] A log module, used to record the log information during the execution of computing tasks in real time, and capture and record abnormal events in time for subsequent maintenance and troubleshooting of computing tasks.
[0117] A data management module, used to record the detailed information of computing tasks and their execution processes, such as execution status, execution time, etc., and provide functions for adding, deleting, modifying and querying data to achieve comprehensive management of computing task data, and record the key task features of computing tasks and manage the feature data set.
[0118] A resource monitoring module, during the operation of the system, is responsible for continuously monitoring the computing resource occupancy, the remaining platform resources and the execution progress of computing tasks of all online computing tasks to ensure the effective utilization of system resources and the smooth execution of computing tasks.
[0119] Specifically, during the period when the user operates the heterogeneous resource hardware platform (including CPU, GPU, NPU, etc.) to execute various computing tasks such as training, deployment, and debugging, the system will implement a series of steps to optimize computing task management and resource allocation. The system will first monitor and record the optimal allocation scheme of historical computing task resources, which includes the characteristics of the computing task queue, including the task key characteristics of each computing task, and at the same time summarize the overall execution duration of the computing task queue. These detailed characteristics are then stored in the feature dataset. The system retrieves these feature data from the feature dataset, and after standardization processing, exports the data in the CSV file format to form the dataset required for the intelligent allocation and scheduling model. The dataset is further divided into a test set and an evaluation set for the training and evaluation verification of the intelligent allocation and scheduling model. The construction of the intelligent allocation and scheduling model is based on machine learning technology, and the training process depends on the training set data. The user submits the basic information of the deep learning computing task through the login interface, and the basic information is used to generate the computing task queue configuration file. After the user imports the computing task queue configuration file into the system, the system will retrieve and summarize the usage status and remaining capacity of all available heterogeneous computing power resources (CPU, GPU, NPU, etc.) in real time. The system parses the computing task queue configuration file, extracts the specific requirements of each computing task, and preliminarily classifies the computing tasks according to the characteristics of the computing tasks (especially the specified hardware platform type). On this basis, the system calculates the total amount of resources required for the computing tasks to be scheduled on each computing platform. If the total amount does not exceed the remaining resources of the platform, the scheduling is directly executed; otherwise, the intelligent allocation and scheduling model is enabled to conduct in-depth analysis based on the task key characteristics of the computing task queue to determine the optimal resource allocation scheme, aiming to complete all computing tasks in the shortest time. According to the obtained optimal resource allocation scheme, the system adopts a dynamic adjustment resource scheduling strategy to accurately schedule the computing tasks to the appropriate heterogeneous hardware platform to achieve efficient computing power allocation and scheduling. During the execution stage of the computing task, the system will continuously record the operation logs of the computing task to ensure traceability. After the computing task is completed, the system will organize and save the output results according to the format and storage path specified in the configuration file. In this process, the task key characteristics of each computing task are collected and recorded again, the total execution duration of all computing tasks in the computing task queue and the remaining resources of the platform during execution are calculated, and these characteristics are written into the feature dataset for updating the intelligent allocation and scheduling model.
[0120] In summary, the above is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent allocation and scheduling method for heterogeneous resources based on deep learning, characterized in that, Specifically, it includes the following steps: Step 1: Obtain the historical calculation task queue, collect the task key features, importance values, and calculation task serial numbers, and calculate the total execution duration of the task queue and the remaining platform resources to construct a feature dataset; the task key features include the priority of executing the calculation, task execution time, execution calculation platform, and calculation resource occupancy; Step 2: Build an intelligent allocation and scheduling model based on DeepFM and LSTM. The input of the intelligent allocation and scheduling model is the task key features, total execution duration of the calculation task queue, and the remaining platform resources, and the output is the calculation task sequence; use the feature dataset constructed in Step 1 to complete the training of the intelligent allocation and scheduling model; Step 3: In actual use, obtain the task key features of the calculation task to be processed, divide the calculation task to be processed into multiple calculation task queues according to the execution calculation platform, and calculate the total resource occupancy of the calculation task queues that need to be executed on each calculation platform; if the total resource occupancy is not greater than the remaining platform resources, then schedule and execute the calculation task queue, otherwise input the task key features, total execution duration, and the remaining platform resources into the intelligent allocation and scheduling model obtained in Step 2 to get the calculation task sequence, and then complete the execution of the calculation task queue on the calculation platform according to the calculation task sequence; The intelligent allocation and scheduling model is formed by cascading the DeepFM layer, LSTM layer, and output layer. Among them, the input of the DeepFM layer is the task key features of the calculation tasks in the calculation task queue, and the output is the importance value of each calculation task; the input of the LSTM layer is the importance value and key features of the calculation tasks in the calculation task queue, the total execution duration of the calculation task queue, and the remaining platform resources, and the output is the timing features of the calculation tasks; the input of the output layer is the timing features of the calculation tasks, and the output is the calculation task sequence.
2. The intelligent allocation and scheduling method for heterogeneous resources according to claim 1, wherein Use the GAUC evaluation method to evaluate the training results of the intelligent allocation and scheduling model.
3. The intelligent allocation and scheduling method for heterogeneous resources according to claim 1, wherein While following the calculation task sequence, adopt a dynamic adjustment resource scheduling strategy to complete the execution of the calculation task queue on the calculation platform.
4. The heterogeneous resource intelligent allocation and scheduling method according to claim 3, characterized in that The dynamic adjustment resource scheduling strategy is to adopt the priority scheduling algorithm of immediate preemption and the fast switching scheduling algorithm to dynamically schedule the calculation tasks.
5. The intelligent allocation and scheduling method for heterogeneous resources according to claim 1, characterized in that The method of scheduling and executing the calculation task queue in Step 3 is: schedule and execute according to the priority of the calculation task.
6. A heterogeneous resource intelligent allocation and scheduling system based on deep learning that adopts the heterogeneous resource intelligent allocation and scheduling method described in claim 1, characterized in that, It includes a model training module, an input module, a parsing module, a heterogeneous computing module, a log module, a data management module, and a resource monitoring module; The model training module is responsible for training the intelligent allocation and scheduling model, including a training sub-module and an evaluation sub-module; the input module is responsible for receiving the basic information of the calculation task and organizing it into a calculation task queue configuration file, and sending the calculation task queue configuration file to the parsing module; The parsing module is used to interpret the calculation task queue configuration file to obtain the task key features, and send the task key features to the heterogeneous computing module; the heterogeneous computing module is used to analyze and obtain the calculation task sequence according to the task key features by using the intelligent allocation and scheduling model; the log module is used to record the log information and abnormal events during the execution of the calculation task in real time; The data management module is used to record the task key features of the computing tasks and provide functions for data modification and query; the resource monitoring module is used by the user to obtain the computing resource occupancy of the computing tasks, the remaining platform resources, and the execution progress of the computing tasks.
7. The heterogeneous resource intelligent allocation and scheduling system according to claim 6, wherein It is deployed and run in the Docker containerized manner.
8. The heterogeneous resource intelligent allocation and scheduling system according to claim 6, wherein The input module uses the command-line tool Curl to send a POST input request to the parsing module to complete the transfer of the computing task queue configuration file.
9. The heterogeneous resource intelligent allocation and scheduling system according to claim 6, wherein The parsing module performs standardization processing on the obtained task key features.
Citation Information
Patent Citations
Computing power flexible combination method and system based on embedded platform
CN116737397A
Cross-domain cloud platform neural network training task scheduling method based on reinforcement learning
CN118550667A