Computing resource scheduling system for distributed AI training tasks

Through the combination of task reception, resource evaluation, deep Q network allocation and LSTM prediction, the problem of inefficient computing resource scheduling in a distributed computing environment is solved, and efficient resource utilization and stable task execution are achieved.

CN120353585AInactive Publication Date: 2025-07-22SHANGHAI YIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510429677.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a distributed computing environment, the existing technology cannot efficiently schedule computing resources, resulting in low resource utilization and low task execution efficiency, affecting the progress and effectiveness of AI training.

Method used

The task receiving module is used for preprocessing, the resource evaluation module evaluates the computing node status in real time, uses the deep Q network to perform task allocation, and optimizes resource allocation through task execution monitoring and resource dynamic adjustment modules, and dynamic adjustment is made in combination with LSTM to predict future resource requirements.

Benefits of technology

It improves the rationality and accuracy of task allocation, improves resource utilization and task execution efficiency, and ensures the stable and continuous execution of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353585A_ABST
    Figure CN120353585A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computing resource management, and particularly discloses a distributed AI training task computing resource scheduling system, which is characterized in that a task receiving module is used for receiving a plurality of AI training tasks; the resource evaluation module is used for evaluating the resource state of each computing node in the distributed computing environment; the task distribution module distributes AI training tasks to appropriate computing nodes by using a deep Q network; the task execution monitoring module is used for monitoring the execution condition of the allocated tasks; and the resource dynamic adjustment module dynamically adjusts the resource allocation of the computing node according to the information fed back by the task execution monitoring module. According to the method, the accurate basis can be provided for task allocation through comprehensive preprocessing of the task and real-time accurate evaluation of the resource state of the computing node. The deep Q network is used for task allocation, so that the system can continuously learn and optimize a task allocation strategy, and the resource utilization rate and the task execution efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computing resource management, and specifically refers to a computing resource scheduling system for distributed AI training tasks. Background Art

[0002] With the rapid development of artificial intelligence technology, the scale and complexity of AI training tasks are constantly increasing. In a distributed computing environment, how to efficiently schedule computing resources to ensure that numerous AI training tasks can be executed quickly and stably has become an urgent problem to be solved. Traditional computing resource scheduling methods often cannot fully consider the diversity of tasks and the dynamic changes in the resource status of computing nodes, resulting in low resource utilization and low task execution efficiency, seriously affecting the progress and effect of AI training.

[0003] Therefore, a computing resource scheduling system for distributed AI training tasks has become an urgent problem for people to solve. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a computing resource scheduling system for distributed AI training tasks, which can achieve efficient and intelligent resource scheduling according to task characteristics and the resource status of computing nodes, and improve resource utilization and task execution efficiency.

[0005] To solve the above technical problem, the technical solution provided by the present invention is: A computing resource scheduling system for distributed AI training tasks, comprising:

[0006] A task receiving module, configured to receive multiple AI training tasks and perform preprocessing, where the preprocessing includes task parsing, priority determination, and task type identification;

[0007] A resource evaluation module, configured to evaluate the resource status of each computing node in the distributed computing environment, and use a combination of periodic evaluation and event-triggered evaluation to update the resource status of each computing node in real time;

[0008] A task allocation module, according to the requirements of the AI training tasks received by the task receiving module and the resource status of each computing node evaluated by the resource evaluation module, uses a deep Q-network to allocate the AI training tasks to appropriate computing nodes;

[0009] A task execution monitoring module, configured to monitor the execution of the allocated tasks, including the execution progress of the tasks, the resource usage, and whether there are any abnormalities;

[0010] A resource dynamic adjustment module, according to the information fed back by the task execution monitoring module, dynamically adjusts the resource allocation of the computing nodes.

[0011] Further, when parsing the task, the task receiving module deeply analyzes the scale of the dataset required by the task, the data format, the algorithm complexity, and the dependent libraries; determines the task priority by considering the business urgency of the task, the impact weight on the overall business, and the estimated execution duration; and identifies the task type according to the type of machine learning algorithm and the training objective adopted by the task.

[0012] Further, for tasks with complex dependencies, the task receiving module presents the sequence and data flow between tasks by constructing a task dependency graph.

[0013] Further, the periodic evaluation of the resource evaluation module comprehensively scans the resources of all computing nodes at fixed time intervals; the event-triggered evaluation immediately updates the resource status information of the corresponding nodes when there are mutations in the resource status of the computing nodes, new nodes are added, or old nodes fail and exit.

[0014] Further, the resource status of the computing node includes the free amounts and usage rates of the CPU, memory, and GPU.

[0015] Further, the method of using a deep Q-network to allocate AI training tasks to appropriate computing nodes is as follows:

[0016] Define the state space: Combine the free amounts and usage rates of the CPU, memory, and GPU of each computing node, as well as the network bandwidth, disk I / O read and write speeds, and the type, priority, and resource requirements of the currently to-be-allocated task into a high-dimensional vector as the state;

[0017] Define the action space: The action represents allocating the current task to a certain computing node, and the action space is the set of all available computing nodes;

[0018] Define the reward function: The reward function is used to measure the quality of each action, and the reward value is comprehensively calculated based on the expected execution duration and resource utilization rate of the task execution; if the overall resource utilization rate can be improved and the expected execution duration can be shortened after the task is allocated, a positive reward is given; otherwise, a negative reward is given;

[0019] Train the DQN network: By continuously interacting with the environment, the agent selects actions according to the current state, obtains a new state and reward after executing the actions, and uses this data to update the parameters of the DQN network to achieve an optimal task allocation strategy.

[0020] Further, the task execution monitoring module also includes an exception handling mechanism. When a task execution exception is detected, the task is automatically restarted; if the restart fails multiple times, the task is reallocated to other appropriate computing nodes according to the real-time resource status provided by the resource evaluation module; at the same time, the execution progress and resource usage of the task are displayed to the system administrator in the form of a chart.

[0021] Furthermore, the resource dynamic adjustment module predicts future resource requirements using LSTM based on historical task execution data and the current resource status, and performs resource pre-allocation in advance. The specific steps for predicting future resource requirements using LSTM are as follows:

[0022] Data preprocessing: Collect the resource usage data of historical tasks, perform normalization processing, and scale the data to the required range;

[0023] Construct an LSTM model: Build a neural network model including an input layer, an LSTM hidden layer, and an output layer. Among them, the input layer receives historical resource usage data, the LSTM hidden layer captures the time series features in the data, and the output layer predicts future resource requirements;

[0024] Model training: Use historical data to train the LSTM model, and adjust the weight parameters of the model through the backpropagation algorithm to minimize the error between the prediction result of the model and the actual value;

[0025] Resource pre-allocation: According to the prediction result of the LSTM model, reserve a certain amount of computing resources in advance for future possible resource requirements. At the same time, during the task execution process, dynamically adjust the resource allocation according to the real-time resource usage situation and task progress to ensure the reasonable allocation of resources among tasks.

[0026] Furthermore, the resource dynamic adjustment module also includes a resource reservation mechanism, which reserves a certain proportion of computing resources in advance for high-priority tasks or tasks with relatively large expected resource requirement growth to ensure the smooth execution of critical tasks.

[0027] The advantages of the present invention compared with the prior art are as follows: Through the comprehensive preprocessing of tasks and the real-time and accurate evaluation of the resource status of computing nodes, the present invention can provide an accurate basis for task allocation, improving the rationality and accuracy of task allocation.

[0028] The present invention uses a deep Q-network for task allocation, enabling the system to continuously learn and optimize the task allocation strategy, adapt to complex and changing task and resource environments, and significantly improve resource utilization and task execution efficiency.

[0029] Through the mutual cooperation of the task execution monitoring module and the resource dynamic adjustment module, the present invention can timely discover and handle problems during task execution. By dynamically adjusting resource allocation, it ensures the continuous and stable execution of tasks, further improving the reliability and overall performance of the system. Brief Description of the Drawings

[0030] Figure 1 is a system block diagram of a computing resource scheduling system for a distributed AI training task of the present invention.

[0031] Figure 2 It is a flowchart of a method for using a deep Q - network to assign AI training tasks to appropriate computing nodes.

[0032] Figure 3 It is a flowchart of using LSTM to predict future resource requirements. Detailed implementation manners

[0033] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.

[0034] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present invention and its application or use.

[0035] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification.

[0036] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of exemplary embodiments may have different values.

[0037] The following further details a computing resource scheduling system for a distributed AI training task of the present invention with reference to the accompanying drawings.

[0038] Combined with the attached Figures 1-3 , the present invention is introduced in detail.

[0039] A computing resource scheduling system for a distributed AI training task, comprising:

[0040] Task receiving module: This module is responsible for receiving multiple AI training tasks and pre - processing them. In the task parsing process, key information such as the scale of the dataset required by the task, data format, algorithm complexity, and dependent libraries will be deeply analyzed. The task priority is determined by comprehensively considering the business urgency of the task, its impact weight on the overall business, and the estimated execution duration. The task type is identified based on the type of machine learning algorithm and training objective used in the task. For tasks with complex dependencies, this module constructs a task dependency graph to clearly show the sequence and data flow between tasks.

[0041] Resource Evaluation Module: This module is used to evaluate the resource status of each computing node in a distributed computing environment. The resource status of a computing node covers the free amount and utilization rate of the CPU, memory, and GPU, and also includes key resource metrics such as network bandwidth and disk I / O read / write speed. It uses a combination of periodic evaluation and event-triggered evaluation to update the resource status of each computing node in real time. Periodic evaluation comprehensively scans the resources of all computing nodes at fixed time intervals to obtain detailed data such as the CPU utilization rate, memory free amount, GPU computing core usage, network bandwidth occupancy rate, and disk I / O read / write rate of each node. Event-triggered evaluation immediately updates the resource status information of the corresponding node when there are mutations in the resource status of a computing node, new nodes are added, or old nodes fail and exit, ensuring that the system always has the latest resource dynamics.

[0042] Task Allocation Module: This module allocates AI training tasks to appropriate computing nodes using a deep Q-network based on the requirements of the AI training tasks received by the task receiving module and the resource status of each computing node evaluated by the resource evaluation module. The specific implementation method is as follows:

[0043] Define the state space: Combine the free amount and utilization rate of the CPU, memory, and GPU of each computing node, as well as the network bandwidth, disk I / O read / write speed, and the type, priority, and resource requirements of the current task to be allocated into a high-dimensional vector as the state.

[0044] Define the action space: An action represents allocating the current task to a certain computing node, and the action space is the set of all available computing nodes.

[0045] Define the reward function: The reward function is used to measure the quality of each action, and the reward value is comprehensively calculated based on the expected execution duration and resource utilization rate of the task. If the overall resource utilization rate can be improved and the expected execution duration can be shortened after the task is allocated, a positive reward is given; otherwise, a negative reward is given.

[0046] Train the DQN network: By continuously interacting with the environment, the agent selects actions based on the current state, obtains a new state and reward after executing the actions, and uses this data to update the parameters of the DQN network to achieve an optimal task allocation strategy. In the process of continuously trying to allocate tasks to different computing nodes, the agent gradually learns how to make the best task allocation decisions according to different tasks and resource states based on the obtained reward feedback.

[0047] Task Execution Monitoring Module: This module is used to monitor the execution status of assigned tasks, including the execution progress of tasks (measured by indicators such as the number of completed training rounds and data processing volume), resource usage (real-time tracking of the occupancy rates of resources such as CPU, memory, and GPU), and whether anomalies occur. This module also includes an anomaly handling mechanism. When a task execution anomaly (such as program crash or resource exhaustion error) is detected, the task is automatically restarted; if multiple restarts fail, according to the real-time resource status provided by the resource evaluation module, the task is re-assigned to other suitable computing nodes. At the same time, the execution progress and resource usage of the task are presented to the system administrator in the form of charts, facilitating the administrator to intuitively understand the task execution status, discover problems in a timely manner, and take measures.

[0048] Resource Dynamic Adjustment Module: This module dynamically adjusts the resource allocation of computing nodes according to the information fed back by the task execution monitoring module. Based on historical task execution data and the current resource status, LSTM is used to predict future resource requirements and perform resource pre-allocation in advance. The specific steps for using LSTM to predict future resource requirements are as follows:

[0049] Data Preprocessing: Collect the resource usage data of historical tasks, perform normalization processing, and scale the data to the required range.

[0050] Construct LSTM Model: Build a neural network model including an input layer, an LSTM hidden layer, and an output layer. The input layer receives historical resource usage data, the LSTM hidden layer captures the time series features in the data, and the output layer predicts future resource requirements.

[0051] Model Training: Use historical data to train the LSTM model, and adjust the weight parameters of the model through the backpropagation algorithm to minimize the error between the prediction results of the model and the actual values. During the training process, continuously optimize the model parameters to improve the prediction accuracy.

[0052] Resource Pre-allocation: According to the prediction results of the LSTM model, reserve a certain amount of computing resources in advance for future possible resource requirements. At the same time, during the task execution process, dynamically adjust the resource allocation according to the real-time resource usage and task progress to ensure the reasonable allocation of resources among tasks. During the task execution process, if it is found that a certain task has low resource utilization efficiency, its resource allocation can be appropriately reduced, and the resources can be allocated to other tasks that need them more. This module also includes a resource reservation mechanism to reserve a certain proportion of computing resources in advance for high-priority tasks or tasks with relatively large expected resource demand growth to ensure the smooth execution of key tasks.

[0053] The specific implementation process of a computing resource scheduling system for distributed AI training tasks of the present invention is as follows:

[0054] System Architecture Deployment

[0055] Task receiving module: Deployed on a high-performance front-end server, it receives AI training tasks submitted by operators or other systems through a network interface.

[0056] Resource evaluation module: Install a resource monitoring agent on each computing node, and the central server is responsible for aggregating and processing this data.

[0057] Task allocation module: Runs on a scheduling server with powerful computing capabilities, which is used to execute the deep Q-network algorithm.

[0058] Task execution monitoring module: Runs a monitoring program on each computing node, and at the same time sets up a monitoring server to collect and display monitoring data.

[0059] Resource dynamic adjustment module: Deployed on a server with data analysis and processing capabilities, it is responsible for running the LSTM model and performing resource pre-allocation and dynamic adjustment.

[0060] Specific implementation process

[0061] 1. Task receiving module

[0062] The operator submits an AI training task to the task receiving module through a web interface or a command-line tool. After receiving the task, the task receiving module immediately performs preprocessing:

[0063] Task parsing: The operator submits an image classification task based on deep learning. The task parsing program will analyze the task script and find that this task requires the use of a dataset containing 100,000 images, the data format is JPEG, and it uses the ResNet-50 convolutional neural network and depends on the PyTorch deep learning framework.

[0064] Priority determination: Considering that this task is an important scientific research project, has a great impact on the overall scientific research progress, and is expected to have a long execution time (about 24 hours), the task receiving module sets its priority to high.

[0065] Task type identification: Determine that this task is an image classification type according to the convolutional neural network and image classification target used in the task script.

[0066] Construct a task dependency graph: If this task has complex dependencies and the image data needs to be preprocessed before model training, the task receiving module will construct a task dependency graph to clearly show the sequence and data flow between the preprocessing task and the training task.

[0067] 2. Resource evaluation module

[0068] Periodic evaluation: The central server of the resource evaluation module sets a timed task that executes once every 5 minutes. Within each period, the central server sends resource data collection requests to the resource monitoring agents of all computing nodes. The agents collect information such as the free amount and usage rate of the CPU, memory, and GPU of the computing nodes and return the data to the central server. For example, the CPU usage rate of computing node A is 30%, the free memory amount is 8GB, and the GPU usage rate is 20%.

[0069] Event-triggered evaluation: When the GPU usage rate of computing node B suddenly soars from 20% to 90%, its resource monitoring agent immediately detects this resource state mutation event and sends the updated resource information to the central server. After receiving the information, the central server immediately updates the resource state of computing node B. In addition, when a new computing node joins the cluster or an old node fails and exits, it will also trigger corresponding resource state updates.

[0070] 3. Task allocation module

[0071] The task allocation module uses a deep Q-network to allocate tasks based on the task requirements provided by the task receiving module and the resource status of the computing nodes provided by the resource evaluation module:

[0072] Define the state space: Combine the free amount and usage rate of the CPU, memory, and GPU of each computing node, as well as the network bandwidth, disk I / O read and write speed, and the type, priority, and resource requirements of the currently to-be-allocated image classification task (such as 2GB of memory and 30% of GPU computing resources) into a high-dimensional vector as the state. For example, the state information of computing node C is: CPU idle rate 50%, free memory amount 10GB, GPU usage rate 10%, remaining network bandwidth 200Mbps, disk I / O read and write speed 80MB / s, and the task information is image classification type, high priority, 2GB memory requirement, 30% GPU requirement.

[0073] Define the action space: Assume there are 10 available computing nodes in the cluster, and the action space is the set of the numbers of these 10 nodes {1, 2, 3, …, 10}.

[0074] Define the reward function: For the action of allocating the image classification task to computing node C, if the overall resource utilization rate is expected to increase from 60% to 70% after allocation and the expected execution duration is shortened from the original 24 hours to 20 hours, a positive reward value (such as +10) is given; conversely, if the resource utilization rate decreases and the execution duration increases after allocation, a negative reward value (such as -5) is given.

[0075] Training the DQN network: The agent selects an action based on the current state (such as allocating a task to computing node C). After executing the action, a new state and reward are obtained. These data are stored in the experience replay pool, and data is periodically sampled from the experience replay pool to train and update the DQN network, continuously optimizing the task allocation strategy.

[0076] 4. Task Execution Monitoring Module

[0077] The monitoring program running on each computing node in the task execution monitoring module will monitor the task execution situation in real time:

[0078] Execution Progress Monitoring: For an image classification task, the monitoring program reads the training log and finds that 50 rounds of training have been completed, with a total plan of 100 rounds. Therefore, it is judged that the task execution progress is 50%.

[0079] Resource Usage Monitoring: The monitoring program tracks the resource usage of computing node C in real time and finds that the CPU usage rate is 40%, the memory usage rate is 60%, and the GPU usage rate is 35%.

[0080] Exception Handling Mechanism: If during the task execution process, the monitoring program detects that the GPU has overheated and caused the program to crash. The monitoring program will automatically restart the task. If the restart fails 3 times in a row, the monitoring program will send the exception information and task information to the task allocation module. The task allocation module, based on the real-time resource status provided by the resource evaluation module, reallocates the task to other suitable computing nodes (such as computing node D).

[0081] Visualization Display: The monitoring program transmits the task execution progress and resource usage data to the monitoring server through the network. The monitoring server uses a data visualization tool (such as Grafana) to plot the data into charts for the system administrator to view.

[0082] 5. Resource Dynamic Adjustment Module

[0083] The resource dynamic adjustment module predicts future resource requirements using LSTM based on historical task execution data and the current resource status, and performs resource pre-allocation and dynamic adjustment:

[0084] Data Preprocessing: The resource dynamic adjustment module collects the resource usage data of all AI training tasks in the past week, including the time series data of CPU, memory, and GPU usage. These data are normalized and scaled to the range of 0 - 1 for processing by the LSTM model.

[0085] Construct an LSTM model: Build a neural network model that includes an input layer, an LSTM hidden layer, and an output layer. The input layer receives historical resource usage data, the LSTM hidden layer captures the time series features in the data, and the output layer predicts the resource requirements for the next 2 hours.

[0086] Model training: Use historical data to train the LSTM model, and continuously adjust the weight parameters of the model through the backpropagation algorithm to minimize the error between the predicted results of the model and the actual values.

[0087] Resource pre-allocation: According to the prediction results of the LSTM model, it is found that the GPU requirements for image classification tasks will increase significantly within the next 2 hours. The resource dynamic adjustment module reserves a certain proportion (such as 40%) of GPU resources in advance for these tasks.

[0088] Dynamic adjustment: During the execution of the tasks, it is monitored that the resource utilization efficiency of another task on computing node D is low, and its GPU utilization rate is only 10%. The resource dynamic adjustment module allocates a part (such as 20%) of the GPU resources of this task to the image classification task through the API interface of the resource management system to ensure the reasonable allocation of resources among tasks. At the same time, a certain proportion of computing resources is reserved for high-priority image classification tasks to ensure the smooth execution of key tasks.

[0089] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative work without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A computing resource scheduling system for distributed AI training tasks, characterized in that: including a task receiving module, configured to receive multiple AI training tasks and perform preprocessing, where the preprocessing includes task parsing, priority determination, and task type identification; a resource evaluation module, configured to evaluate the resource status of each computing node in a distributed computing environment, and update the resource status of each computing node in real time by combining periodic evaluation and event-triggered evaluation; a task allocation module, which, according to the requirements of the AI training tasks received by the task receiving module and the resource status of each computing node evaluated by the resource evaluation module, uses a deep Q-network to allocate the AI training tasks to appropriate computing nodes; a task execution monitoring module, configured to monitor the execution of the allocated tasks, including the execution progress of the tasks, resource usage, and whether anomalies occur; a resource dynamic adjustment module, which dynamically adjusts the resource allocation of computing nodes according to the information fed back by the task execution monitoring module.

2. The computing resource scheduling system for a distributed AI training task according to claim 1, wherein: When parsing tasks, the task receiving module deeply analyzes the scale of the dataset required by the task, data format, algorithm complexity, and dependent libraries; determines the task priority by considering the business urgency of the task, the impact weight on the overall business, and the estimated execution duration; and identifies the task type according to the type of machine learning algorithm and training objective adopted by the task.

3. The computing resource scheduling system for a distributed AI training task according to claim 2, characterized in that: For tasks with complex dependencies, the task receiving module presents the sequence and data flow between tasks by constructing a task dependency graph.

4. The computing resource scheduling system for a distributed AI training task according to claim 3, wherein: The periodic evaluation of the resource evaluation module comprehensively scans the resources of all computing nodes at fixed time intervals; the event-triggered evaluation immediately updates the resource status information of the corresponding nodes when there are mutations in the resource status of computing nodes, new nodes are added, or old nodes fail and exit.

5. The computing resource scheduling system for a distributed AI training task according to claim 4, characterized in that: The resource status of the computing nodes includes the free amount and usage rate of the CPU, memory, and GPU.

6. The computing resource scheduling system for a distributed AI training task according to claim 5, wherein: The method of using a deep Q-network to allocate AI training tasks to appropriate computing nodes is as follows: Define the state space: Combine the free amount and usage rate of the CPU, memory, and GPU of each computing node, as well as network bandwidth, disk I / O read and write speed, and the type, priority, and resource requirements of the currently to-be-allocated task into a high-dimensional vector as the state; Define the action space: An action represents allocating the current task to a certain computing node, and the action space is the set of all available computing nodes; Define the reward function: The reward function is used to measure the quality of each action, and the reward value is comprehensively calculated based on the expected execution duration and resource utilization rate of the task execution; if the overall resource utilization rate can be improved and the expected execution duration can be shortened after the task is allocated, a positive reward is given; otherwise, a negative reward is given; Train the DQN network: By continuously interacting with the environment, the agent selects actions according to the current state, obtains a new state and reward after executing the actions, and uses this data to update the parameters of the DQN network to achieve an optimal task allocation strategy.

7. The computing resource scheduling system for a distributed AI training task according to claim 6, wherein: The task execution monitoring module further includes an exception handling mechanism. When a task execution exception is detected, the task is automatically restarted. If the restart fails multiple times, the task is reallocated to other suitable computing nodes according to the real-time resource status provided by the resource evaluation module. At the same time, the execution progress and resource usage of the task are presented to the system administrator in the form of a chart.

8. The computing resource scheduling system for a distributed AI training task according to claim 7, wherein: The resource dynamic adjustment module predicts future resource requirements using LSTM based on historical task execution data and the current resource status, and performs resource pre-allocation in advance. The specific steps for predicting future resource requirements using LSTM are as follows: Data preprocessing: Collect the resource usage data of historical tasks, perform normalization processing, and scale the data to the required range. Construct an LSTM model: Build a neural network model including an input layer, an LSTM hidden layer, and an output layer. Among them, the input layer receives historical resource usage data, the LSTM hidden layer captures the time series features in the data, and the output layer predicts future resource requirements. Model training: Use historical data to train the LSTM model, and adjust the weight parameters of the model through the backpropagation algorithm to minimize the error between the prediction result of the model and the actual value. Resource pre-allocation: According to the prediction result of the LSTM model, reserve a certain amount of computing resources in advance for future possible resource requirements. At the same time, during the task execution process, dynamically adjust the resource allocation according to the real-time resource usage and task progress to ensure the reasonable allocation of resources among tasks.

9. The computing resource scheduling system for a distributed AI training task according to claim 8, characterized in that: The resource dynamic adjustment module further includes a resource reservation mechanism, which reserves a certain proportion of computing resources in advance for high-priority tasks or tasks with relatively large expected resource requirement growth to ensure the smooth execution of critical tasks.

Citation Information

Cited By

  • GPU resource dynamic allocation method and system based on load awareness

    CN120832243A