Large model training-oriented GPU (Graphics Processing Unit) cluster computing power optimization architecture

By real-time monitoring and managing the resource status of GPU clusters, combined with multi-factor evaluation and dynamic optimization strategies, the problem of low resource utilization efficiency of GPU clusters in large model training is solved, and efficient, stable and economical computing power management is achieved.

CN120448030AInactive Publication Date: 2025-08-08XIANGTAN UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510495318.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

During the training of large-scale model, the computing power resource state of the GPU cluster is complex and changeable. How to monitor and manage these resources in real time to ensure efficient utilization, and how to dynamically allocate resources and optimize computing power according to task needs to improve overall utilization efficiency is the difficulty of the existing technology.

Method used

The computing power resource management module, task scheduling module, computing power optimization module and result feedback module are adopted to monitor the GPU status through high-precision sensors, combine multi-factor evaluation and allocation strategies, support dynamic adjustment and load balancing, realize multi-task parallel processing and fault warning, and use blockchain technology to ensure data security.

Benefits of technology

It improves the computing power utilization efficiency of the GPU cluster, reduces training costs, enhances system stability and ease of use, and improves the effects of fault warning and user interaction interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448030A_ABST
    Figure CN120448030A_ABST
Patent Text Reader

Abstract

A GPU cluster computing power optimization architecture oriented to large model training is characterized in that a computing power resource management module is used for monitoring and managing the computing power resource state of each GPU in a GPU cluster in real time, a task scheduling module is used for allocating GPU resources according to the requirements of training tasks, and a computing power optimization module is used for dynamically optimizing the computing power of the GPU cluster in the task execution process. And the result feedback module is used for collecting and analyzing the optimized computing power use condition so as to adjust a subsequent task scheduling strategy. According to the method, the computing power resource state of the GPU cluster is monitored and managed in real time, the GPU resources are dynamically allocated according to the demand of the training task, and the computing power of the GPU cluster is dynamically optimized in the task execution process, so that the computing power utilization efficiency of the GPU cluster is effectively improved, and the training cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of GPU large model training technology, and specifically provides a GPU cluster computing power optimization architecture for large model training. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, large-scale model training has become a crucial application scenario in high-performance computing (HPC). However, the computing power resources of GPU clusters often face significant challenges during large-scale model training. The complex and ever-changing nature of computing resources within GPU clusters necessitates real-time monitoring and management of these resources to ensure their efficient utilization. Furthermore, the diverse demands of training tasks necessitate dynamic allocation of GPU resources based on these requirements and dynamic optimization of computing power during task execution to improve overall computing power utilization efficiency, both hot topics and challenges in current research. Summary of the Invention

[0003] In view of the above situation, in order to overcome the shortcomings of the existing technology, the present invention provides a GPU cluster computing power optimization architecture for large model training to at least partially solve the above technical problems.

[0004] The technical solution adopted by the present invention is as follows:

[0005] This paper proposes a GPU cluster computing power optimization architecture for large model training, including:

[0006] Computing power resource management module, task scheduling module, computing power optimization module and result feedback module; the computing power resource management module is used to monitor and manage the computing power resource status of each GPU in the GPU cluster in real time, the task scheduling module is used to allocate GPU resources according to the requirements of training tasks, the computing power optimization module is used to dynamically optimize the computing power of the GPU cluster during task execution, and the result feedback module is used to collect and analyze the optimized computing power usage to adjust subsequent task scheduling strategies;

[0007] The computing resource management module also includes:

[0008] The computing resource management module has a built-in efficient computing power monitoring unit. This unit uses high-precision sensors and advanced data acquisition technology to obtain real-time status information such as computing power load, memory usage, temperature, and power consumption of each GPU in the GPU cluster, ensuring data accuracy and real-time performance.

[0009] The computing power allocation unit in the computing power resource management module adopts a computing power allocation strategy based on a comprehensive evaluation of multiple factors. This computing power allocation strategy not only considers the computing power load and memory usage of the GPU, but also combines status information such as GPU temperature and power consumption, as well as factors such as the urgency of the training task, the amount of computing power required, and the expected execution time. It allocates corresponding GPU resources to each task to maximize the utilization of GPU resources.

[0010] The computing resource management module supports dynamic adjustment and load balancing of GPU resources. When the computing load of a GPU is too high or the temperature is abnormal, the computing resource allocation unit can automatically adjust the allocation of GPU resources and migrate some tasks to other idle or less loaded GPUs to achieve load balancing of the GPU cluster, avoiding waste of computing resources and performance degradation caused by GPU overheating.

[0011] The computing resource management module also has the functions of reserving and elastically expanding GPU resources. It can reserve a certain number of GPU resources in advance according to the requirements of training tasks, and dynamically increase GPU resources when needed to meet the needs of high-performance computing scenarios such as large-scale model training.

[0012] The computing resource management module adopts a data security and privacy protection mechanism based on blockchain technology to ensure the security and traceability of the computing resource status information and task scheduling information of each GPU in the GPU cluster, and prevent data leakage and malicious attacks.

[0013] In one embodiment of the present invention, the task scheduling module adopts a priority-based task scheduling algorithm. The task scheduling algorithm not only considers factors such as the urgency of the training task, the amount of computing resources required, and the estimated execution time, but also combines the computing resource status and load balancing of the GPU cluster to allocate corresponding GPU resources to each task and generate a task execution plan.

[0014] The task scheduling module further includes:

[0015] The task scheduling module supports multi-task parallel processing. It can schedule multiple training tasks to execute in parallel based on the computing resource status and load balancing of the GPU cluster, thereby improving the computing power utilization efficiency of the GPU cluster.

[0016] The task scheduling module has the function of dynamically adjusting task priorities. It can dynamically adjust the priority of tasks according to the urgency and execution progress of training tasks to ensure that urgent tasks are handled first.

[0017] The task scheduling module also has a task preemption mechanism. When a new urgent task appears, it can preempt the low-priority task being executed to ensure that the urgent task can be processed in time.

[0018] The task scheduling module supports custom task scheduling strategies. Users can customize task scheduling strategies according to actual needs to meet the requirements of specific application scenarios.

[0019] In one embodiment of the present invention, the task scheduling module adopts a priority-based task scheduling algorithm to allocate corresponding GPU resources to each task and generate a task execution plan based on factors such as the urgency of the training task, the size of the required computing resources and the expected execution time.

[0020] In one embodiment of the present invention, the computing power optimization module includes a computing power prediction unit, a computing power scheduling unit and a computing power adjustment unit; the computing power prediction unit is used to predict the computing power demand of the GPU cluster in the future period of time, and the computing power scheduling unit dynamically adjusts the allocation of GPU resources according to the computing power demand and the current computing power status of the GPU cluster; the computing power adjustment unit is used to fine-tune the computing power of the GPU according to the real-time computing power load during task execution to improve the utilization efficiency of the computing power.

[0021] In one embodiment of the present invention, the computing power prediction unit uses a machine learning algorithm to train historical computing power usage data to establish a computing power demand prediction model. The computing power demand prediction model can predict the computing power demand in the future based on the current training task information and GPU cluster status.

[0022] In one embodiment of the present invention, the result feedback module includes a data analysis unit and a strategy adjustment unit; the data analysis unit is used to collect and analyze the optimized computing power usage, including the computing power utilization rate of each task, the task execution time and the overall computing power load of the GPU cluster, etc. The strategy adjustment unit adjusts the subsequent computing power allocation strategy and task scheduling strategy according to the analysis results to improve the computing power utilization efficiency of the GPU cluster.

[0023] In one embodiment of the present invention, a fault warning module is further included; the fault warning module is used to monitor the operating status of each GPU in the GPU cluster in real time. When a GPU failure or abnormality is detected, a warning message is issued in a timely manner and a corresponding fault handling process is started.

[0024] In one embodiment of the present invention, the fault warning module uses a deep learning-based fault prediction algorithm to train historical fault data to establish a fault prediction model. The fault prediction model can predict possible future fault conditions based on current GPU operating status information.

[0025] In one embodiment of the present invention, a user interaction interface is also included; the user interaction interface is used to display the computing resource status, task execution status, optimization results and fault warning information of the GPU cluster, and provide corresponding operation options for users to intervene and adjust.

[0026] In one embodiment of the present invention, the user interaction interface adopts a graphical interface design, which can intuitively display the computing resource distribution, task execution progress and optimization effect of the GPU cluster, and at the same time provide convenient operation tools so that users can remotely monitor and manage the GPU cluster.

[0027] The beneficial effects of the technical solution of the present invention are:

[0028] The present invention monitors and manages the computing power resource status of the GPU cluster in real time, dynamically allocates GPU resources according to the requirements of the training task, and dynamically optimizes the computing power of the GPU cluster during task execution, thereby effectively improving the computing power utilization efficiency of the GPU cluster and reducing training costs. At the same time, through the setting of the fault warning module and the user interaction interface, the stability and usability of the GPU cluster are further improved.

[0029] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0031] Figure 1 A schematic diagram of a GPU cluster computing power optimization architecture for large model training proposed in an embodiment of the present invention;

[0032] Figure 2 A schematic diagram of a computing resource management module proposed in an embodiment of the present invention;

[0033] Figure 3 This is a schematic diagram of a task scheduling module proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0035] The following describes a GPU cluster computing power optimization architecture for large model training according to an embodiment of the present invention with reference to the accompanying drawings.

[0036] like Figures 1 to 3 As shown, an embodiment of the present invention provides a GPU cluster computing power optimization architecture for large model training, including: a computing power resource management module, a task scheduling module, a computing power optimization module and a result feedback module; the computing power resource management module is used to monitor and manage the computing power resource status of each GPU in the GPU cluster in real time, the task scheduling module is used to allocate GPU resources according to the requirements of the training task, the computing power optimization module is used to dynamically optimize the computing power of the GPU cluster during the task execution process, and the result feedback module is used to collect and analyze the optimized computing power usage to adjust the subsequent task scheduling strategy;

[0037] The computing resource management module also includes:

[0038] The computing resource management module has a built-in efficient computing power monitoring unit. Through high-precision sensors and advanced data acquisition technology, the computing power monitoring unit obtains real-time status information such as computing power load, memory usage, temperature, and power consumption of each GPU in the GPU cluster, ensuring the accuracy and real-time nature of the data.

[0039] The computing power allocation unit in the computing power resource management module adopts a computing power allocation strategy based on a comprehensive evaluation of multiple factors. The computing power allocation strategy not only considers the computing power load and memory usage of the GPU, but also combines status information such as GPU temperature and power consumption, as well as factors such as the urgency of the training task, the amount of computing power required, and the expected execution time. It allocates appropriate GPU resources to each task to maximize the utilization of GPU resources.

[0040] The computing resource management module supports dynamic adjustment and load balancing of GPU resources. When the computing load of a GPU is too high or the temperature is abnormal, the computing resource allocation unit can automatically adjust the allocation of GPU resources and migrate some tasks to other idle or less loaded GPUs to achieve load balancing of the GPU cluster, avoiding waste of computing resources and performance degradation caused by GPU overheating.

[0041] The computing resource management module also has the functions of reserving and elastically expanding GPU resources. It can reserve a certain number of GPU resources in advance according to the requirements of training tasks, and dynamically increase GPU resources when needed to meet the needs of high-performance computing scenarios such as large-scale model training.

[0042] The computing resource management module adopts a data security and privacy protection mechanism based on blockchain technology to ensure the security and traceability of the computing resource status information and task scheduling information of each GPU in the GPU cluster, and prevent data leakage and malicious attacks.

[0043] In specific applications of the embodiments of the present invention, the computing power resource management module serves as the basis of the architecture. Through the built-in computing power monitoring unit, it uses high-precision sensors and advanced data acquisition technology to obtain key status information such as the computing power load, memory usage, temperature and power consumption of each GPU in the GPU cluster in real time and accurately.

[0044] After obtaining real-time status information, the computing power allocation unit intelligently allocates GPU resources for each training task based on a computing power allocation strategy based on a comprehensive evaluation of multiple factors. This strategy comprehensively considers the GPU's current load, memory usage, temperature, power consumption, and other status, as well as the urgency of the training task, the required computing power resources, and the estimated execution time. This strategy ensures maximum utilization of GPU resources while avoiding performance bottlenecks and computing power waste caused by improper resource allocation.

[0045] During task execution, the computing power optimization module dynamically adjusts and optimizes the computing power resources of the GPU cluster based on real-time monitoring of GPU status information and task execution status. For example, when the computing power load of a GPU is too high or the temperature is abnormal, the module will trigger the dynamic adjustment mechanism of computing power resources, migrating some tasks to other idle or less loaded GPUs, thereby achieving load balancing of the GPU cluster and avoiding excessive concentration of computing power resources and performance degradation caused by GPU overheating.

[0046] In addition, the computing resource management module also has the function of reserving and elastically expanding GPU resources. Based on the demand forecast of the training task, the module can reserve a certain amount of GPU resources in advance to ensure a quick response when the task arrives. At the same time, during the execution of the task, if there is a shortage of computing resources, the module can also dynamically increase GPU resources to meet the needs of high-performance computing scenarios such as large-scale model training. To ensure the security and traceability of the computing resource status information and task scheduling information of each GPU in the GPU cluster, the computing resource management module adopts a data security and privacy protection mechanism based on blockchain technology. The mechanism ensures the confidentiality, integrity and traceability of information through encrypted storage and distributed ledger technology, effectively preventing the risk of data leakage and malicious attacks.

[0047] In one possible implementation, the task scheduling module uses a priority-based task scheduling algorithm. This algorithm not only considers factors such as the urgency of the training task, the amount of computing resources required, and the estimated execution time, but also combines the computing resource status and load balancing of the GPU cluster to allocate corresponding GPU resources to each task and generate a task execution plan.

[0048] The task scheduling module also includes:

[0049] The task scheduling module supports multi-task parallel processing. It can schedule multiple training tasks to execute in parallel based on the computing resource status and load balancing of the GPU cluster, thereby improving the computing power utilization efficiency of the GPU cluster.

[0050] The task scheduling module has the function of dynamically adjusting task priorities. It can dynamically adjust the priority of tasks according to the urgency and execution progress of training tasks to ensure that urgent tasks are handled first.

[0051] The task scheduling module also has a task preemption mechanism. When a new urgent task appears, it can preempt the low-priority task being executed to ensure that the urgent task can be processed in time.

[0052] The task scheduling module supports custom task scheduling strategies. Users can customize task scheduling strategies according to actual needs to meet the requirements of specific application scenarios.

[0053] In specific applications, the task scheduling algorithm of the embodiments of the present invention not only comprehensively considers key factors such as the urgency of the training task, the amount of computing resources required, and the estimated execution time, but also deeply analyzes the current computing resource status of the GPU cluster, including the computing load, memory usage, temperature, and power consumption of each GPU, as well as the overall load balancing of the cluster. Based on this information, the algorithm can accurately match the most appropriate GPU resources for each training task, ensuring efficient and stable execution of the task.

[0054] The task scheduling module supports multi-task parallel processing. It can schedule multiple training tasks for parallel execution based on the GPU cluster's computing resource status and load balancing, improving the GPU cluster's computing power utilization efficiency and shortening the overall task execution time. To further enhance task flexibility and responsiveness, the task scheduling module also features dynamic task priority adjustment. This intelligently adjusts task priorities in real time based on the urgency and execution progress of training tasks. When urgent tasks arise, the system can respond quickly to ensure they receive priority processing, thus meeting the diverse needs of practical application scenarios.

[0055] The task scheduling module also features a built-in task preemption mechanism. When a new urgent task is inserted into the system, it can preempt GPU resources occupied by currently executing low-priority tasks, ensuring that urgent tasks are processed promptly, improving system flexibility and responsiveness, and ensuring efficient execution of urgent tasks. To meet the specific needs of different users, the task scheduling module also supports customizable task scheduling policies. Users can customize task scheduling policies based on their application scenarios and requirements, such as setting specific task priority rules and defining conditions for task preemption.

[0056] In one possible implementation, the task scheduling module adopts a priority-based task scheduling algorithm to allocate corresponding GPU resources to each task and generate a task execution plan based on factors such as the urgency of the training task, the amount of computing resources required, and the expected execution time.

[0057] The computing power optimization module includes a computing power prediction unit, a computing power scheduling unit, and a computing power adjustment unit; the computing power prediction unit is used to predict the computing power demand of the GPU cluster in the future; the computing power scheduling unit dynamically adjusts the allocation of GPU resources based on the computing power demand and the current computing power status of the GPU cluster; the computing power adjustment unit is used to fine-tune the GPU computing power according to the real-time computing power load during task execution to improve the utilization efficiency of computing power.

[0058] The computing power prediction unit uses a machine learning algorithm to train historical computing power usage data to establish a computing power demand prediction model. The computing power demand prediction model can predict the computing power demand in the future based on the current training task information and GPU cluster status.

[0059] In specific applications of the embodiments of the present invention, the computing power prediction unit uses advanced machine learning algorithms to conduct in-depth training and learning of historical computing power usage data, successfully building a computing power demand prediction model. This model can accurately predict computing power demand in the future based on current training task information and GPU cluster status. Based on the prediction data provided by the computing power prediction unit, the computing power scheduling unit can grasp the changing trends of future computing power demand in real time. Based on these trends, combined with the current computing power status of the GPU cluster, the computing power scheduling unit can dynamically adjust the GPU resource allocation strategy.

[0060] During task execution, the computing power adjustment unit is responsible for real-time monitoring of the GPU's computing power load. Once it finds that the actual computing power load does not meet expectations, the computing power adjustment unit will immediately activate the fine-tuning mechanism to accurately adjust the GPU's computing power. The real-time fine-tuning function ensures that the GPU resources can always maintain the best working state, thereby further improving the utilization efficiency of computing power.

[0061] In one possible implementation, the result feedback module includes a data analysis unit and a strategy adjustment unit; the data analysis unit is used to collect and analyze the optimized computing power usage, including the computing power utilization rate of each task, the task execution time and the overall computing power load of the GPU cluster, etc. The strategy adjustment unit adjusts the subsequent computing power allocation strategy and task scheduling strategy based on the analysis results to improve the computing power utilization efficiency of the GPU cluster.

[0062] In specific applications of the embodiments of the present invention, the data analysis unit is responsible for comprehensively collecting and analyzing the usage of the optimized GPU cluster computing power, including but not limited to the computing power utilization rate of each training task, the actual execution time of the task, and the overall computing power load status of the GPU cluster. Through in-depth analysis of these detailed data, the data analysis unit can accurately reveal the current status of computing power usage and its potential problems. The data analysis unit will use advanced statistical analysis methods and data mining technology to conduct a comprehensive analysis of the collected data. This process not only includes quantitative evaluation of various indicators, but also involves the prediction and interpretation of computing power usage trends. Through these in-depth analysis work, the data analysis unit can provide scientific and accurate data support for the strategy adjustment unit.

[0063] Based on the detailed analysis results provided by the data analysis unit, the policy adjustment unit will take on the important task of adjusting the subsequent computing power allocation strategy and task scheduling strategy. This process aims to further improve the computing power utilization efficiency of the GPU cluster through intelligent policy adjustment. The policy adjustment unit will formulate targeted improvement plans based on the problems and deficiencies revealed in the analysis results. For example, for tasks with low computing power utilization, the policy adjustment unit will consider reducing the GPU resources allocated to them or merging them with other tasks. For tasks with long execution times, their execution time will be shortened by optimizing algorithms or increasing parallelism.

[0064] At the same time, the policy adjustment unit also closely monitors the overall computing load of the GPU cluster to ensure that the cluster always maintains an efficient and stable working state. When the cluster load is too high, the policy adjustment unit will promptly activate the load balancing mechanism to reduce the load through reasonable task scheduling and resource allocation; when the cluster load is too low, it will increase the load level by increasing the number of tasks or increasing the complexity of tasks.

[0065] In one possible implementation, a fault warning module is further included; the fault warning module is used to monitor the operating status of each GPU in the GPU cluster in real time. When a GPU failure or abnormality is detected, a warning message is issued in a timely manner and a corresponding fault handling process is initiated.

[0066] The fault warning module uses a deep learning-based fault prediction algorithm to train historical fault data to establish a fault prediction model. The fault prediction model can predict possible future faults based on the current GPU operating status information.

[0067] In the specific application of the embodiment of the present invention, the fault warning module lies in the fault prediction algorithm based on deep learning it adopts. This algorithm first conducts in-depth mining and analysis of historical fault data to train an accurate fault prediction model. The model has strong learning and generalization capabilities and can make intelligent predictions of future fault conditions based on the current GPU operating status information, including but not limited to key indicators such as temperature, power consumption, load rate, and memory occupancy.

[0068] During actual operation, the fault warning module continuously receives operating status data from each node in the GPU cluster. After preprocessing, this data is input into the trained fault prediction model. Through complex calculations and analysis, the model can evaluate the health status of each GPU in real time and predict the type and probability of faults that may occur in the future.

[0069] Once the model predicts a potential failure risk for a GPU, the Fault Warning Module immediately triggers an alert mechanism. This mechanism includes sending an alert message to the system administrator or designated monitoring platform. The message details the predicted failure, the failure type, the scope of impact, and recommended countermeasures. Furthermore, the Fault Warning Module automatically or manually initiates emergency response mechanisms based on pre-defined fault handling procedures, such as switching to a backup GPU, reducing load, or restarting the node, to ensure stable cluster operation and data security.

[0070] The fault warning module doesn't exist in isolation; instead, it works closely with other modules within the entire GPU cluster's computing power optimization architecture. For example, when the warning module issues a warning, the task scheduling module can promptly adjust task allocation strategies to avoid assigning new training tasks to the faulty GPU. The computing power optimization module can dynamically adjust computing resource allocation based on the fault's severity, ensuring that the cluster's overall computing power output is not affected.

[0071] In one possible implementation, a user interaction interface is further included; the user interaction interface is used to display the computing resource status, task execution status, optimization results and fault warning information of the GPU cluster, and provide corresponding operation options for users to intervene and adjust.

[0072] The user interaction interface adopts a graphical interface design, which can intuitively display the computing resource distribution, task execution progress and optimization effect of the GPU cluster, and provide convenient operation tools so that users can remotely monitor and manage the GPU cluster.

[0073] In specific applications of the embodiments of the present invention, the user interaction interface, through careful layout and color matching, enables users to grasp the computing power resource distribution, task execution progress, and optimization effects of the GPU cluster at a glance. In terms of displaying the computing power resource status, the user interaction interface can present key information such as the computing power utilization, temperature, and power consumption of each GPU in the cluster in real time, allowing users to fully understand the operating status of the cluster. At the same time, the interface also displays the distribution of computing power resources in a graphical manner, such as through heat maps or bar charts, to intuitively reflect the computing power differences between different GPUs.

[0074] In terms of displaying task execution status, the interface provides a detailed list of currently executing tasks, including information such as task name, execution time, required computing resources, and execution progress. Users can click on a task entry to further view task details or make necessary adjustments. Regarding the display of optimization results, the user interface intuitively presents indicators such as the comparison of computing power utilization before and after optimization and the reduction in task execution time, allowing users to clearly see the actual effects of the optimization. The interface also provides detailed explanations and suggestions for optimization strategies to help users better understand and apply optimization solutions.

[0075] In terms of fault warning information display, once a fault or abnormality occurs in the GPU cluster, the interface will immediately pop up an alert window, displaying information such as the fault type, scope of impact, and recommended countermeasures. Users can quickly initiate the fault handling process or make other necessary interventions by clicking on the operation options in the alert window. In addition, the user interface also provides convenient operation tools, such as shortcut buttons or menu options for tasks scheduling, resource allocation, troubleshooting, and other functions, allowing users to easily remotely monitor and manage the GPU cluster, greatly improving the cluster's operation and maintenance efficiency and flexibility.

[0076] When the present invention is used in practice, Example 1: Construction of basic computing power optimization architecture

[0077] Implementation of computing resource management module:

[0078] Computing power monitoring unit: This unit uses a high-performance sensor array to monitor key indicators such as computing power load, memory usage, operating temperature, and power consumption of each GPU in the GPU cluster in real time. The data is aggregated to the central management server through a high-speed network interface to achieve real-time data updates and analysis.

[0079] Computing Power Allocation Unit: Based on a multi-factor comprehensive evaluation algorithm, the algorithm dynamically adjusts GPU resource allocation, taking into account factors such as the GPU's current load, historical performance, life expectancy, and task urgency. For example, for GPUs with high loads or approaching temperature thresholds, the system automatically reduces the amount of tasks assigned to them to protect the hardware and extend its lifespan.

[0080] Implementation of task scheduling module:

[0081] Task Priority Queue: A dynamic priority queue is established based on factors such as task urgency, required computing resources, and estimated execution time. High-priority tasks are given priority access to GPU resources, ensuring that critical tasks are completed quickly.

[0082] Multi-task parallel scheduling: Leveraging the parallel computing capabilities of GPU clusters, we can schedule multiple tasks for parallel execution. Intelligent algorithms are used to predict resource conflicts between tasks, enabling us to schedule resources in advance and reduce task waiting times.

[0083] Implementation of computing power optimization module:

[0084] Computing power prediction unit: Based on machine learning models, it predicts the computing power demand trend of GPU clusters over a period of time. The prediction results are used to adjust resource allocation in advance to ensure a balance between computing power supply and demand.

[0085] Computing power scheduling unit: Dynamically adjusts GPU resource allocation strategies based on computing power forecasts. If a surge in computing power demand is predicted during a certain period, the system will reserve or increase GPU resources in advance to meet the upcoming high load demand.

[0086] Computing power adjustment unit: During task execution, it monitors the GPU's computing power load in real time and fine-tunes the GPU's computing power based on real-time data to ensure optimal resource utilization.

[0087] Implementation of result feedback module:

[0088] Performance monitoring and reporting: Regularly collect and analyze GPU cluster performance data, including key metrics such as computing power utilization, task execution time, and resource bottlenecks. Generate detailed performance reports for administrators to reference and optimize system configurations.

[0089] Fault warning and user interaction interface:

[0090] Fault Warning System: Integrates deep learning models to monitor the health of GPU clusters in real time, predict potential failures, and issue early warnings. It also provides troubleshooting advice and guides administrators to respond quickly.

[0091] User Interface: A graphical user interface is used to intuitively display the GPU cluster's computing resource status, task execution status, optimization results, and fault warning information. Convenient operation tools are provided to support remote monitoring and management.

[0092] Example 2: Optimizing the performance of the architecture in practical applications

[0093] To verify the effectiveness of this invention, a large cluster consisting of 100 GPUs was selected for testing. The test scenario involved large-scale deep learning model training tasks involving multiple fields such as image recognition and natural language processing.

[0094] Improved computing power utilization: By implementing this invention, the computing power utilization efficiency of GPU clusters has increased by an average of over 30%. Especially during high-load periods, the dynamic scheduling and optimization strategies for computing power resources effectively alleviate resource bottlenecks and improve overall training speed.

[0095] Cost savings: Due to the significant improvement in computing power utilization, the time required for training tasks is significantly shortened, thereby reducing GPU usage time and energy costs. According to preliminary estimates, this method can save approximately 20% of operating costs compared to traditional management methods.

[0096] Enhanced system stability: The introduction of a fault warning system effectively reduces GPU failure rates and improves system stability and reliability. During testing, no training tasks were interrupted due to GPU failures.

[0097] Optimized user experience: The user-friendly interface design and convenient operation tools greatly enhance the user experience. Administrators can intuitively monitor and manage GPU clusters and quickly respond to various abnormal situations.

[0098] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0099] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. A GPU cluster computing power optimization architecture for large model training, characterized by: include: Computing power resource management module, task scheduling module, computing power optimization module and result feedback module; The computing power resource management module is used to monitor and manage the computing power resource status of each GPU in the GPU cluster in real time. The task scheduling module is used to allocate GPU resources according to the requirements of the training task. The computing power optimization module is used to dynamically optimize the computing power of the GPU cluster during task execution. The result feedback module is used to collect and analyze the usage of the optimized computing power to adjust the subsequent task scheduling strategy. The computing resource management module also includes: The computing resource management module has a built-in efficient computing power monitoring unit. This unit uses high-precision sensors and advanced data acquisition technology to obtain real-time status information such as computing power load, memory usage, temperature, and power consumption of each GPU in the GPU cluster, ensuring data accuracy and real-time performance. The computing power allocation unit in the computing power resource management module adopts a computing power allocation strategy based on a comprehensive evaluation of multiple factors. This computing power allocation strategy not only considers the computing power load and memory usage of the GPU, but also combines status information such as GPU temperature and power consumption, as well as factors such as the urgency of the training task, the amount of computing power required, and the expected execution time. It allocates corresponding GPU resources to each task to maximize the utilization of GPU resources. The computing resource management module supports dynamic adjustment and load balancing of GPU resources. When the computing load of a GPU is too high or the temperature is abnormal, the computing resource allocation unit can automatically adjust the allocation of GPU resources and migrate some tasks to other idle or less loaded GPUs to achieve load balancing of the GPU cluster, avoiding waste of computing resources and performance degradation caused by GPU overheating. The computing resource management module also has the functions of reserving and elastically expanding GPU resources. It can reserve a certain number of GPU resources in advance according to the requirements of training tasks, and dynamically increase GPU resources when needed to meet the needs of high-performance computing scenarios such as large-scale model training. The computing resource management module adopts a data security and privacy protection mechanism based on blockchain technology to ensure the security and traceability of the computing resource status information and task scheduling information of each GPU in the GPU cluster, and prevent data leakage and malicious attacks.

2. The GPU cluster computing power optimization architecture for large model training according to claim 1 is characterized in that: The task scheduling module adopts a priority-based task scheduling algorithm. The task scheduling algorithm not only considers factors such as the urgency of the training task, the amount of computing resources required, and the estimated execution time, but also combines the computing resource status and load balancing of the GPU cluster to allocate corresponding GPU resources to each task and generate a task execution plan; The task scheduling module further includes: The task scheduling module supports multi-task parallel processing. It can schedule multiple training tasks to execute in parallel based on the computing resource status and load balancing of the GPU cluster, thereby improving the computing power utilization efficiency of the GPU cluster. The task scheduling module has the function of dynamically adjusting task priorities. It can dynamically adjust the priority of tasks according to the urgency and execution progress of training tasks to ensure that urgent tasks are handled first. The task scheduling module also has a task preemption mechanism. When a new urgent task appears, it can preempt the low-priority task being executed to ensure that the urgent task can be processed in time. The task scheduling module supports custom task scheduling strategies. Users can customize task scheduling strategies according to actual needs to meet the requirements of specific application scenarios.

3. The GPU cluster computing power optimization architecture for large model training according to claim 1 is characterized in that: The task scheduling module adopts a priority-based task scheduling algorithm to allocate corresponding GPU resources to each task according to factors such as the urgency of the training task, the amount of computing power resources required, and the expected execution time, and generates a task execution plan.

4. The GPU cluster computing power optimization architecture for large model training according to claim 1 is characterized in that: The computing power optimization module includes a computing power prediction unit, a computing power scheduling unit and a computing power adjustment unit; the computing power prediction unit is used to predict the computing power demand of the GPU cluster in the future period of time; the computing power scheduling unit dynamically adjusts the allocation of GPU resources according to the computing power demand and the current computing power status of the GPU cluster; the computing power adjustment unit is used to fine-tune the GPU computing power according to the real-time computing power load during task execution to improve the utilization efficiency of computing power.

5. The GPU cluster computing power optimization architecture for large model training according to claim 4 is characterized in that: The computing power prediction unit uses a machine learning algorithm to train historical computing power usage data to establish a computing power demand prediction model. The computing power demand prediction model can predict the computing power demand in the future based on the current training task information and GPU cluster status.

6. The GPU cluster computing power optimization architecture for large model training according to claim 1 is characterized in that: The result feedback module includes a data analysis unit and a strategy adjustment unit; the data analysis unit is used to collect and analyze the optimized computing power usage, including the computing power utilization rate of each task, the task execution time and the overall computing power load of the GPU cluster, etc. The strategy adjustment unit adjusts the subsequent computing power allocation strategy and task scheduling strategy based on the analysis results to improve the computing power utilization efficiency of the GPU cluster.

7. The GPU cluster computing power optimization architecture for large model training according to any one of claims 1 to 6, characterized in that: It also includes a fault warning module; the fault warning module is used to monitor the operating status of each GPU in the GPU cluster in real time. When a GPU failure or abnormality is detected, it will promptly issue a warning message and start the corresponding fault handling process.

8. The GPU cluster computing power optimization architecture for large model training according to claim 7 is characterized in that: The fault warning module uses a deep learning-based fault prediction algorithm to train historical fault data to establish a fault prediction model. The fault prediction model can predict possible future faults based on the current GPU operating status information.

9. The GPU cluster computing power optimization architecture for large model training according to any one of claims 1 to 8, characterized in that: It also includes a user interaction interface; the user interaction interface is used to display the computing resource status, task execution status, optimization results and fault warning information of the GPU cluster, and provide corresponding operation options for users to intervene and adjust.

10. The GPU cluster computing power optimization architecture for large model training according to claim 9, characterized in that: The user interaction interface adopts a graphical interface design, which can intuitively display the computing resource distribution, task execution progress and optimization effect of the GPU cluster, and at the same time provide convenient operation tools so that users can remotely monitor and manage the GPU cluster.

Citation Information

Cited By

  • Parallel test task scheduling method oriented to high-computing-power GPU (Graphics Processing Unit) chip

    CN120832279A

  • A parallel test task scheduling method for a large-computing-power GPU chip

    CN120832279B

  • Method for dynamically using computing power of reasoning server through model superposition technology

    CN121072759A

  • Fault processing method and device

    CN121187868A

  • GPU cluster, redundancy optimization method of GPU cluster, electronic equipment, storage medium and computer program product

    CN121326655A