Method and system for regulating and controlling energy consumption of GPU BOX, and computer device

By analyzing the task characteristics of GPU BOX and building a performance model, real-time monitoring and dynamic regulation of energy consumption are carried out, solving the problems of complex energy consumption management and resource waste in existing GPU servers, and achieving efficient energy consumption management and resource utilization.

CN120704889AActive Publication Date: 2025-09-26BEIJING RONGXIN ZHIYUAN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510853644.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-26
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing GPU server energy consumption management method is simple and static, and cannot be dynamically adjusted according to task requirements, resulting in high energy consumption, resource waste and complex management, affecting overall energy efficiency and stability.

Method used

By analyzing the characteristics of the tasks to be processed, identifying task characteristic information, building a GPU BOX performance model, monitoring and dynamically adjusting energy consumption in real time, and rationally allocating tasks to appropriate GPU BOXes, dynamic energy consumption regulation is achieved.

Benefits of technology

It improves GPU resource utilization, reduces operating costs, ensures stable system operation, avoids resource waste and task backlogs, and meets energy conservation and emission reduction requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704889A_ABST
    Figure CN120704889A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU BOX energy consumption regulation and control method and system and a computer device. The method comprises the following steps: performing feature analysis on a received to-be-processed task, and identifying task feature information; obtaining historical performance data of all GPU BOXs of the same node corresponding to the task type, constructing a GPU BOX performance model according to state information, obtained in real time, of all GPU BOXs, and obtaining a GPU BOX distribution strategy used for executing the task in the same node according to the model; determining a target GPU BOX according to the GPU BOX distribution strategy, and executing the to-be-processed task; and monitoring the operation information of the target GPU BOX in real time, obtaining the dynamic regulation and control strategy of the energy consumption of the GPU BOX according to the operation information, and executing the dynamic regulation and control strategy of the energy consumption of the GPU BOX. According to the method, dynamic energy consumption regulation and control of the GPU BOX in different working modes can be achieved, efficient utilization of energy consumption is achieved, the service life of equipment is prolonged, energy is saved, and cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method and system for controlling energy consumption of a GPU BOX, and a computer device. Background Art

[0002] In existing technologies, the demand for AI applications, primarily fine-tuning and inference, has exploded, significantly increasing the frequency of use and computing load of GPU servers. When running, the primary power consumption of GPU servers comes from the various GPU modules. During large-model pre-training and fine-tuning, if a GPU involved in a task encounters a problem, the entire task will fail and need to be restarted. This not only wastes a significant amount of computing resources but also increases unnecessary energy consumption. Under low-load conditions, GPUs maintain high energy consumption. Even if the server handles fewer tasks, they still consume a significant amount of electricity, severely impacting overall energy efficiency and keeping operational costs high. Furthermore, GPUs generate significant heat when operating under high load for extended periods, which can damage their internal electronic components and reduce their performance and stability. Frequent replacement of aging GPUs not only increases hardware costs but also impacts the normal operation of the server.

[0003] Currently, energy consumption management for GPU servers mostly adopts a simple cumulative approach, manually controlling the power on and off of GPU servers. This fails to dynamically activate and use the GPU on demand based on actual task requirements. Furthermore, the energy consumption control interfaces of different GPU manufacturers are not unified, making the management process complex and inefficient. Manual management is not only prone to operational errors, but also unable to adjust the server's energy consumption status in a timely manner according to task changes. Summary of the Invention

[0004] In view of this, the embodiments of the present disclosure provide a GPU BOX energy consumption control method and system, as well as a computer device, which can solve the problems existing in the prior art, such as high energy consumption of GPU BOX servers, low energy utilization efficiency, and inability to dynamically adjust the energy consumption of multiple GPU BOX servers under different task states.

[0005] The present disclosure provides a method for controlling energy consumption of a GPU BOX, including: Performing feature analysis on the received tasks to be processed to identify task characteristic information, wherein the task characteristic information includes task type; Get the historical performance data of all GPU boxes on the same node corresponding to the task type; Building a GPU BOX performance model based on the real-time acquired status information of all GPU BOXes and the historical performance data; Analyze the task characteristic information according to the GPU BOX performance model to obtain a GPU BOX allocation strategy for executing the task in the same node; Determine a target GPU BOX according to the GPU BOX allocation strategy, and execute the task to be processed based on the target GPU BOX; The operating information of the target GPU BOX is monitored in real time, a dynamic control strategy for energy consumption of the GPU BOX is acquired according to the operating information, and the dynamic control strategy for energy consumption of the GPU BOX is executed.

[0006] In a second aspect, the present application discloses a GPU BOX energy consumption control system, which is used to execute the GPU BOX energy consumption control method, including: The server is configured to receive tasks to be processed, perform feature analysis on the tasks to be processed, and identify task characteristic information, wherein the task characteristic information includes task type, computing requirements, memory usage, priority information, and estimated task execution time; A historical performance data acquisition module is used to obtain historical performance data of all GPU BOXes on the same node corresponding to the task type; A model building module is used to build a GPU BOX performance model based on the status information of all GPU BOXs obtained in real time and the historical performance data; An allocation strategy acquisition module, configured to analyze the task characteristic information according to the GPU BOX performance model and acquire a GPU BOX allocation strategy for executing the task in the same node; an execution module, configured to determine a target GPU BOX according to the GPU BOX allocation strategy, and execute the task to be processed based on the target GPU BOX; The dynamic control module is used to monitor the operating information of the target GPU BOX in real time, obtain the GPU BOX energy consumption dynamic control strategy according to the operating information, and execute the GPU BOX energy consumption dynamic control strategy.

[0007] In a third aspect, the embodiments of the present disclosure further provide a computer device that adopts the following technical solution: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute any of the above-mentioned methods for controlling energy consumption of the GPU BOX.

[0008] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute any of the above-mentioned methods for controlling energy consumption of a GPU BOX.

[0009] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.

[0010] The GPU Box energy consumption control method disclosed in this application can achieve real-time energy consumption control for each GPU. By analyzing task characteristic information and constructing a GPU Box performance model, it can rationally allocate tasks to appropriate GPU Boxes, fully leveraging the performance advantages of each GPU Box, avoiding resource waste and task backlogs, and thus improving overall task execution efficiency. The method monitors the operating information of the target GPU Box in real time and dynamically adjusts energy consumption based on the operating status, preventing the GPU Box from consuming excessive electricity unnecessarily, reducing the operating costs of the data center and meeting the requirements of energy conservation and emission reduction. Reasonable task allocation strategies and energy consumption control strategies can fully utilize all GPU Boxes in the same node, avoiding situations where some GPU Boxes are overloaded while others are idle, thereby improving resource utilization. Through real-time monitoring and dynamic control of GPU Boxes, problems that arise during GPU Box operation, such as overheating and overload, can be promptly discovered and resolved, ensuring stable system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flow chart of the energy consumption control method of the GPU BOX provided in an embodiment of the present disclosure.

[0013] Figure 2 This is a flowchart of a method for obtaining historical performance data of all GPU BOXes on the same node corresponding to a task type provided by an embodiment of the present disclosure.

[0014] Figure 3 A flowchart of a method for constructing a GPU BOX performance model provided in an embodiment of the present disclosure.

[0015] Figure 4A flowchart of a method for obtaining a GPU BOX allocation strategy provided by an embodiment of the present disclosure.

[0016] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0018] Reference Figure 1 In a first aspect, the present application discloses a method for controlling energy consumption of a GPU BOX, comprising: S100: Receive tasks to be processed.

[0019] Specifically, the server receives tasks to be processed through a network interface and can use common network protocols such as HTTP, TCP, etc.; for example, in a distributed computing system, the client sends a task request to a specific port of the server, and the server uses a corresponding network programming library (such as Python's socket library or Flask framework) to listen to the port. After receiving the task request, it parses it into the format of the task to be processed.

[0020] S200 , performing feature analysis on the task to be processed and identifying task characteristic information.

[0021] Among them, task characteristic information includes task type, computing requirements, memory usage, priority information and estimated task execution time.

[0022] The priority information includes high priority (such as real-time tasks) and low priority (such as batch tasks).

[0023] By identifying task characteristic information, we can gain a deeper understanding of the task's requirements and characteristics, providing a basis for subsequent selection of the appropriate GPU Box and formulation of energy consumption control strategies. Different types of tasks, with varying computing needs and memory usage, have different performance requirements for the GPU Box. Accurate feature analysis helps improve the efficiency and accuracy of task execution.

[0024] Furthermore, task type identification can be performed based on the task's file extension, task description, or execution command. Computational requirements analysis can estimate computational requirements based on factors such as the task's code structure and input data size. Memory usage estimation can be performed by analyzing factors such as the variables and data structures involved in the task's code. Priority information can be used to determine the priority of a task based on factors such as the task submitter and the task's urgency. Estimated runtime can be achieved by combining historical task runtime data with the current task's characteristic information using machine learning algorithms (such as linear regression).

[0025] S300: Obtain historical performance data of all GPU boxes on the same node corresponding to the task type.

[0026] As each GPU Box executes a task, its performance data, including computing speed, memory bandwidth, temperature, and power consumption, is recorded and stored in a database. Depending on the task type, historical performance data for all GPU Boxes on the same node can be queried from the database using SQL statements or a database query interface.

[0027] Historical performance data can reflect the actual performance of each GPU Box when processing specific types of tasks, providing a reference for building performance models. By analyzing historical data, we can understand the strengths and weaknesses of each GPU Box and make more reasonable decisions when allocating tasks.

[0028] S400 builds a GPU BOX performance model based on the real-time status information and historical performance data of all GPU BOXes.

[0029] Through the constructed model, we can understand the performance of each GPU BOX in advance, avoiding resource waste or inefficient task execution caused by blindly allocating tasks.

[0030] Specifically, intelligent energy consumption management software can be deployed in the server to synchronize the historical data and real-time status of the energy consumption management agents of each GPU server, store them in a time series database in chronological order, and synchronize them with the intelligent energy consumption management software.

[0031] S500: Analyze task characteristic information according to the GPU BOX performance model to obtain a GPU BOX allocation strategy for executing the task in the same node.

[0032] A reasonable task allocation strategy can fully utilize the performance advantages of each GPU BOX, improve the overall task execution efficiency, and avoid the situation where some GPU BOXes are overloaded while others are idle due to random task allocation, thereby improving resource utilization.

[0033] S600: Determine a target GPU BOX according to a GPU BOX allocation policy, and execute the task to be processed based on the target GPU BOX.

[0034] This step clarifies the specific device for task execution, ensuring that the task can be executed smoothly according to the allocation strategy; matching the task with the appropriate GPU BOX improves the accuracy and efficiency of task execution.

[0035] S700 monitors the operating information of the target GPU BOX in real time, obtains the GPU BOX energy consumption dynamic control strategy based on the operating information, and executes the GPU BOX energy consumption dynamic control strategy.

[0036] Through real-time monitoring and dynamic regulation, energy consumption can be flexibly adjusted according to the actual operation of the GPU BOX, avoiding the situation where the GPU BOX still consumes a lot of electricity under low load, thereby reducing overall energy consumption and improving energy utilization efficiency.

[0037] The GPU Box energy consumption control method disclosed in this application can achieve real-time energy consumption control for each GPU. By analyzing task characteristic information and constructing a GPU Box performance model, it can rationally allocate tasks to appropriate GPU Boxes, fully leveraging the performance advantages of each GPU Box, avoiding resource waste and task backlogs, and thus improving overall task execution efficiency. The method monitors the operating information of the target GPU Box in real time and dynamically adjusts energy consumption based on the operating status, preventing the GPU Box from consuming excessive electricity unnecessarily, reducing the operating costs of the data center and meeting the requirements of energy conservation and emission reduction. Reasonable task allocation strategies and energy consumption control strategies can fully utilize all GPU Boxes in the same node, avoiding situations where some GPU Boxes are overloaded while others are idle, thereby improving resource utilization. Through real-time monitoring and dynamic control of GPU Boxes, problems that arise during GPU Box operation, such as overheating and overload, can be promptly discovered and resolved, ensuring stable system operation.

[0038] Reference Figure 2 The method for obtaining the historical performance data of all GPU BOXes of the same node corresponding to the task type in S300 specifically includes: S310: Obtain performance data of each GPU BOX during the execution of historical tasks and store the data in a target database.

[0039] Among them, performance data includes computing speed corresponding to different tasks, memory bandwidth usage information, temperature information, power consumption information, load requirement information, and historical execution time.

[0040] For example, in a large data center, multiple GPU boxes are used for different tasks, such as deep learning training, video rendering, and scientific computing. For each GPU box, a corresponding monitoring system collects real-time performance data as it executes historical tasks. For example, when a GPU box executes a deep learning training task, the monitoring system records its computing speed (floating-point operations per second, e.g., 10 teraflops); memory bandwidth usage (e.g., the amount of data read and written from memory per second); temperature information (using a temperature sensor to obtain the real-time temperature of the GPU); power consumption information (using a power meter to measure the real-time power consumption of the GPU box); load requirements (e.g., the percentage of GPU cores occupied by the task); and historical execution duration (e.g., the time elapsed from task start to finish). This collected performance data is organized into a specific format and stored in a target database. The target database can be a relational database, storing data by fields such as GPU box number, task type, and timestamp to facilitate subsequent querying and management.

[0041] Recording performance data helps fully understand the actual performance of each GPU Box under different tasks. By analyzing this data, we can identify the strengths and weaknesses of each GPU Box. For example, a GPU Box may have high computing speed but high power consumption when processing a specific task type. This helps make more reasonable decisions when allocating tasks. Storing data in the target database facilitates long-term data preservation and management. The database's structured storage method allows for easy data query, statistics, and analysis, laying the foundation for subsequent queries of relevant performance data based on task types.

[0042] S320: According to the task type, query the target database for performance data of all GPUBOXs of the same node that have executed the task type.

[0043] This step allows targeted acquisition of performance data related to the current task type, providing an accurate reference for subsequent task allocation and energy consumption control. By comparing the performance of different GPU Boxes on the same node under this task type, the GPU Box most suitable for executing the current task can be selected, improving task execution efficiency, effectively avoiding interference from irrelevant data, reducing the workload of data processing, and querying only data related to the current task type, thereby improving the efficiency and accuracy of data acquisition.

[0044] Reference Figure 3The method of S400 "building a GPU BOX performance model based on real-time acquired status information and historical performance data of all GPU BOXes" is a method for building a GPU BOX performance model, specifically including: S410, obtains status information of all GPU BOXes in real time.

[0045] The status information includes the current load, current temperature, current power consumption, and current available memory of each GPU BOX.

[0046] Specifically, you can use GPU BOX management tools (such as NVIDIA SMI and AMD ROCm SMI) to obtain the status information of all GPU BOXs in real time.

[0047] Furthermore, in a data center with multiple GPU boxes, dedicated monitoring software and hardware sensors can be used to obtain real-time status information. For example, NVIDIA management tools (such as NVIDIA SMI) can be used to obtain the current GPU box load; temperature sensors can measure the current temperature; power consumption sensors can monitor current power consumption in real time; and system software can query the current available memory. Real-time status information reflects the current operating status of the GPU box, providing the latest and most accurate data foundation for subsequent availability assessment and performance model construction. Only based on real-time data can decisions be made that are consistent with current conditions, avoiding improper task allocation due to the use of outdated information.

[0048] S420 , respectively determining availability evaluation results of the current load, current temperature, current power consumption, and current available memory.

[0049] By evaluating each key indicator individually, we can clearly understand the availability of the GPU BOX in all aspects. This helps identify potential problems with the GPU BOX, such as excessive load or high temperature, and provides a detailed basis for subsequent decision-making, avoiding the impact of normal task execution due to a certain indicator exceeding the reasonable range.

[0050] S430: Determine a preset load weight, a preset temperature weight, a preset power consumption weight, and a preset memory weight based on historical performance data of all GPU BOXes.

[0051] Determining weights based on historical performance data allows for a more scientific consideration of the impact of various indicators when comprehensively evaluating GPU Box availability. Different tasks may have different sensitivities to different indicators. Reasonable weight settings can more accurately reflect the actual performance of the GPU Box and provide a more reasonable basis for task allocation.

[0052] Specifically, the performance of all GPU BOXes in historical tasks is analyzed. For example, in a large number of previous tasks, it is found that load has a greater impact on task execution speed, while temperature has a relatively smaller impact on task execution speed. Through statistical analysis and empirical judgment, for example, the preset load weight can be determined to be 0.5, the preset temperature weight to be 0.1, the preset power consumption weight to be 0.2, and the preset memory weight to be 0.2. These weights represent the importance of each indicator in evaluating GPU BOX performance.

[0053] S440: When all availability evaluation results are passed, a weighted sum of the current load, current temperature, current power consumption, and current available memory is performed according to a preset load weight, a preset temperature weight, a preset power consumption weight, and a preset memory weight to obtain an availability score for the single GPU BOX.

[0054] The availability score for a single GPU Box is calculated as follows: preset load weight × current load + preset temperature weight × current temperature + preset power consumption weight × current power consumption + preset memory weight × current available memory. This weighted summation of the availability score comprehensively considers the impact of various indicators, converting multiple GPU Box performance metrics into a single value, making it easier to compare and rank different GPU Boxes. When assigning tasks, GPU Boxes with higher availability scores can be prioritized based on the availability score, improving task execution efficiency.

[0055] S450 builds a GPUBOX performance model based on the availability scores of all GPU BOXes and the corresponding historical execution times of all GPU BOXes.

[0056] Specifically, the availability scores of all GPU boxes are sorted in descending order. The sorted list is the GPU box performance model. The higher the score, the higher it is in the list, which means its performance is relatively better.

[0057] Alternatively, the number of clusters is determined; the availability scores of all GPU BOXs are clustered according to preset rules to build a cluster-based GPU BOX performance model; wherein determining the number of clusters includes classifying the GPU BOXes into three categories: high performance, medium performance, and low performance).

[0058] Alternatively, a performance model can be constructed using a machine learning algorithm (such as linear regression) with the availability score as the independent variable and the historical execution time as the dependent variable. Suppose that by training data from multiple GPU boxes, a linear regression equation is obtained: historical execution time = a × availability score + b, where a and b are model parameters. This model can be used to predict the approximate execution time of a GPU box based on its availability score. This constructed performance model can quantitatively predict the performance of a GPU box, providing a more scientific basis for task allocation and scheduling. During task allocation, the GPU box with the shortest execution time can be selected based on the model's predictions, improving overall task execution efficiency. Furthermore, this model can be used to dynamically adjust task execution during execution, optimizing task allocation strategies based on actual conditions.

[0059] The method for determining the availability evaluation results of the current load, the current temperature, the current power consumption, and the current available memory in S420 includes: A100, determining whether the current load is lower than a preset load threshold; if so, determining that the availability evaluation result of the current load is passed; if not, determining that the availability evaluation result of the current load is failed; A200 determines whether the current temperature is lower than a preset temperature threshold. If so, determines that the availability evaluation result of the current temperature passes. If not, determines that the availability evaluation result of the current temperature fails. A300 determines whether the current power consumption is lower than a preset power consumption threshold. If so, determines that the availability evaluation result of the current power consumption passes; if not, determines that the availability evaluation result of the current power consumption fails. A400 determines whether the current available memory is higher than a preset memory threshold. If so, determines that the availability evaluation result of the current available memory is passed. If not, determines that the availability evaluation result of the current available memory is failed.

[0060] By evaluating the availability of four key metrics—current load, temperature, power consumption, and available memory—we can promptly identify potential GPU Box issues and avoid assigning tasks to GPU Boxes in poor condition, thereby ensuring stable operation of the entire system and reducing task failures due to hardware failures or insufficient resources. By selecting GPU Boxes in good condition and with sufficient resources for task execution, we can fully utilize the performance and resources of each GPU Box, avoid resource waste, and improve the overall resource utilization efficiency of the system. While ensuring system performance, we prioritize GPU Boxes with lower power consumption for task execution, helping to reduce energy consumption and cooling costs, thereby lowering data center operating costs. Based on the availability evaluation results of each GPU Box, we can develop a more scientific and reasonable task allocation strategy, accurately assigning tasks to the most suitable GPU Box, and improving task execution efficiency and quality.

[0061] Reference Figure 4 Regarding the method of S500 "analyzing task characteristic information according to the GPU BOX performance model to obtain a GPU BOX allocation strategy for executing tasks in the same node", the method for obtaining the GPU BOX allocation strategy includes: S510 , based on the GPU BOX performance model, obtain a GPU BOX whose historical execution time is not longer than the estimated task execution time and whose availability score is not lower than a preset score threshold, and record it as a qualified GPU BOX.

[0062] This step can quickly screen out GPU boxes that meet basic requirements, narrow the search scope of subsequent allocation strategies, avoid allocating tasks to GPU boxes with long historical execution times or low availability scores, and improve the quality and efficiency of task allocation.

[0063] S520 encodes the task allocation status of each qualified GPU BOX into a chromosome.

[0064] Suppose we have three qualified GPU boxes (A, B, and C). Each GPU box has two task allocation states: assigned (indicated by 1) and unassigned (indicated by 0). Possible chromosome encodings are as follows: [1, 0, 1] indicates that tasks are assigned to GPU boxes A and C, but not to B; [0, 1, 0] indicates that tasks are assigned to GPU boxes B, but not to A or C. Encoding task allocation as chromosomes provides a suitable data structure for subsequent optimization using genetic algorithms. Genetic algorithms can easily operate on chromosomes, simulating the process of biological evolution to find the optimal solution.

[0065] S530: Determine a fitness function for evaluating the quality of each chromosome, where the fitness function is the inverse of the estimated task execution time.

[0066] For a task allocation scheme corresponding to a chromosome, if the GPU BOX performance model estimates that the task execution time under this scheme is 8 hours, then the fitness value of this chromosome is 1 / 8 = 0.125. The fitness function provides the genetic algorithm with a criterion for evaluating the quality of each task allocation scheme. The inverse of the estimated task execution time is used as the fitness function, which gives schemes with shorter execution times higher fitness. This guides the genetic algorithm to evolve towards the shortest execution time, thereby finding the optimal task allocation strategy.

[0067] S540 , based on the value of the fitness function, select a chromosome whose fitness is greater than a preset fitness threshold to perform crossover and mutation operations to generate a new chromosome.

[0068] Assume the preset fitness threshold is 0.1. Currently, the fitness of chromosome [1, 0, 1] is 0.125, and the fitness of chromosome [0, 1, 0] is 0.08. The chromosome [1, 0, 1] with a fitness greater than 0.1 is selected for crossover and mutation operations. A crossover operation involves exchanging some genes in [1, 0, 1] with another chromosome, for example, by crossing it with [0, 1, 0] to obtain a new chromosome [1, 1, 1]. A mutation operation involves randomly changing a gene in a chromosome, for example, changing the second gene in [1, 0, 1] from 0 to 1 to obtain [1, 1, 1]. Crossover and mutation operations simulate the inheritance and mutation processes in biological evolution. By operating on chromosomes with higher fitness to generate new chromosomes, the crossover increases the diversity of the search space and helps find a more optimal task allocation solution.

[0069] S550, based on the new chromosome, repeatedly perform chromosome selection, crossover and mutation operations until the optimal chromosome is found when the maximum number of iterations is reached or the fitness value converges.

[0070] The difference between the optimal fitness values ​​of two adjacent generations is calculated. If the difference is less than a preset threshold, it is considered that the fitness value has converged.

[0071] Through multiple iterations and evolutions, the task allocation plan is continuously optimized to increase the probability of finding the optimal solution. The setting of the maximum number of iterations and the fitness value convergence conditions can prevent the algorithm from falling into an infinite loop and ensure that the algorithm ends within a reasonable time.

[0072] S560: Obtain a GPU BOX allocation strategy for executing tasks in the same node according to the optimal chromosome.

[0073] In the previous step, we encoded the task allocation of each qualified GPU Box into a chromosome. Specifically, a chromosome is a binary list, where each element corresponds to a qualified GPU Box. A value of 0 indicates that the GPU Box does not participate in task execution, while a value of 1 indicates that the GPU Box participates in task execution. When the optimal chromosome is found through the genetic algorithm, this chromosome represents the optimal GPU Box allocation solution. We simply traverse the optimal chromosome and select the qualified GPU Boxes corresponding to the elements with a value of 1. These GPU Boxes are the optimal GPU Boxes we need.

[0074] Assume the optimal chromosome is [1, 0, 1], indicating that the task is assigned to GPU Boxes A and C, but not to B. This is the GPU Box allocation strategy for executing tasks on the same node. After optimization using the genetic algorithm, the task allocation strategy corresponding to the optimal chromosome is optimized to minimize task execution time while meeting basic requirements. This effectively improves task execution efficiency and achieves optimal resource allocation.

[0075] In this embodiment, the task allocation scheme is optimized through a genetic algorithm to find the GPU BOX allocation strategy that minimizes the task execution time, which can significantly improve the task execution efficiency and reduce the overall task execution time. The optimal allocation scheme is found among many qualified GPU BOXes, making full use of the resources of each GPU BOX, avoiding resource waste, and improving resource utilization.

[0076] The method of S700 for "real-time monitoring of target GPU BOX operation information, obtaining a GPU BOX energy consumption dynamic control strategy based on the operation information, and executing the GPU BOX energy consumption dynamic control strategy" includes: S710 monitors the target GPU BOX's operating information in real time.

[0077] The operation information includes the actual computing load, actual temperature information, and actual power consumption information when executing the task to be processed.

[0078] Real-time acquisition of GPU Box operating information is the basis for subsequent energy consumption control. The computing load reflects the current workload of the GPU Box, while the temperature and power consumption are closely related to the performance and energy consumption of the GPU Box. By monitoring this information in real time, we can promptly understand the status of the GPU Box and provide accurate data support for subsequent energy consumption control decisions.

[0079] S720 , obtain the GPU BOX corresponding to the actual computing load below the load lower limit threshold, record it as the GPU BOX to be adjusted, and obtain the actual low-load operation time of each GPU BOX to be adjusted that continuously runs at a load below the load lower limit threshold.

[0080] The lower load threshold is 15% of the rated load of the corresponding GPU BOX.

[0081] Screening out GPU boxes running at low loads can help us focus on devices that may be wasting energy. The actual duration of low-load operation is an important basis for determining whether energy consumption control is necessary and what control strategy to adopt. If a GPU box is in a low-load operation state for a long time, it is necessary to control its energy consumption to reduce unnecessary energy consumption.

[0082] S730: Determine a corresponding GPU BOX energy consumption dynamic control strategy based on the actual low-load operation time, and perform GPU BOX energy consumption control based on the GPU BOX energy consumption dynamic control strategy.

[0083] Formulating different energy consumption control strategies based on the actual low-load operation time can achieve precise energy consumption management. Different low-load operation times reflect the different usage states of the GPU BOX. Targeted control strategies can minimize energy consumption and improve energy utilization efficiency while ensuring the normal operation of the GPU BOX.

[0084] Furthermore, the method of S730 of "determining a corresponding GPU BOX energy consumption dynamic control strategy based on the actual low-load operation time, and performing GPU BOX energy consumption control according to the GPU BOX energy consumption dynamic control strategy" specifically includes: S731, determine whether the total number of GPU BOXes to be adjusted is greater than N / 2. If so, determine whether the average load of all target GPU BOXes is lower than the load lower limit threshold. If so, randomly select The target GPU BOX is put into sleep mode. If not, a random The GPU BOX to be adjusted is put into hibernation.

[0085] in, ; , N is the total number of all target GPU BOXes, The total number of all GPU BOXes to be adjusted.

[0086] When there are a large number of GPU boxes to be adjusted, the average load of all target GPU boxes can be determined to provide a more comprehensive understanding of the operating status of the entire system. If the average load is also low, it means that the entire system is in a low-load state. Randomly selecting target GPU boxes for hibernation can quickly reduce system energy consumption. If the average load is not low, only randomly selecting the GPU boxes to be adjusted for hibernation can accurately target low-load devices, avoid affecting the normal operation of high-load devices, and reduce energy consumption while ensuring system performance.

[0087] S732: If the total number of GPU BOXes to be adjusted is not greater than N / 2, adjust the power consumption of all GPU BOXes to be adjusted to 5%-13% of the rated load of the corresponding GPU BOXes.

[0088] Furthermore, it is preferred that the power consumption of all GPU BOXs to be adjusted is adjusted to 10% of the rated load of the corresponding GPU BOX.

[0089] When there are a small number of GPU boxes to be adjusted, adjusting their power consumption to a lower level can effectively reduce the energy consumption of these low-load devices without affecting the overall system performance. Selecting a range of 5%-13% can ensure that the GPU boxes maintain basic operating status so that they can quickly resume high-load operation when needed, while also achieving significant energy consumption reduction.

[0090] In this embodiment, different energy consumption control strategies are adopted based on the number of GPU BOXes to be adjusted and the average system load. This allows for precise energy consumption optimization for different system states, avoiding the performance loss or energy waste that may result from a one-size-fits-all control approach. By hibernating low-load GPU BOXes or reducing their power consumption, unnecessary energy consumption is reduced, improving the energy efficiency of the entire data center and reducing operating costs. The overall operating state and performance requirements of the system are fully considered during the energy consumption control process. Randomly selecting hibernating devices or adjusting power consumption does not significantly impact the normal operation of the system, ensuring system stability and reliability. This solution can dynamically adjust energy consumption strategies based on real-time monitored GPU BOX operating information, adapting to changes in system load and ensuring that the system is always operating efficiently and energy-efficiently.

[0091] Furthermore, when all target GPU BOXes in the same node are divided into several groups, the target GPU BOXes in each group are recorded as subgroups, and N is the total number of all target GPU BOXes in each subgroup. The total number of all GPU BOXes to be adjusted in each subgroup; The operating mode in which the total number of GPU BOXes to be adjusted is greater than N / 2 and the average load of all target GPU BOXes is lower than the load lower limit threshold is recorded as the target mode; Determine all subgroups with the target pattern and record them as target groups; When there are at least two target groups, the load is migrated to one target group for processing, and all GPU boxes in the target group without tasks are put to sleep.

[0092] Group management allows for more refined GPU Box management. Different subgroups may have different operating characteristics and load conditions. Grouping allows for more precise energy control and task allocation tailored to each subgroup's actual conditions, improving management efficiency. Defining target modes can help quickly identify subgroups with significant potential for energy optimization. When the majority of GPU Boxes within a subgroup are underloaded, this indicates potential energy waste within that subgroup, and defining this as a target mode facilitates subsequent centralized processing. Clearly defining target groups allows for focused energy control on subgroups, avoiding unnecessary operations across all subgroups and improving the relevance and efficiency of energy control. Through load migration and hibernation operations, resources can be further consolidated, reducing the number of GPU Boxes under load and lowering overall system energy consumption. By concentrating tasks within a subgroup, resource utilization within that subgroup can be improved, boosting overall system performance.

[0093] This embodiment centralizes tasks in low-load subgroups and hibernates idle GPU Boxes, reducing unnecessary energy consumption and significantly lowering overall data center energy consumption, thus saving operating costs. Load migration concentrates tasks within certain subgroups, improving resource utilization within those subgroups and avoiding resource fragmentation and waste, thereby enhancing overall system efficiency. Grouping GPU Boxes on the same node and centralizing processing for targeted groups simplifies energy consumption control and task management, reducing management complexity. This solution dynamically adjusts task allocation and energy consumption strategies based on the actual operating status of subgroups, enabling the system to better adapt to varying loads and enhancing system flexibility and adaptability.

[0094] In this application, the operating information of the target GPU BOX is monitored in real time, a dynamic GPU energy consumption control strategy is obtained based on the operating information, and the dynamic GPU energy consumption control strategy is executed, which also includes: B100, obtains the historical computing power demand of the same node corresponding to the task type during the non-working period and the computing power demand of the same node during the peak period.

[0095] Here, the same node refers to a single server.

[0096] Suppose we have a server node used for machine learning training tasks. Using the server's logging system, we can obtain the historical computing power demand of this node during non-operating hours (e.g., 12:00 PM to 6:00 AM daily) over the past month. Statistics show that the average computing power demand during non-operating hours is 1,000 GFLOPS (floating-point operations per second). Similarly, we obtain the computing power demand of this node during peak hours (e.g., 10:00 AM to 4:00 PM daily) from the logs and find that the average peak computing power demand is 5,000 GFLOPS. Understanding computing power demand during non-operating and peak hours provides foundational data for developing energy consumption control strategies. Different task types have significantly different computing power requirements during different time periods. Accurately obtaining this data allows us to more precisely adjust the GPU Box's energy consumption based on actual demand, avoiding energy waste during low-demand periods while ensuring sufficient computing power during high-demand periods.

[0097] B200 determines the level of historical computing power demand during non-working hours.

[0098] when When the historical computing power demand level during the non-working period is determined to be the first level; when When the historical computing power demand level during the non-working period is determined to be the second level; The historical computing power demand of the same node during non-working period. It is the computing power demand of the same node during peak period.

[0099] Dividing computing power requirements into levels makes the formulation of energy consumption control strategies more intuitive and simple; different levels correspond to different energy consumption control strategies, so that the appropriate control method can be quickly determined based on the approximate range of computing power requirements, thereby improving decision-making efficiency.

[0100] B300 determines the energy consumption control strategy of all GPU BOXes in the same node according to their levels, and dynamically controls the energy consumption of all GPU BOXes in the same node according to the energy consumption control strategy.

[0101] Formulating energy consumption control strategies based on computing power demand levels enables precise energy management. Reducing the power consumption of some GPU Boxes during low-demand periods not only meets basic computing power requirements but also significantly reduces energy consumption and saves costs. Maintaining the normal operation of all GPU Boxes during high-demand periods ensures the system can provide sufficient computing power to process tasks and maintain the normal operation of business.

[0102] The methods disclosed in B100-B300 analyze and categorize computing power requirements in different time periods and adjust the energy consumption of GPU boxes accordingly. This can avoid energy waste during non-working or low-demand periods, effectively reducing the overall energy consumption of the server and lowering operating costs. Dynamically adjusting the working status of the GPU boxes based on actual computing power requirements allows for more rational allocation and utilization of resources, fully utilizing the computing power of all GPU boxes during peak periods and reasonably reducing the power consumption of some GPU boxes during non-working periods, thereby improving resource utilization efficiency. While meeting computing power requirements in different time periods, it can ensure that the server has sufficient computing power to process tasks during peak periods, preventing the normal operation of the business from being affected by energy consumption regulation, thus ensuring business continuity and stability. This solution can dynamically adjust based on the historical computing power requirements of different task types, allowing the server system to better adapt to different business scenarios and working time periods, improving the system's flexibility and adaptability.

[0103] The B300 method of "determining the energy consumption control strategy for all GPU Boxes in the same node based on their levels, and dynamically controlling the energy consumption of all GPU Boxes in the same node based on the energy consumption control strategy" specifically includes: When the level is the first level, the energy consumption control strategy for all GPU BOXes in the same node is to disconnect the power supply of at least 1 / 3 of the GPU BOXes in the same node; When the level is the second level, the energy consumption control strategy for all GPU BOXes in the same node is determined to include disconnecting the power supply of at least two-thirds of the GPU BOXes in the same node, or randomly selecting two GPU BOXes to retain power supply; According to the energy consumption control strategy of all GPU BOXes in the same node, the power supply switch of the corresponding GPU BOX is controlled through the I²C chip.

[0104] At the first level of computing power demand, the current computing power demand is relatively low. Disconnecting the power of at least 1 / 3 of the GPU boxes can directly reduce the number of GPU boxes in operation, thereby significantly reducing the overall energy consumption of the node. At the same time, keeping some GPU boxes running can also ensure that the system maintains a certain level of computing power to cope with a small number of tasks that may arise.

[0105] The second level of computing power requires even lower computing power. Further reducing the number of running GPU Boxes, disconnecting at least two-thirds of the GPU Boxes or leaving only two powered, can lower energy consumption to a lower level, achieving even greater energy conservation. Both approaches offer flexibility, allowing users to choose the most appropriate strategy based on specific circumstances (such as performance differences between different GPU Boxes and anticipated subsequent tasks). The I²C chip is a common chip used for inter-device communication. In a server node, the power supply switch for each GPU Box is connected to the I²C bus. Once a power control strategy is determined, the server's control system can send control signals to the corresponding power supply switch via the I²C protocol. The I²C chip enables precise control of each GPU Box's power supply switch, accurately disconnecting or connecting the power supply to a specific GPU Box based on the energy control strategy, ensuring accurate energy control.

[0106] The solution disclosed in this embodiment precisely adjusts the operating status of the GPU BOX according to different computing power demand levels, minimizing energy consumption and saving electricity costs in different low-computing power demand scenarios. Different energy consumption control strategies are provided for different levels of computing power demand, and two optional methods are provided at the second level, allowing the system to flexibly adjust according to actual conditions and better adapt to various changes in computing power demand. Automated control of the power supply switch is achieved through the I²C chip, reducing manual intervention, improving the efficiency and accuracy of energy consumption control, and also reducing the risk of errors caused by human operation.

[0107] Furthermore, the energy consumption control method of the GPU BOX disclosed in the present application further includes: when several target GPU BOXes executing a task continuously run at no less than 75% of the rated load of the corresponding GPU for more than a preset high-load running time threshold, starting other target GPU BOXes to execute the task.

[0108] For example, the three selected target GPU BOXes execute the assigned tasks. When the three target GPU BOXes continue to run at a high load (75%) for more than a preset time (such as 30 seconds), the other two dormant / low-power GPUs are started through the VLLM inference framework to join in executing the unprocessed tasks.

[0109] When some GPU Boxes operate under high load for extended periods, their performance may be limited, or even lead to overheating and other issues, impacting the overall system's task processing speed. Enabling additional GPU Boxes to participate in task execution can help offload the workload from the high-loaded GPU Box, avoiding system performance bottlenecks and ensuring efficient and stable task completion. Increasing the number of running GPU Boxes means the system can handle more computing tasks simultaneously, improving parallel computing capabilities. Tasks previously handled by a few GPU Boxes can now be shared by more, significantly shortening task processing time and improving overall work efficiency.

[0110] Prolonged high-load operation accelerates GPU Box hardware aging and increases the risk of hardware failure. By activating additional GPU Boxes to offload tasks, the duration of sustained high-load operation on a single GPU Box can be reduced, reducing hardware wear and tear, extending the GPU Box's lifespan, and reducing hardware maintenance and replacement costs. High-load operation generates significant heat, which, if not dissipated promptly, can cause overheating and damage. By activating more GPU Boxes to offload tasks, the load on each GPU Box is reduced, and heat generation is correspondingly reduced, helping to maintain the GPU Box's operating temperature within a safe range and protecting the hardware. In some cases, some GPU Boxes may be in a dormant or low-power state, underutilizing their resources. When existing GPU Boxes are operating at high load, activating these idle GPU Boxes can fully utilize the system's computing resources and improve overall system resource utilization.

[0111] This strategy dynamically adjusts the number of GPU boxes involved in task execution based on the actual task load. When the task load increases, more GPU boxes are automatically activated to cope with the situation. When the task load decreases, some GPU boxes can be adjusted to low power or sleep mode, achieving dynamic resource optimization and making the system more flexible to adapt to different task requirements.

[0112] The following GPU server built with hot-swappable technology includes several GPU BOXes, and a group of five GPU BOXes is used as an example for explanation.

[0113] Case 1 (Basically Idle): When the average load of a group of GPU Boxes is below 15% and persists for a time exceeding a threshold, the energy management agent is invoked to control multiple GPU Boxes in that group to hibernate. For example, three or four GPU Boxes can be hibernated, saving at least 3*U2² / R watts of energy in real time. Case 2 (Inference Mode with Few Concurrent Sessions): When two GPU Boxes in a group have high loads and the loads of the other three remain below 15% for a time exceeding a threshold, the energy management agent is invoked to control two GPU Boxes in that group to hibernate, saving 2*U2² / R watts of energy in real time. Case 3 (Multiple GPU Groups Basically Idle): When Case 1 or 2 occurs across multiple GPU Boxes on the same server, the load is migrated to a single GPU Box, and all GPUs in the group with no tasks are hibernated.

[0114] Case 4 (mode with few inference tasks): For inference, normal usage will fall within the range of cases 1 and 2, requiring automatic energy consumption control. Non-working time periods (such as commuting hours, weekends, and holidays) can be set. When the energy management agent detects that the GPU server is idle in real time, the aforementioned strategy can be used to control the server's GPUs to only power on one or two GPUs within two to three threshold time periods, significantly saving energy. Case 5 (fine-tuning and pre-training mode): If the server is used for fine-tuning or pre-training, it will generally be in high-occupancy mode. The idle time threshold can be set relatively long, and even after tasks are completed, the idle time will also be relatively long. Apply the pre-processing strategy corresponding to case 4, modifying parameters such as the time threshold and the number of active GPUs in the group according to actual needs. Case 6 (low-power inference mode): For scenarios with high response requirements, during GPU box wakeup and task overload, if many GPUs are in sleep mode, there will be slight response delays or lags. In this case, a larger number of active GPUs can be set. In this case, the idle GPUs that are not in sleep mode can be adjusted to low-power mode. Case 7 (task increase mode): When the working GPU load continues to be high and reaches the high load time threshold, the low power consumption of the same group will be adjusted to normal power consumption to meet the computing power demand. If it still cannot meet the demand, the dormant GPUs will be gradually enabled to work normally.

[0115] Furthermore, this application also includes setting consumption values ​​for actions such as sleep scheduling, energy-saving scheduling, and wake-up scheduling. The higher the value, the higher the accuracy required for the corresponding action to achieve effective energy savings. For example, out-of-group sleep scheduling involving cross-switch chip task migration requires higher accuracy than in-group sleep scheduling, which requires higher accuracy than low-power scheduling within a group. However, the higher the value, the better the energy savings. Preprocessing is performed based on the task patterns and energy consumption records of existing control policies. This preprocessing collects statistics based on time periods, task volume patterns within time periods, real-time energy consumption, overall time period energy consumption, and scheduling accuracy. The statistical results are generated and reported, and saved. An intelligent algorithm regularly analyzes the statistical results and analyzes historical task characteristics for long periods of peak usage, such as task models, initiating IP addresses, or user names, to gradually reduce energy-saving scheduling that consistently consumes peak hours. The algorithm also analyzes task time period characteristics, analyzes the scheduling success rate and energy-saving effect of tasks with these time period characteristics, implements load forecasting, and forms an optimized timeline-based dynamic control strategy. For stable computing environments with relatively regular tasks, repeated recursion of the above method can form a relatively practical dynamic control strategy that can significantly exceed the energy savings achieved through manual intervention.

[0116] In this application, the GPU BOX includes an internal PCIe interconnect module, a GPU BOX module, a control module, a storage module, and a network module, all of which are connected to the internal PCIe interconnect module via the PCIe protocol. The GPU BOX adopts a modular structure to facilitate the rapid replacement and installation of GPUs, ensuring physical stability and safety during the insertion and removal process.

[0117] Specifically, the GPU Box includes a housing with a slot for mounting a GPU module. A control motherboard (i.e., a circuit board) and a network module are mounted within the housing. The control motherboard integrates a control module (preferably an integrated system-on-chip), a storage module (preferably an NVMe SSD hard drive), and a PCIe interconnect module (preferably a PCIe switch chip). The GPU module, control module, storage module, and network module are all connected to the PCIe interconnect module via the PCIe Class 1 protocol. A P2P DMA channel is established between the GPU module and the storage module, enabling direct access between the storage device and the GPU module's video memory, supporting direct read and write access to the disk. Compared to traditional data access methods, in which data is first transferred from disk to system memory and then processed and scheduled by the CPU for transfer to GPU video memory, the GPU Box disclosed in this embodiment provides a direct access method that significantly reduces data transmission latency and improves data processing efficiency. This solution bypasses the CPU, enabling direct communication between the GPU and storage device, significantly improving data transmission efficiency and speed.

[0118] The control motherboard has a PCIe slot for installing the GPU module, a hard drive interface for installing the storage module, and a backplane connection interface. The backplane connection interface is installed with a BP connector, which connects to the network module. The BP connector is used to output the network module's network signal to the OSFP interface on the server CPU control board. The PCIe slot can adapt to the PCIe interface of different models of GPU cards.

[0119] In this application, "all GPU Boxes in the same node" refers to all GPU Boxes belonging to the same CPU server, each of which contains a separate GPU card. Each GPU Box is independently configured, acting like a standardized component. During server or computing system deployment, it simply needs to be installed into the corresponding interface, eliminating the need for complex wiring and debugging. This plug-and-play feature significantly shortens system deployment time and improves efficiency. When computing needs increase, independent GPU Boxes can be easily added to boost the system's computing power. Whether increasing the number of GPUs within a single server or expanding across a cluster of multiple servers, this can be easily achieved without requiring major modifications to the existing system. Each GPU Box is independently configured, so if one fails, it will not affect the normal operation of the others. The system can quickly identify and isolate the failed unit, allowing computing tasks to continue using the remaining functioning GPUs, thereby improving the reliability and availability of the entire system and reducing downtime caused by hardware failures. Independent GPU Boxes offer greater operability during maintenance. If a unit experiences a problem, it can be removed from the system for repair or replacement without affecting the normal operation of other units or the entire system. This makes maintenance simpler and more efficient, reducing maintenance costs and difficulty. Because each GPU is independent, its performance can be monitored and managed individually. The monitoring software provides real-time information on each unit's operating status, performance indicators, and other information, allowing for timely identification of potential problems and implementation of adjustments and optimizations, ensuring the entire system is always operating optimally.

[0120] For different GPU BOXs, each GPU BOX on the same node is connected to the external PCIe interconnect module through the second-class PCIe protocol; the version of the second-class PCIe protocol is lower than the version of the first-class PCIe protocol; a GPUDirect P2P connection channel is established between the GPUs in each GPU BOX on the same node, so that the GPUs can directly use the PCIe bus for P2P DMA communication. Specifically, the GPU module is integrated with a DMA controller connected to the GPU card to control the data reading and writing of the GPU card and control the data interaction with the PCIe Switch chip. The storage module is integrated with a DMA controller connected to the storage chip to control the data reading and writing of the storage chip and control the data interaction with the PCIe Switch chip. The channel that supports P2P DMA is the transmission channel between the DMA controller inside the GPU module and the DMA controller inside the storage module.

[0121] Each GPU Box on a different node (i.e., on a different CPU server) connects to an RDMA switch module via the PCIe Class 3 protocol. The version of the PCIe Class 3 protocol is consistent with the version of the PCIe Class 1 protocol. GPUDirect RDMA (GDR) connections are established between GPU modules in every GPU Box on different nodes, enabling cross-node GPU interconnection via the RDMA switch module. The CPU (i.e., server) and GPU Box are not co-located on the same board, but are connected to each other via the PCIe Class 4 protocol, which is no later than the version of the PCIe Class 2 protocol. Specifically, the CPU manages the GPU modules, network modules, and storage modules via the PCIe bus. This non-co-located design eliminates reliance on components such as PCIe interface switches and retimers.

[0122] A hot-swappable connection channel is established between the GPU BOX and the CPU. The GPU BOX is integrated into the GPU BOX and connected to the server. A hot-swappable connection channel is established between the two. This channel is a dynamically pluggable interface link based on a specific high-speed data transmission protocol. This channel provides high-speed, low-latency data transmission capabilities, enabling real-time data exchange and command transmission between the GPU BOX and the CPU. Furthermore, when the server is operating normally, the GPU BOX can be plugged in and out at any time. During hot-swappable operations, there is no electrical damage to the server system or the GPU BOX, ensuring system stability and data integrity. This channel utilizes a reliable interface design to ensure stable data transmission and prevent signal interference and electrical problems that may occur during the plug-in and plug-out process.

[0123] In this application, from the technical control layer, the hot-swap function of the GPU BOX is realized through structural design and hardware circuits, supporting dynamic adjustment of power management to improve the flexibility and energy efficiency of the system and ensure that the GPU can be replaced or managed without shutting down the server. The sleep and wake-up of the GPU is controlled by IIC, and the corresponding state machine is designed to ensure that the GPU automatically enters low-power mode when idle for a long time, thereby reducing energy consumption. When the task is resumed, the current state of the GPU is restored through the save and load mechanism to ensure that the task can continue seamlessly. A persistent storage solution is designed to save the current computing state and important data when the GPU is dormant, so that work can be quickly resumed after waking up, ensuring the security and integrity of the state saving process, and avoiding data loss and computing interruption.

[0124] Furthermore, the system also includes a hot-swap fault monitoring mechanism and a safety protection mechanism. The hot-swap fault monitoring mechanism is used to implement fault monitoring during the hot-swap process, ensuring timely identification and alarming of problems such as abnormal temperatures and poor connections during GPU insertion and removal. The designed self-diagnosis function can automatically detect and record fault information during the hot-swap process for subsequent analysis and improvement. Protection circuits such as overcurrent protection and short-circuit protection are incorporated into the circuit design to ensure the safety of the device during hot-swap operations. User permission management and operation logging are designed to ensure that only authorized users can perform hot-swap operations, avoiding system failures caused by incorrect operations.

[0125] In this application, the energy management agent layer, based on the GPU hot-swappable driver and service, can manage and control all GPUs on the same server (i.e., the same node). Furthermore, users can set different energy management policies based on specific application requirements, including high-efficiency mode, performance mode, and power-saving mode, to flexibly address different tasks and scenarios. The system automatically switches GPU operating modes based on the current task load and user-defined policies to achieve optimal energy efficiency. The system regularly evaluates the effectiveness of different energy management policies and continuously optimizes policy settings by analyzing energy consumption data and performance feedback, ensuring that the system remains optimal under varying load conditions. An effective sleep and wakeup mechanism optimizes GPU energy consumption during low-load or idle states, minimizing energy consumption without compromising service quality. When tasks are idle for extended periods, the GPU can be automatically placed in a low-power or sleep state to reduce unnecessary energy consumption. The system determines whether the GPU can enter sleep mode based on activity monitoring and task scheduling information. Before sleep mode, the software saves the current task state and GPU configuration to memory or persistent storage for rapid restoration upon wakeup. This mechanism is implemented through IIC control, ensuring fast response. When a new task arrives or an existing task needs to continue, the GPU can be automatically woken up and restored to the previous working state to ensure task continuity and user experience. Task priorities can also be classified and managed to ensure that critical tasks are given priority, thereby improving the overall efficiency and responsiveness of the system. Specifically, priorities are set for different tasks and graded according to factors such as the urgency, importance, and resource requirements of the tasks (such as high, medium, and low priority). When scheduling tasks, high-priority tasks are given priority and assigned to GPUs with appropriate loads to ensure the timely completion of critical tasks and optimize resource utilization. Based on real-time monitoring data and task execution status, the system is allowed to dynamically adjust task priorities. For example, when a high-priority task requires resources, its scheduling priority can be temporarily increased.

[0126] This application can also provide in-depth insights into GPU and system operations through comprehensive status monitoring and logging, supporting troubleshooting, performance optimization, and decision support. It can continuously monitor the operating status of each GPU, including load, temperature, energy consumption, fault status, etc., and feed this information back to the management system in real time. It can record all important events and status changes, including task launches, completions, GPU status changes, fault events, etc. The logs should include timestamps, event types, and related parameters to facilitate subsequent analysis. It can also generate regular or on-demand reports by analyzing recorded status and log data to help administrators understand system operating efficiency, energy consumption performance, and potential problems, and support optimization and adjustment.

[0127] In this application, from the perspective of the global intelligent energy consumption management layer, the energy efficiency of each GPUBOX can be optimized through comprehensive scheduling and intelligent algorithms, thereby improving energy utilization efficiency and reducing operating costs. Specifically, it includes: 1) Energy consumption monitoring and analysis mechanism, which is used to monitor and analyze the energy consumption and status of the GPU cluster in real time, provide data support for decision-making, ensure efficient operation of the system and optimize resource utilization; 2) Intelligent task scheduling mechanism, which is used to optimize the scheduling of tasks on the GPU through intelligent algorithms, ensure the efficiency and stability of system operation, and at the same time improve resource utilization efficiency and reduce energy consumption; 3) Cross-node resource coordination mechanism, which is used to coordinate and optimize resources in a multi-node GPU server environment to improve overall energy efficiency and ensure the coordination between nodes. 4) User configuration and custom policy mechanism, which provides flexible user configuration options and enables users to customize energy management policies according to their needs to meet specific application requirements and performance goals; 5) Fault detection and automatic recovery mechanism, which ensures system stability and reliability by monitoring GPU status in real time, quickly responds to faults and performs automatic recovery; 6) Reporting and visualization mechanism, which generates visual energy consumption reports and monitoring dashboards to help users understand the system's energy efficiency performance and facilitate management decisions; 7) Feedback mechanism, which evaluates the overall performance of the GPU cluster through comprehensive data and provides optimization suggestions to help users optimize resource utilization and energy efficiency.

[0128] To effectively manage GPU energy consumption, the following approaches can be adopted for energy consumption monitoring and analysis: First, conduct real-time monitoring. Utilize specialized monitoring tools and sensors to regularly collect key metrics such as energy consumption, temperature, and load for each GPU. Ensure data collection is frequent enough, with updates every second or minute to reflect GPU status in real time. Design a distributed monitoring architecture, utilizing message queues or stream processing frameworks (such as Kafka or Flume) to transmit collected data to processing systems in real time, ensuring efficient data collection and aggregation within large clusters. Next, consider data storage and management. Store collected monitoring data in a time-series database (such as InfluxDB or Prometheus). Design a suitable storage strategy to ensure long-term data preservation and rapid access. Clean and preprocess the collected data to remove noise and outliers. Data analysis is then performed, with regular, automated energy consumption reports generated. These reports include energy consumption trends, temperature fluctuations, load conditions, and key performance indicators (KPIs) such as average energy consumption, peak energy consumption, and GPU utilization for each GPU, ensuring that the reports are easy to read and useful. Data visualization tools (such as Grafana or Tableau) are used to visualize the monitored data, presenting information intuitively through charts and dashboards. Interactive interfaces facilitate in-depth analysis of the performance of specific time periods or GPUs. Energy efficiency bottlenecks are then identified by using data mining techniques to analyze historical energy consumption data. By comparing energy consumption and performance across different time periods or tasks, improper resource allocation or configuration issues can be identified. Improvement opportunities are then identified by analyzing the relationship between energy consumption trends and workload, recommending energy efficiency improvements under high load conditions, and using machine learning algorithms to predict future energy consumption trends to proactively identify potential issues. Finally, decision support is provided, providing management with optimization recommendations based on the monitoring and analysis results, such as adjusting task scheduling strategies, reducing unnecessary GPU usage, and optimizing power management. A feedback mechanism is established to continuously monitor the effectiveness of policy adjustments, ensuring the effectiveness of measures and continuously optimizing energy management based on new data.

[0129] The intelligent task scheduling mechanism consists of four main components: 1) Feature-based scheduling: First, task characteristics are analyzed. Feature analysis is performed on pending tasks to identify attributes such as computational requirements, memory usage, priority, and runtime estimates. Tasks are then categorized into high-priority, real-time tasks and low-priority, batch tasks. GPU status is also evaluated, with each GPU's load, temperature, power consumption, and available memory monitored in real time. Historical performance data is combined to evaluate its performance in processing specific types of tasks and establish a GPU performance model. Intelligent scheduling algorithms, such as those based on genetic algorithms, ant colony algorithms, or deep learning, are then designed and implemented. These algorithms intelligently select the optimal GPU for task allocation based on task characteristics and GPU status, while also considering task dependencies and adhering to the task execution order and resource competition logic.

[0130] 2) Dynamic Policy Adjustment: Continuously analyze real-time data to track GPU load and energy consumption. Leveraging real-time data stream analysis to determine GPU utilization efficiency and energy consumption, the system uses real-time feedback to identify periods of high and low load. During periods of high load, the system automatically increases GPU power to improve performance and meet real-time requirements. During periods of low load, the system automatically reduces power or places idle GPUs in power-saving mode to reduce energy consumption. A user interface is also provided to configure dynamic policy parameters, such as load thresholds and power caps, for flexible scheduling adjustments.

[0131] 3) Forecasting and Planning: Based on historical data analysis, we collect and analyze historical load and energy consumption data, apply statistical analysis methods to identify load patterns and trends, and build load forecasting models using techniques such as time series analysis and regression analysis to predict future GPU load and energy consumption trends. We use machine learning algorithms such as linear regression, decision trees, or neural networks to train historical data to generate load forecasting models. Based on the model's predictions, we plan resource allocation and scheduling strategies in advance. Based on these predictions, we reserve or adjust GPU resources in advance to cope with periods of high load and schedule tasks in advance during periods of low load to optimize resource utilization.

[0132] 4) Scheduling Performance Evaluation: After each task is scheduled, we monitor task execution performance indicators, including task completion time, energy consumption, and GPU load changes, evaluate scheduling effectiveness, and generate feedback reports. We combine scheduling performance feedback with intelligent scheduling algorithms to continuously optimize scheduling strategies and algorithms. We regularly evaluate and update scheduling algorithms to adapt to changing workloads and system environments.

[0133] For cross-node resource coordination, the following implementation methods can be used to achieve efficient cross-node GPU resource coordination: 1) Global resource view: Deploy resource monitoring tools to collect real-time information such as GPU status, CPU load, memory usage, and network bandwidth of each node. This data is integrated to form a global resource view, and the real-time status of all nodes is displayed through a visual interface. Design and implement a load balancing algorithm to dynamically adjust the task allocation of each node based on the data collected from the global view. Strategies such as polling, minimum number of connections, and load weighting are used to ensure even task distribution and avoid node overload or idleness. Set resource usage thresholds for each node. When the node load exceeds the threshold, the system automatically assigns new tasks to the lower-loaded node. Intelligent algorithms are used to continuously optimize the threshold strategy to adapt to different workloads and resource requirements.

[0134] 2) Task Migration: Continuously monitor resource usage across nodes. Based on task characteristics, apply decision trees or other intelligent algorithms for real-time data analysis to identify tasks requiring migration and determine their necessity and feasibility. Design effective task migration strategies to ensure data consistency and task integrity during the migration process, minimize the impact on system performance, and provide a bidirectional migration model. Develop a task migration scheduling system to automatically handle task migration requests, including task suspension, data transfer, and resumption of execution on the target node. Accurately record and manage task status, allowing users to track task progress in real time.

[0135] 3) Coordination Strategy Optimization: Collect and analyze historical task allocation data, evaluate task migration performance and resource utilization, and use machine learning to identify optimal resource allocation and migration patterns under different loads to optimize future scheduling and migration strategies. A user feedback interface is provided to allow users to evaluate migration strategies and results. By analyzing user feedback, the system's decision-making logic is optimized to better align task scheduling and migration with actual needs.

[0136] 4) Performance Monitoring and Evaluation: After each migration task, we monitor node performance changes before and after the migration, such as task completion time, energy consumption, and load changes. We generate migration effectiveness evaluation reports to help the team understand migration performance in different scenarios and further adjust and optimize strategies. We establish a regular evaluation and optimization mechanism to continuously improve cross-node resource coordination strategies and algorithms based on changes in system performance and user needs. We also employ agile development methods to rapidly iterate system functionality to address evolving workloads and market demands.

[0137] Regarding user-configured and customized policy mechanisms, energy management can be implemented from two perspectives: policy setting and custom policy storage and management. For policy setting, the user interface design begins with a user-friendly configuration interface, providing a simple, wizard-style setup process. This allows users to intuitively set and adjust energy management policies and quickly get started. Task priority is also set, with multiple priority levels available to support highly customized task processing strategies, ensuring sufficient resources for critical tasks. For time period control, users can define policies for different time periods based on their needs, such as improving GPU performance during peak hours and reducing energy consumption during off-peak hours. Scheduled tasks can also be used to set regularly executed energy management policies to automatically adjust resource utilization. Furthermore, power capping allows users to customize power limits for each GPU to control overall energy consumption. Task scheduling rules, such as the number of concurrently executed tasks and task migration conditions, can be set to optimize resource allocation. Custom policy storage and management allows users to create and save multiple energy management policy templates. Policy version management facilitates viewing, restoring, and managing different versions of policy configurations. It also provides a strategy testing function, allowing users to simulate strategy effects in a non-production environment and verify the rationality and effectiveness of the settings. Users can adjust strategies based on the test results to ensure that the strategies ultimately applied to the production environment are optimized.

[0138] To effectively address GPU failures and ensure stable system operation, the following implementation methods can be used for fault detection and automatic recovery: 1) Fault Detection: Deploy advanced monitoring tools to monitor GPU temperature, load, power consumption, and operating status in real time. Set fault thresholds. Once any metric exceeds the preset threshold, the system automatically triggers the fault detection mechanism. Leverage machine learning algorithms to analyze historical data and establish a baseline for normal conditions. This automatically identifies potential failure modes and establishes an alert mechanism to immediately notify management and system administrators upon detection of a failure.

[0139] 2) Fault Handling Mechanism: When a GPU failure is detected, the system automatically switches to an available backup GPU, ensuring that task status is preserved during the switchover process and minimizing business impact. An intelligent decision-making system is developed to automatically execute the fault recovery process based on real-time monitoring data and historical fault records. The fault and recovery process is recorded in detail for subsequent analysis and improvement.

[0140] 3) Task Rescheduling: After a failure occurs, the system automatically identifies affected tasks and intelligently selects the most appropriate healthy GPU for rescheduling based on task priority and resource availability. It also automatically generates a failure recovery report, documenting the failure time, cause, and recovery process, providing a basis for subsequent failure analysis. A visual interface allows users to track the failure and recovery process in real time, enhancing system transparency.

[0141] To effectively manage and monitor GPU energy consumption, status, and performance, the following reporting and visualization mechanisms are available: First, intuitive dashboards are designed using visualization tools to display key metrics such as GPU utilization, temperature, and power consumption, allowing users to monitor system status in real time. Users can also customize dashboard layouts, selected display metrics, and chart types to suit their needs. Data stream processing technology ensures real-time dashboard data updates, and an alerting function provides visual alerts to users of anomalies. Second, a scheduled reporting function is provided. The system automatically generates energy consumption and performance reports at a user-defined frequency (such as daily, weekly, or monthly). These reports cover key performance indicators, trend analysis, and anomaly summaries, and users can customize the frequency and content of report generation. Furthermore, the system features a report distribution function, allowing users to send reports to specific team members or managers for collaborative analysis and decision-making. Reports can also be exported in various formats, such as PDF and Excel, for further analysis and archiving.

[0142] To optimize and improve system performance, the following implementation methods can be used for feedback mechanisms: 1) Performance benchmarking: Establish a regular performance benchmarking schedule to evaluate the system's energy efficiency, response time, and task success rate under varying loads. Collect and analyze test data, generating detailed performance reports to inform subsequent optimization efforts. Develop or integrate existing benchmarking tools to ensure that testing covers different types of tasks and load conditions, comprehensively assessing system capabilities, and supporting user-defined test scenarios to simulate diverse workloads.

[0143] 2) Optimization Recommendation Generation: Based on monitoring data and performance evaluation results, data analysis and mining techniques are used to identify bottlenecks and resource waste in the system. Based on task usage patterns, targeted resource allocation and scheduling optimization recommendations are generated. An intelligent optimization system is designed to automatically analyze historical data and generate recommendations for adjusting resource allocation, modifying energy consumption strategies, and task scheduling. Furthermore, simulation tools are provided to enable users to predict the effects of recommendations before implementing them, helping them make more informed decisions.

[0144] The implementation of this technical solution significantly improves GPU energy management under varying task loads. Preliminary predictions and test results indicate that the solution achieves the following energy management benefits: Actual tests show that the application of the intelligent scheduling algorithm reduces overall GPU energy consumption by approximately 20%-30%. In continuously powered inference mode, energy consumption is expected to drop by over 10%. By optimizing power management, GPU power consumption is reduced under low load conditions, extending GPU lifespan by approximately 15%-25%. Dynamically plugging and unplugging GPUs improves system resource utilization by approximately 25%. Performance is maintained when high computing power is required, while energy conservation is achieved under low load conditions. These results demonstrate that this solution not only addresses the shortcomings of existing energy management technologies but also provides an efficient and flexible solution for intelligent GPU management, demonstrating high practical value and market potential. Quantitative testing and analysis, along with practical application results, validate the effectiveness and feasibility of this solution, laying the foundation for further optimization and promotion.

[0145] The GPU BOX energy consumption control method disclosed in this application comprehensively and systematically addresses a series of problems existing in the prior art, starting from task analysis, GPU allocation, and real-time monitoring and control. Specifically, by performing feature analysis on received pending tasks and identifying task characteristics (including task type), this enables the system to gain a deep understanding of the nature and requirements of the tasks, avoiding blind allocation of GPU resources. For example, for some simple inference tasks, if GPU resources suitable for large model pre-training are mistakenly used, it will result in wasted resources and increased energy consumption. However, by accurately identifying task characteristics, the GPU can be more rationally scheduled to execute the task.

[0146] The system obtains historical performance data of all GPU boxes on the same node corresponding to the task type and combines it with the status information of all GPU boxes obtained in real time to build a GPU box performance model. This model can comprehensively consider historical performance and current status, providing a more scientific basis for task allocation. During large-scale model pre-training and fine-tuning, the system can accurately evaluate the performance and reliability of each GPU box based on the model, giving priority to GPU boxes with stable performance and good status to execute tasks, thereby reducing the probability of the entire task failing due to GPU problems and needing to start over, thereby reducing the waste of computing resources and unnecessary energy consumption.

[0147] Based on the GPU Box performance model, task characteristics are analyzed to determine the GPU Box allocation strategy for executing tasks within the same node. Under low-load conditions, the system dynamically allocates the appropriate number and performance of GPU Boxes based on the actual task requirements, preventing all GPUs from maintaining high energy consumption. For example, when a server is processing a small number of tasks, the system can activate only a few GPU Boxes with appropriate performance to execute the tasks, while placing other GPUs in low-power mode. This effectively reduces overall energy consumption, improves energy efficiency, and reduces operational costs.

[0148] Monitor the target GPU BOX's operating information in real time, obtain the GPU BOX energy consumption dynamic control strategy based on the operating information, and execute the strategy. When the GPU generates a lot of heat under long-term high-load operation, the system can adjust the GPU's working status in time based on the monitored temperature and other operating information, such as reducing its computing frequency and task load, to reduce heat generation, protect the electronic components inside the GPU, improve its performance and stability, reduce the frequent replacement of GPUs due to aging, reduce hardware costs, and ensure the normal operation of the server.

[0149] The method of the present application breaks away from the existing simple cumulative manual control method, and realizes the dynamic on-demand activation and use of GPU according to actual task requirements. It can automatically allocate resources and regulate energy consumption according to the characteristics of the task and the status of the GPU without manual intervention, avoiding the mistakes that are easy to occur in manual operation and improving management efficiency. Although the energy consumption control interfaces of different GPU manufacturers are not unified, the method of the present application realizes the unified management of GPU BOXes of different manufacturers by constructing a universal GPU BOX performance model and dynamic regulation strategy. The system focuses on the overall performance and energy consumption of the GPU BOX, and does not rely on specific manufacturer interfaces, thereby simplifying the management process, improving the flexibility and efficiency of management, and being able to adjust the energy consumption status of the server in a timely manner according to task changes.

[0150] In the current AI field, the demand for applications based on fine-tuning and reasoning has exploded, which has led to a significant increase in the frequency of use and computing load of GPU servers. However, there are many drawbacks in the existing GPU server energy consumption management, such as problems with the GPU involved in the task will lead to task failure, the GPU still maintains high energy consumption when the load is low, the GPU is damaged by long-term high-load work, and the energy consumption management method is backward. These problems have caused adverse consequences such as waste of resources and increased costs. The solution disclosed in this application not only needs to be optimized from the hardware structure level, but also has the ability of software intelligent management. With the help of this technology, multi-GPU BOX servers can automatically carry out dynamic intelligent management and optimization of energy consumption under different task states. Through the application of this technology, on the one hand, the energy consumption of GPU BOX servers can be significantly reduced, energy utilization efficiency can be improved, thereby reducing operating costs; on the other hand, it can also effectively extend the service life of GPU BOX, which is of great significance for promoting the sustainable development of the AI ​​field. In a second aspect, the present application discloses a GPU BOX energy consumption control system, which is used to execute the GPU BOX energy consumption control method, including: The server is used to receive pending tasks, perform feature analysis on the pending tasks, and identify task characteristic information, including task type, computing requirements, memory usage, priority information, and estimated task execution time; The historical performance data acquisition module is used to obtain the historical performance data of all GPU boxes on the same node corresponding to the task type; The model building module is used to build a GPU BOX performance model based on the status information and historical performance data of all GPU BOXes obtained in real time; The allocation strategy acquisition module is used to analyze the task characteristic information according to the GPU BOX performance model and obtain the GPU BOX allocation strategy for executing tasks in the same node; An execution module is used to determine the target GPU BOX according to the GPU BOX allocation policy and execute the pending task based on the target GPU BOX; The dynamic control module is used to monitor the operating information of the target GPU BOX in real time, obtain the GPUBOX energy consumption dynamic control strategy based on the operating information, and execute the GPUBOX energy consumption dynamic control strategy.

[0151] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.

[0152] The processor can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is used to execute the computer-readable instructions stored in the memory, causing the computer device to execute all or part of the steps of the GPU BOX energy consumption control method described in various embodiments of the present disclosure.

[0153] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience, this embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the scope of protection of this disclosure.

[0154] like Figure 5 The present invention provides a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 5 The computer device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0155] like Figure 5 As shown, a computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) or programs loaded from a storage device into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0156] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes and hard disks; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 5A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0157] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the energy consumption control method of the GPU BOX of the embodiment of the present disclosure are performed.

[0158] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the energy consumption control method of the GPU BOX in the above-mentioned embodiments of the present disclosure are executed.

[0159] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).

[0160] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

Claims

1. A method for controlling energy consumption of a GPU BOX, characterized in that: include: Performing feature analysis on the received tasks to be processed to identify task characteristic information, wherein the task characteristic information includes task type; Get the historical performance data of all GPU boxes on the same node corresponding to the task type; Building a GPU BOX performance model based on the real-time acquired status information of all GPU BOXes and the historical performance data; Analyze the task characteristic information according to the GPU BOX performance model to obtain a GPU BOX allocation strategy for executing the task in the same node; Determine a target GPU BOX according to the GPU BOX allocation strategy, and execute the task to be processed based on the target GPU BOX; The operating information of the target GPU BOX is monitored in real time, a dynamic control strategy for energy consumption of the GPU BOX is acquired according to the operating information, and the dynamic control strategy for energy consumption of the GPU BOX is executed.

2. The method for controlling energy consumption of a GPU BOX according to claim 1, wherein: The acquisition of historical performance data of all GPU BOXes on the same node corresponding to the task type includes: Obtain performance data from each GPU BOX during its historical task execution and store it in the target database; the performance data includes computing speed, memory bandwidth usage, temperature, power consumption, load requirements, and historical execution time corresponding to different tasks; According to the task type, the performance data of all GPU BOXes of the same node that have executed the task type are queried from the target database.

3. The method for controlling energy consumption of a GPU BOX according to claim 2, wherein: The step of constructing a GPU BOX performance model based on the real-time acquired status information of all GPU BOXes and the historical performance data includes: Obtain status information of all GPU BOXes in real time, including the current load, current temperature, current power consumption, and current available memory of each GPU BOX; respectively determining availability evaluation results of the current load, the current temperature, the current power consumption, and the current available memory; Determine a preset load weight, a preset temperature weight, a preset power consumption weight, and a preset memory weight based on the historical performance data of all GPU BOXs; When all the availability evaluation results are passed, the current load, the current temperature, the current power consumption, and the current available memory are weighted and summed according to the preset load weight, the preset temperature weight, the preset power consumption weight, and the preset memory weight to obtain an availability score for a single GPU BOX; A GPU BOX performance model is constructed based on the availability scores of all GPU BOXs and the historical execution times corresponding to all GPU BOXs.

4. The method for controlling energy consumption of a GPU BOX according to claim 3, wherein: The respectively determining the availability evaluation results of the current load, the current temperature, the current power consumption, and the current available memory includes: Determine whether the current load is lower than a preset load threshold; if so, determine that the availability evaluation result of the current load is passed; if not, determine that the availability evaluation result of the current load is failed; Determine whether the current temperature is lower than a preset temperature threshold, if so, determine that the availability evaluation result of the current temperature is passed; if not, determine that the availability evaluation result of the current temperature is failed; Determine whether the current power consumption is lower than a preset power consumption threshold; if so, determine that the availability evaluation result of the current power consumption is passed; if not, determine that the availability evaluation result of the current power consumption is failed; Determine whether the current available memory is higher than a preset memory threshold; if so, determine that the availability evaluation result of the current available memory is passed; if not, determine that the availability evaluation result of the current available memory is failed.

5. The method for controlling energy consumption of a GPU BOX according to claim 3, wherein: The analyzing the task characteristic information according to the GPU BOX performance model to obtain a GPU BOX allocation strategy for executing the task in the same node includes: According to the GPU BOX performance model, a GPU BOX whose historical execution time is not longer than the estimated task execution time and whose availability score is not lower than a preset score threshold is obtained and recorded as a qualified GPU BOX; Encode the task allocation of each qualified GPU BOX into a chromosome; Determine a fitness function for evaluating the quality of each chromosome, where the fitness function is the inverse of the estimated task execution time; According to the value of the fitness function, the chromosomes with fitness greater than the preset fitness threshold are selected for crossover and mutation operations to generate new chromosomes; Based on the new chromosome, the chromosome selection, crossover and mutation operations are repeated until the optimal chromosome is found when the fitness value converges; The GPU BOX allocation strategy for executing tasks in the same node is obtained based on the optimal chromosome.

6. The method for controlling energy consumption of a GPU BOX according to claim 1, wherein: The real-time monitoring of the target GPU BOX's operating information, obtaining a GPU BOX energy consumption dynamic control strategy based on the operating information, and executing the GPU BOX energy consumption dynamic control strategy includes: Real-time monitoring of the operating information of the target GPU BOX; the operating information includes actual computing load, actual temperature information, and actual power consumption information when executing the task to be processed; Obtaining the GPU Box corresponding to the actual computing load that is lower than the load lower limit threshold, recording it as the GPU Box to be adjusted, and obtaining the actual low-load operation time of each GPU Box to be adjusted; wherein the load lower limit threshold is 15% of the rated load of the corresponding GPU Box; According to the actual low-load operation time, a corresponding GPU BOX energy consumption dynamic control strategy is determined, and GPU BOX energy consumption is controlled according to the GPU BOX energy consumption dynamic control strategy.

7. The method for controlling energy consumption of a GPU BOX according to claim 6, wherein: Determining a corresponding GPU BOX energy consumption dynamic control strategy according to the actual low-load operation time, and performing GPU BOX energy consumption control according to the GPU BOX energy consumption dynamic control strategy, includes: Determine whether the total number of the GPU BOX to be adjusted is greater than N / 2. If so, determine whether the average load of all the target GPU BOX is lower than the load lower limit threshold. If so, randomly select The target GPU BOX is put into sleep mode. If not, a random The GPU BOX to be adjusted is put into hibernation; in, ; , N is the total number of all the target GPU BOXes, The total number of all GPU BOXes to be adjusted; If the total number of the GPU BOXes to be adjusted is not greater than N / 2, the power consumption of all the GPU BOXes to be adjusted is adjusted to 5%-13% of the rated load of the corresponding GPU BOXes.

8. The method for controlling energy consumption of a GPU BOX according to claim 7, wherein: When all the target GPU BOXes in the same node are divided into several groups, the target GPU BOXes in each group are recorded as a subgroup, and N is the total number of all the target GPU BOXes in each subgroup. The total number of all the GPU BOXes to be adjusted in each subgroup; The operating mode in which the total number of the GPU BOXes to be adjusted is greater than N / 2 and the average load of all the target GPU BOXes is lower than the load lower limit threshold is recorded as the target mode; Determine all the subgroups containing the target pattern and record them as target groups; When there are at least two target groups, the load is migrated to one target group for processing, and all GPU BOXes in the target group without tasks are put into hibernation.

9. The method for controlling energy consumption of a GPU BOX according to claim 1, wherein: The real-time monitoring of the operating information of the target GPU BOX, obtaining a GPU energy consumption dynamic control strategy according to the operating information, and executing the GPU energy consumption dynamic control strategy further includes: Obtain the historical computing power demand of the same node corresponding to the task type during the non-working period and the computing power demand of the same node during the peak period; Determine the historical level of computing power demand during non-working hours; The energy consumption control strategy of all GPU BOXes in the same node is determined according to the level, and the energy consumption of all GPU BOXes in the same node is dynamically controlled according to the energy consumption control strategy.

10. The method for controlling energy consumption of a GPU BOX according to claim 9, wherein: Determining the level of historical computing power demand during the non-working period includes: when When the historical computing power demand level during the non-working period is determined to be the first level; when When the historical computing power demand level during the non-working period is determined to be the second level; in, The historical computing power demand of the same node during non-working period. It is the computing power demand of the same node during peak period.

11. The method for controlling energy consumption of a GPU BOX according to claim 10, wherein: Determining the energy consumption control strategy of all GPU BOXes in the same node according to the level, and dynamically controlling the energy consumption of all GPU BOXes in the same node according to the energy consumption control strategy, includes: When the level is the first level, determining the energy consumption control strategy for all GPU BOXes in the same node to disconnect the power supply of at least 1 / 3 of the GPU BOXes in the same node; When the level is the second level, determining the energy consumption control strategy for all GPU BOXes in the same node includes disconnecting the power supply of at least two-thirds of the GPU BOXes in the same node, or randomly selecting two GPU BOXes to retain power supply; According to the energy consumption control strategy of all GPU BOXes in the same node, the power supply switch of the corresponding GPU BOX is controlled through the I²C chip.

12. The method for controlling energy consumption of a GPU BOX according to claim 11, wherein: Also includes: When several target GPU BOXes executing tasks continuously run at no less than 75% of the rated load of the corresponding GPU for more than a preset high-load running time threshold, other target GPU BOXes are started to execute the tasks.

13. A GPU BOX energy consumption control system, configured to execute the GPU BOX energy consumption control method according to any one of claims 1 to 12, characterized in that: include: The server is configured to receive tasks to be processed, perform feature analysis on the tasks to be processed, and identify task characteristic information, wherein the task characteristic information includes task type, computing requirements, memory usage, priority information, and estimated task execution time; A historical performance data acquisition module is used to obtain historical performance data of all GPU BOXes on the same node corresponding to the task type; A model building module is used to build a GPU BOX performance model based on the status information of all GPU BOXs obtained in real time and the historical performance data; An allocation strategy acquisition module, configured to analyze the task characteristic information according to the GPU BOX performance model and acquire a GPU BOX allocation strategy for executing the task in the same node; an execution module, configured to determine a target GPU BOX according to the GPU BOX allocation policy, and execute the task to be processed based on the target GPU BOX; The dynamic control module is used to monitor the operating information of the target GPU BOX in real time, obtain the GPU BOX energy consumption dynamic control strategy according to the operating information, and execute the GPU BOX energy consumption dynamic control strategy.

14. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the energy consumption control method of the GPU BOX according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the energy consumption control method of the GPU BOX according to any one of claims 1 to 12.

16. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • GPU internal energy consumption optimization method based on task balance scheduling

    CN109992385A

  • Computing system and method for GPU (Graphics Processing Unit) computing power scheduling

    CN119645661A

  • Network resource hierarchical coordination method and system based on multiple agents

    CN120091245A

Cited By

  • Performance analysis and tuning system based on memory access density program

    CN121681308A