Efficient GPU computing power server resource optimization method and system
By evaluating the GPU computing power server and dynamically adjusting the Batch size, optimizing the DataLoader and model architecture, and dynamically adjusting resource allocation, the problem that traditional GPU resource management methods cannot fully utilize the GPU parallel processing capabilities, achieving more efficient resource utilization and training speed.
Patent Information
- Application Number
- CN202510145314.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Traditional GPU resource management methods cannot fully utilize the GPU's parallel processing capabilities, resulting in resource waste and performance bottlenecks, especially when dynamic changes during task execution are not effectively considered.
By conducting a preliminary evaluation of the hardware configuration of the GPU computing power server, the Batch size is dynamically adjusted, the DataLoader performance is optimized, the model architecture is adjusted, and the allocation of CPU, GPU, memory and storage resources is dynamically adjusted according to the real-time load conditions, so that all components can work together.
It improves the initial utilization efficiency of resources, reduces the idleness or overload of resources caused by inappropriate Batch size, significantly improves the overall training speed and model inference speed and accuracy, and avoids resource waste and performance bottlenecks.
Smart Images

Figure CN120147101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to an efficient method and system for optimizing the resources of a GPU computing power server. Background Art
[0002] A GPU server is a server composed of a Graphics Processing Unit (GPU) for large-scale data computing and processing, which provides strong support for complex computing tasks with its efficient parallel processing ability. However, with the continuous growth of computing requirements, how to more effectively utilize and optimize the resources of a GPU computing power server and improve the training speed and efficiency has become an urgent problem in the industry.
[0003] Traditional GPU resource management methods are sometimes based on static configuration, that is, various parameters and resource allocations are preset before the task starts. This method ignores the dynamic changes during the task execution, such as changes in model complexity, data flow efficiency, and real-time load conditions, etc. Therefore, it sometimes cannot fully utilize the parallel processing ability of the GPU, resulting in resource waste and performance bottlenecks. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an efficient method and system for optimizing the resources of a GPU computing power server, which can improve the initial utilization efficiency of resources.
[0005] To solve the above technical problem, the technical solution of the present invention is as follows:
[0006] In a first aspect, an efficient method for optimizing the resources of a GPU computing power server, the method includes:
[0007] Step 1, preliminarily evaluate the hardware configuration of the GPU computing power server to obtain an evaluation result, including GPU model, quantity, and memory capacity; configure initial training parameters according to the evaluation result, including Batch size;
[0008] Step 2, based on the evaluation result, combine the complexity of the current training task and the parallel processing ability of the GPU, and dynamically adjust the Batch size to obtain an adjustment result;
[0009] Step 3, optimize the performance of the DataLoader according to the adjustment result to obtain the optimized data flow situation;
[0010] Step 4, adjust the model architecture according to the optimized data flow situation to match it with the parallel processing ability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain an adjusted model;
[0011] Step 5: Dynamically adjust the allocation of CPU, GPU, memory, and storage resources according to the adjusted model and the real-time load during the training process, so that each component can work collaboratively.
[0012] Furthermore, conduct a preliminary evaluation of the hardware configuration of the GPU computing power server to obtain the evaluation results, including GPU model, quantity, and memory capacity, including:
[0013] Define an evaluation function to quantify the performance of different hardware configuration schemes;
[0014] Determine the optional range of the hardware configuration, including different models of GPUs, the range of GPU quantities, and the range of memory capacities. The optional range of the hardware configuration constitutes the solution space, that is, all configuration schemes that can be searched in the simulated annealing algorithm;
[0015] Set the initial temperature of the simulated annealing algorithm, the speed of temperature decrease, and the temperature at termination; at each temperature, set the number of new solutions of the simulated annealing algorithm;
[0016] Randomly select an initial hardware configuration scheme from the configuration space as the current solution;
[0017] At the current temperature, generate a neighboring solution of the current solution; use the evaluation function to calculate the performance of the neighboring solution; according to the Metropolis criterion, decide whether to accept the neighboring solution as the new current solution, and perform iterative operations until the number of iterations at the current temperature is reached, then lower the temperature and repeat the entire iterative process until the temperature drops to the termination temperature to obtain the final hardware configuration scheme output by the simulated annealing algorithm.
[0018] Furthermore, the calculation formula of the evaluation function is:
[0019]
[0020] where P i is the performance score of the i-th GPU, P b is the benchmark performance value; M t is the total memory capacity; M b is the benchmark memory capacity; B t is the total memory bandwidth; B b is the benchmark memory bandwidth; D gg is the data transfer rate between GPUs; D gc is the data transfer rate between the GPU and the CPU; D b is the benchmark data transfer rate; R is the read speed; W is the write speed; N is the network bandwidth; I b is the benchmark I / O performance value; w 1 、w 2 、w 3 and w4 is the weight coefficient of each performance index.
[0021] Furthermore, based on the evaluation results, in combination with the complexity of the current training task and the parallel processing ability of the GPU, dynamically adjust the Batch size to obtain the adjustment result, including:
[0022] According to the result of the evaluation function to obtain the computing power of the GPU;
[0023] Set an initial Batch according to the computing power of the GPU. During the training process, dynamically adjust the Batch size according to the real-time monitored metrics, as follows:
[0024] If the GPU utilization rate is significantly lower than its maximum capacity, gradually increase the Batch;
[0025] If insufficient memory occurs during the training process, reduce the Batch.
[0026] Furthermore, according to the optimized data flow situation, adjust the model architecture to match the parallel processing ability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain the adjusted model, including:
[0027] Evaluate the performance of the current model when processing the optimized data flow, including training speed, memory occupancy, and accuracy metrics;
[0028] Analyze the data batch size, transmission speed provided by the optimized DataLoader, and the utilization rate of the GPU; according to the analysis results, locate the performance bottlenecks in the model architecture; set the adjustment target;
[0029] Adjust the number of layers of the model according to the parallel processing ability of the GPU and the data flow situation;
[0030] Use the ResNet style to improve the information flow and gradient flow, and adjust the number of channels of the convolutional layer;
[0031] Use different widths of convolution kernels in different layers to capture features of different scales;
[0032] Select the corresponding convolution kernel size according to the task requirements and GPU characteristics;
[0033] Conduct experimental verification on the model after each adjustment to obtain the verification result;
[0034] According to the verification results, iteratively adjust the model architecture until an efficiency balance is achieved to obtain the adjusted model.
[0035] Furthermore, according to the adjusted model and the real-time load situation during the training process, dynamically adjust the allocation of CPU, GPU, memory, and storage resources to enable the components to work together, including:
[0036] Based on the monitoring data, identify the performance bottlenecks of the system; predict future resource requirements based on historical data and the current load situation;
[0037] Dynamically adjust the parallelism of data preprocessing and model training tasks according to the CPU usage rate;
[0038] Adjust the batch size according to the GPU utilization rate;
[0039] Monitor the memory occupancy and dynamically adjust the data caching strategy according to the memory usage;
[0040] Optimize the data reading and writing strategies according to the storage I / O waiting time; utilize Docker technology to achieve dynamic resource allocation and isolation;
[0041] Collect the performance metrics during the training process and continuously iterate and optimize the resource allocation strategy according to the performance feedback and resource usage.
[0042] Furthermore, the performance metrics include training speed and accuracy.
[0043] In the second aspect, an efficient GPU computing power server resource optimization system includes:
[0044] An evaluation module for preliminarily evaluating the hardware configuration of the GPU computing power server to obtain an evaluation result, including GPU model, quantity, and memory capacity; configure initial training parameters according to the evaluation result, including the Batch size;
[0045] A dynamic adjustment module for dynamically adjusting the Batch size based on the evaluation result, combined with the complexity of the current training task and the parallel processing ability of the GPU, to obtain an adjustment result;
[0046] An optimization module for optimizing the performance of the DataLoader according to the adjustment result to obtain the optimized data flow situation;
[0047] An adjustment module for adjusting the model architecture according to the optimized data flow situation to match it with the parallel processing ability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain an adjusted model;
[0048] An allocation module for dynamically adjusting the allocation of CPU, GPU, memory, and storage resources according to the adjusted model and the real-time load situation during the training process to enable the components to work together.
[0049] In a third aspect, a computing device includes:
[0050] One or more processors;
[0051] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method described above.
[0052] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the method described above.
[0053] The above solution of the present invention has at least the following beneficial effects:
[0054] By initially evaluating the hardware configuration of the GPU computing power server and configuring initial training parameters according to the evaluation results, the present invention can ensure that the training task matches the hardware resources at the beginning, thereby improving the initial utilization efficiency of resources.
[0055] The present invention dynamically adjusts the Batch size in combination with the complexity of the current training task and the parallel processing ability of the GPU. This strategy not only improves the data throughput during the training process but also effectively reduces the situation of resource idling or overloading caused by an inappropriate Batch size.
[0056] According to the adjustment result of the Batch size, the present invention optimizes the performance of the DataLoader, which can significantly reduce the time for data loading and preprocessing, optimize the data flow situation, ensure that the data can be supplied to the GPU for calculation in a timely and efficient manner, and thus improve the overall training speed.
[0057] The present invention adjusts the model architecture according to the optimized data flow situation, including adjusting the number of layers, width, and convolution kernel size of the model, making the model more adaptable to the parallel processing ability of the GPU. This improvement in the matching degree not only accelerates the training process of the model but also helps to improve the inference speed and accuracy of the model.
[0058] By dynamically adjusting the allocation of CPU, GPU, memory, and storage resources according to the adjusted model and the real-time load situation during the training process, the present invention ensures the collaborative work among components. This dynamic resource allocation mechanism can respond to the load changes during the training process in real time, avoid waste of resources and the occurrence of performance bottlenecks, and thus significantly improve the overall performance and stability of the GPU computing power server. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic flowchart of an efficient GPU computing power server resource optimization method provided by an embodiment of the present invention.
[0060] Figure 2 It is a schematic diagram of an efficient GPU computing power server resource optimization system provided by an embodiment of the present invention. Specific implementation manners
[0061] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0062] As Figure 1 shown, an embodiment of the present invention proposes an efficient GPU computing power server resource optimization method, and the method includes the following steps:
[0063] Step 1, preliminarily evaluate the hardware configuration of the GPU computing power server to obtain an evaluation result, including the GPU model, quantity, and memory capacity; configure initial training parameters according to the evaluation result, including the Batch size; wherein, Batc-size is an important hyperparameter in deep learning and machine learning, which determines the amount of data processed by the model at one time during training;
[0064] Step 2, based on the evaluation result, dynamically adjust the Batch size in combination with the complexity of the current training task and the parallel processing ability of the GPU to obtain an adjustment result;
[0065] Step 3, optimize the performance of the DataLoader according to the adjustment result to obtain the optimized data flow situation; wherein, DataLoader is an iterable object in the pytorch framework, which can load data from a given dataset and provide functions such as batch processing, data shuffling, and multi-process loading;
[0066] Step 4, according to the optimized data flow situation, adjust the model architecture to match the parallel processing ability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain an adjusted model;
[0067] Step 5, dynamically adjust the allocation of CPU, GPU, memory, and storage resources according to the adjusted model and the real-time load situation during the training process so that the components work together.
[0068] In the embodiments of the present invention, by initially evaluating the hardware configuration of the GPU computing power server and configuring the initial training parameters according to the evaluation results, the present invention can ensure that the training task matches the hardware resources at the beginning, thereby improving the initial utilization efficiency of the resources. The present invention combines the complexity of the current training task and the parallel processing ability of the GPU to dynamically adjust the Batch size. This strategy not only improves the data throughput during the training process but also effectively reduces the situation of resource idle or overload caused by inappropriate Batch size. According to the adjustment result of the Batch size, the performance of the DataLoader is optimized. The present invention can significantly reduce the time of data loading and preprocessing, optimize the data flow situation, ensure that the data can be supplied to the GPU for calculation in a timely and efficient manner, and thus improve the overall training speed. The present invention adjusts the model architecture according to the optimized data flow situation, including adjusting the number of layers, width, and convolution kernel size of the model, making the model more adaptable to the parallel processing ability of the GPU. This improvement in matching degree can not only accelerate the training process of the model but also help improve the inference speed and accuracy of the model. By dynamically adjusting the allocation of CPU, GPU, memory, and storage resources according to the adjusted model and the real-time load situation during the training process, the present invention ensures the collaborative work among components. This dynamic resource allocation mechanism can respond to the load changes during the training process in real time, avoid the waste of resources and the emergence of performance bottlenecks, and thus significantly improve the overall performance and stability of the GPU computing power server.
[0069] In the embodiments of the present invention, the hardware configuration of the GPU computing power server is initially evaluated to obtain evaluation results, including the GPU model, quantity, and memory capacity, including:
[0070] Define an evaluation function for quantifying the performance of different hardware configuration schemes;
[0071] Determine the optional range of the hardware configuration, including different models of GPUs, the quantity range of GPUs, and the range of memory capacity. The optional range of the hardware configuration constitutes the solution space, that is, all configuration schemes that can be searched in the simulated annealing algorithm;
[0072] Set the initial temperature of the simulated annealing algorithm, the speed of temperature decrease, and the temperature at termination; at each temperature, set the number of new solutions of the simulated annealing algorithm;
[0073] Randomly select an initial hardware configuration scheme from the configuration space as the current solution;
[0074] At the current temperature, generate a neighboring solution of the current solution; use the evaluation function to calculate the performance of the neighboring solution; according to the Metropolis criterion (a rule for optimization algorithms), decide whether to accept the neighboring solution as the new current solution, and perform iterative operations until the number of iterations at the current temperature is reached, then lower the temperature and repeat the entire iterative process until the temperature drops to the termination temperature to obtain the final hardware configuration solution output by the simulated annealing algorithm.
[0075] In the embodiment of the present invention, by defining the evaluation function, the performance of different hardware configuration solutions can be quantified, making the configuration selection process more objective and accurate. The simulated annealing algorithm can search for the globally optimal or approximately optimal configuration solution in the solution space, avoiding the problem of being trapped in local optima in traditional methods. This method constructs a flexible solution space by determining the optional range of hardware configurations, including different models of GPUs, the range of the number of GPUs, the range of memory capacities, etc. This means that this method is not only applicable to specific models of GPU servers, but also can be extended to hardware configurations with different specifications and performances according to actual needs, having wide applicability. The simulated annealing algorithm can, while accepting worse solutions, jump out of local optima with a certain probability by simulating the physical annealing process, thus having the opportunity to find the globally optimal solution. This search strategy performs well in dealing with large-scale and complex hardware configuration problems and can quickly converge to high-quality solutions. By setting parameters such as the initial temperature, the temperature drop rate, and the termination temperature of the simulated annealing algorithm, the search accuracy and convergence speed of the algorithm can be adjusted according to the actual situation. At the same time, this method can also be combined with other optimization algorithms to form a more flexible and efficient hybrid optimization strategy. By optimizing the hardware configuration solution, it is possible to ensure the best performance improvement within a limited budget. This not only helps to reduce the operating costs of enterprises, but also can improve the return on investment of computing power servers.
[0076] In the embodiment of the present invention, the calculation formula of the evaluation function is:
[0077]
[0078] where P i is the performance score of the i-th GPU, and P b is the benchmark performance value; M t is the total memory capacity; M b is the benchmark memory capacity; B t is the total memory bandwidth; B b is the benchmark memory bandwidth; D gg is the data transfer rate between GPUs; D gc is the data transfer rate between the GPU and the CPU; D b is the benchmark data transfer rate; R is the read speed; W is the write speed; N is the network bandwidth; Ib is the reference I / O performance value; w 1 and w 2 and w 3 and w 4 are the weight coefficients of each performance metric.
[0079] In the embodiments of the present invention, the evaluation function covers multiple aspects such as GPU performance, memory capacity and bandwidth, data transfer rate, and I / O performance, ensuring a comprehensive evaluation of the server hardware performance. This helps to discover potential performance bottlenecks, thereby guiding a more reasonable selection of hardware configurations. By introducing the weight coefficients w 1 and w 2 and w 3 and w 4 , the evaluation function can flexibly balance each performance metric according to actual needs. For example, for compute-intensive tasks, the weight of GPU performance can be increased; while for data transfer-intensive tasks, the weight of data transfer rate can be increased. The design of the evaluation function allows for expansion and customization according to technological development and user requirements. With the emergence of new technology metrics, they can be easily incorporated into the evaluation function, and their importance can be reflected by adjusting the weight coefficients. By quantifying each performance metric and comparing it with the reference value, the evaluation function can objectively evaluate the performance of different hardware configurations. This avoids the bias caused by subjective judgment and improves the accuracy of the evaluation results. The evaluation function provides strong support for the optimization decision of hardware configurations. By calculating the evaluation scores of different configuration schemes, their performance advantages and disadvantages can be intuitively compared, so as to select the best or most suitable configuration scheme for the current needs.
[0080] In the embodiments of the present invention, based on the evaluation results, in combination with the complexity of the current training task and the parallel processing ability of the GPU, the Batch size is dynamically adjusted to obtain the adjustment result, including:
[0081] According to the result of the evaluation function to obtain the computing power of the GPU;
[0082] According to the computing power of the GPU, set an initial Batch, and during the training process, dynamically adjust the Batch size according to the real-time monitored metrics, specifically as follows:
[0083] If the GPU utilization rate is significantly lower than its maximum capacity, gradually increase the Batch;
[0084] If memory shortage occurs during the training process, reduce the Batch.
[0085] In the embodiments of the present invention, by dynamically adjusting the Batch size, the method can ensure the full utilization of GPU resources. When the GPU utilization rate is significantly lower than its maximum capacity, gradually increasing the Batch size can effectively increase the load of the GPU and avoid resource waste. During the training process, if the Batch is set too large, it may cause insufficient GPU memory. By monitoring in real time and dynamically reducing the Batch size, the method can effectively prevent training interruption caused by memory overflow and ensure the stability and continuity of training. An appropriate Batch size is crucial for both training speed and model convergence. Dynamically adjusting the Batch size can improve the amount of data processed per iteration as much as possible while ensuring training stability, thereby improving training efficiency. Different training tasks and hardware configurations may require different Batch size settings. The method can automatically adjust the Batch size according to the actual situation without manual intervention by the user, improving the adaptability and usability of the method. By dynamically adjusting the Batch size, the method can find a balance between computing power and memory resources, fully utilizing the computing power of the GPU while avoiding waste or overconsumption of memory resources.
[0086] In the embodiments of the present invention, according to the optimized data flow situation, the model architecture is adjusted to match the parallel processing ability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain an adjusted model, including:
[0087] Evaluate the performance of the current model when processing the optimized data flow, including training speed, memory occupancy, and accuracy metrics;
[0088] Analyze the data batch size, transmission speed provided by the optimized DataLoader, and the utilization rate of the GPU; according to the analysis results, locate the performance bottlenecks in the model architecture; set adjustment targets;
[0089] Adjust the number of layers of the model according to the parallel processing ability of the GPU and the data flow situation;
[0090] Use the ResNet style to improve the information flow and gradient flow, and adjust the number of channels in the convolutional layer;
[0091] Use different widths of convolution kernels in different layers to capture features of different scales;
[0092] Select the corresponding convolution kernel size according to the task requirements and GPU characteristics;
[0093] Conduct experimental verification on the model after each adjustment to obtain verification results;
[0094] According to the verification results, iteratively adjust the model architecture until an efficiency balance is achieved to obtain an adjusted model.
[0095] In the embodiments of the present invention, by adjusting the number of layers, width, and convolution kernel size of the model, the model can be made more suitable for the parallel processing characteristics of the GPU, thereby improving the training speed. This optimization can ensure smoother data flow, reduce the idle time of computing resources, and accelerate the training process. A reasonable adjustment of the model architecture can reduce the memory occupancy during the training process of the model. By streamlining the number of model layers and channels and selecting an appropriate size of the convolution kernel, the memory consumption can be effectively reduced, enabling the processing of larger-scale datasets or more complex models with limited GPU memory resources. The adjustment of the model architecture not only considers computational efficiency but also takes into account the expressive ability of the model. By using convolution kernels of different widths in different layers to capture features at different scales, the feature extraction ability of the model can be improved, which may further improve the accuracy of the model. Flexibly adjusting the model architecture according to the task requirements and GPU characteristics enables the model to better adapt to different application scenarios. This flexibility enables the model to remain efficient and accurate when facing diverse tasks. By iteratively adjusting the model architecture and conducting experimental verification, while ensuring performance, the optimization of resource utilization can be achieved. This balance not only improves the training efficiency of the model but also ensures the rational use of resources and avoids unnecessary waste.
[0096] In the embodiments of the present invention, according to the adjusted model and the real-time load situation during the training process, the allocation of CPU, GPU, memory, and storage resources is dynamically adjusted to enable the components to work together, including:
[0097] Based on the monitoring data, identify the performance bottlenecks of the system; predict future resource requirements based on historical data and the current load situation;
[0098] According to the CPU usage rate, dynamically adjust the parallelism of data preprocessing and model training tasks;
[0099] According to the GPU utilization rate, adjust the batch size;
[0100] Monitor the memory occupancy situation, and dynamically adjust the data caching strategy according to the memory usage;
[0101] According to the storage I / O waiting time, optimize the data reading and writing strategy; utilize Docker technology to achieve dynamic resource allocation and isolation;
[0102] Collect performance metrics during the training process, and continuously iterate and optimize the resource allocation strategy according to the performance feedback and resource usage; the performance metrics include training speed and accuracy.
[0103] In the embodiments of the present invention, by dynamically adjusting the allocation of resources, it can be ensured that each component (CPU, GPU, memory, and storage) in the system is utilized most efficiently. This avoids waste of resources, especially when the load is low or the demand changes. By real-time monitoring and identifying the performance bottlenecks of the system, the resource allocation can be quickly adjusted to alleviate these bottlenecks. This helps to keep the system running efficiently and avoid affecting the overall performance due to insufficient resources in a certain link. Adjusting the parallelism and batch size of tasks according to the real-time utilization rates of the CPU and GPU can ensure the full utilization of computing resources during the model training process, thereby improving the training speed. At the same time, reasonable resource allocation also helps to improve the accuracy of the model because the training process is more stable and efficient. By monitoring the memory occupancy and storage I / O waiting time and adjusting the data cache and I / O policies accordingly, the risk of memory overflow and storage bottlenecks can be reduced, thereby enhancing the stability of the system. Using Docker technology to achieve dynamic allocation and isolation of resources not only improves the flexibility of the system but also makes the system easier to expand. This means that in the face of training tasks of different scales or demand changes, the system can quickly adapt and provide the required resources. By collecting the performance metrics during the training process and continuously optimizing the resource allocation strategy according to the performance feedback and resource usage, it can be ensured that the system is always in the best state. This continuous optimization process helps to improve the overall performance and efficiency of the system.
[0104] An efficient GPU computing power server resource optimization system, comprising:
[0105] An evaluation module, used to preliminarily evaluate the hardware configuration of the GPU computing power server to obtain an evaluation result, including the GPU model, quantity, and memory capacity; configure initial training parameters according to the evaluation result, including the Batch size;
[0106] A dynamic adjustment module, used to dynamically adjust the Batch size based on the evaluation result, combined with the complexity of the current training task and the parallel processing ability of the GPU, to obtain an adjustment result;
[0107] An optimization module, used to perform performance optimization on the DataLoader according to the adjustment result to obtain the optimized data flow situation;
[0108] An adjustment module, used to adjust the model architecture according to the optimized data flow situation to match the parallel processing ability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain an adjusted model;
[0109] An allocation module, used to dynamically adjust the allocation of CPU, GPU, memory, and storage resources according to the adjusted model and the real-time load situation during the training process, so that each component works in coordination.
[0110] An embodiment of the present invention further provides a computer-readable storage medium storing instructions, which, when run on a computer, cause the computer to execute the method described above. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
[0111] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An efficient GPU computing server resource optimization method, characterized in that: The method comprises: Step 1: Perform a preliminary evaluation of the hardware configuration of the GPU computing server to obtain the evaluation results, including the GPU model, quantity, and memory capacity; configure the initial training parameters, including the batch size, based on the evaluation results; Step 2: Based on the evaluation results, combined with the complexity of the current training task and the parallel processing capability of the GPU, dynamically adjust the batch size to obtain the adjustment result; Step 3: According to the adjustment results, optimize the performance of DataLoader to obtain the optimized data flow; Step 4: According to the optimized data flow, adjust the model architecture to match the parallel processing capability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain the adjusted model. Step 5: Dynamically adjust the allocation of CPU, GPU, memory, and storage resources based on the adjusted model and the real-time load during training to enable the components to work together.
2. The efficient GPU computing server resource optimization method according to claim 1, characterized in that: Conduct a preliminary assessment of the hardware configuration of the GPU computing server to obtain the assessment results, including the GPU model, quantity, and memory capacity, including: Define an evaluation function to quantify the performance of different hardware configurations; Determine the optional range of hardware configuration, including different models of GPUs, the range of GPU quantities, and the range of memory capacity. The optional range of hardware configuration constitutes the solution space, that is, all configuration solutions that can be searched in the simulated annealing algorithm; Set the initial temperature of the simulated annealing algorithm, set the speed of temperature drop and the temperature at the end; at each temperature, set the number of new solutions of the simulated annealing algorithm; Randomly select an initial hardware configuration solution from the configuration space as the current solution; At the current temperature, generate a neighboring solution of the current solution; use the evaluation function to calculate the performance of the neighboring solution; decide whether to accept the neighboring solution as the new current solution according to the Metropolis criterion, iterate until the number of iterations at the current temperature is reached, lower the temperature, and repeat the entire iterative process until the temperature drops to the termination temperature to obtain the final hardware configuration solution output by the simulated annealing algorithm.
3. The efficient GPU computing server resource optimization method according to claim 2, characterized in that: The calculation formula of the evaluation function is: Among them, P i is the performance score of the ith GPU, P b is the benchmark performance value; M t is the total memory capacity; M b is the base memory capacity; B t is the total memory bandwidth; B b is the baseline memory bandwidth; D gg is the data transfer rate between GPUs; D gc is the data transfer rate between GPU and CPU; D b is the benchmark data transfer rate; R is the read speed; W is the write speed; N is the network bandwidth; I b is the benchmark I / O performance value; w1, w2, w3, and w4 are the weight coefficients of each performance indicator.
4. The efficient GPU computing server resource optimization method according to claim 3, characterized in that: Based on the evaluation results, combined with the complexity of the current training task and the parallel processing capability of the GPU, the batch size is dynamically adjusted to obtain the adjustment results, including: According to the result of the evaluation function, the computing power of the GPU is obtained; According to the computing power of the GPU, an initial Batch is set. During the training process, the Batch size is dynamically adjusted according to the real-time monitoring indicators, as follows: If the GPU utilization is significantly lower than its maximum capacity, gradually increase the batch size. If insufficient memory occurs during training, reduce the batch size.
5. The efficient GPU computing server resource optimization method according to claim 4, characterized in that: According to the optimized data flow, adjust the model architecture to match the parallel processing capability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain the adjusted model, including: Evaluate the performance of the current model when processing the optimized data stream, including training speed, memory usage, and accuracy metrics; Analyze the data batch size, transmission speed, and GPU utilization provided by the optimized DataLoader; locate the performance bottlenecks in the model architecture based on the analysis results; and set adjustment goals; Adjust the number of model layers based on the GPU's parallel processing capabilities and data flow; Use ResNet style to improve information flow and gradient flow, and adjust the number of channels in the convolutional layer; Use convolution kernels of different widths in different layers to capture features of different scales; Select the corresponding convolution kernel size according to task requirements and GPU characteristics; Conduct experimental verification on each adjusted model to obtain verification results; According to the verification results, the model architecture is iteratively adjusted until the efficiency balance is achieved to obtain the adjusted model.
6. The efficient GPU computing server resource optimization method according to claim 5, characterized in that: Dynamically adjust the allocation of CPU, GPU, memory, and storage resources based on the adjusted model and the real-time load during training to enable components to work together, including: Identify system performance bottlenecks based on monitoring data; predict future resource requirements based on historical data and current load conditions; Dynamically adjust the parallelism of data preprocessing and model training tasks based on CPU usage; Adjust the batch size based on GPU utilization; Monitor memory usage and dynamically adjust data cache strategies based on memory usage; Optimize data reading and writing strategies based on storage I / O waiting time; use Docker technology to achieve dynamic allocation and isolation of resources; Collect performance indicators during the training process, and continuously iterate and optimize resource allocation strategies based on performance feedback and resource usage.
7. The efficient GPU computing server resource optimization method according to claim 6, characterized in that: Performance indicators include training speed and accuracy.
8. An efficient GPU computing server resource optimization system, the system implementing the method according to any one of claims 1 to 7, characterized in that: include: The evaluation module is used to conduct a preliminary evaluation of the hardware configuration of the GPU computing server to obtain the evaluation results, including the GPU model, quantity and memory capacity; configure the initial training parameters according to the evaluation results, including the batch size; The dynamic adjustment module is used to dynamically adjust the batch size based on the evaluation results, the complexity of the current training task and the parallel processing capability of the GPU to obtain the adjustment result; The optimization module is used to optimize the performance of DataLoader according to the adjustment results to obtain the optimized data flow; An adjustment module is used to adjust the model architecture according to the optimized data flow to match the parallel processing capability of the GPU. Adjusting the model architecture includes adjusting the number of layers, width, and convolution kernel size of the model to obtain an adjusted model. The allocation module is used to dynamically adjust the allocation of CPU, GPU, memory, and storage resources based on the adjusted model and the real-time load during training so that the components can work together.
9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
GPU computing power pool intelligent management method and system
CN117971475A
Method for evaluating real-time performance of computing power network based on analytic hierarchy process
CN118897952A
GPU cluster data sharing method for AI model training
CN119149209A
Resource dynamic analysis and sample load scheduling optimization method oriented to edge distributed training
CN119149247A
Heterogeneous computing power resource allocation optimization method for deep reinforcement learning model training
CN119271398A
Cited By
Distributed large model reasoning optimization method and device based on dynamic micro-batch scheduling
CN120723311A
A distributed large model inference optimization method and device based on dynamic micro-batch scheduling
CN120723311B
Automatic safety monitoring system for computing power server
CN121434026A