GPU port switching method and device
By collecting and preprocessing real-time performance data of GPU ports, using prediction models to generate load prediction results, and formulating and executing decision strategies, the problem of uneven resource allocation in traditional GPU port management methods is solved, and system performance and resource utilization are improved.
Patent Information
- Application Number
- CN202510362020.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional GPU port allocation and management methods lack dynamic and foresight, and cannot effectively respond to complex and changeable computing needs, resulting in uneven resource allocation, performance bottlenecks and low utilization.
By regularly collecting real-time performance data of each GPU port, using prediction models to generate load prediction results, formulating decision strategies and executing migration tasks, and achieving intelligent scheduling and optimization of GPU ports.
It improves the overall performance and resource utilization of the system, and realizes intelligent scheduling and optimization of GPU port resources.
Smart Images

Figure CN120335950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of port switching, and specifically to a GPU port switching method and device. Background Art
[0002] In the era of rapid digital development today, big data, artificial intelligence, and graphics-intensive applications have sprung up rapidly. In the field of big data, the scale of storage, processing, and analysis of massive data is unprecedentedly large. From the daily hundreds of millions of transaction data on e-commerce platforms to the massive text, pictures, and video information generated by users in social networks, all require efficient computing resources for mining and analysis to obtain valuable business insights and user behavior patterns. In terms of artificial intelligence, the training and inference processes of deep learning models have almost demanding requirements for computing power. In application scenarios such as speech recognition, image recognition, and natural language processing, extremely complex neural network structures need to be processed, with billions or even trillions of floating-point operations per second being carried out. Graphics-intensive applications are no less impressive. Whether it is the realistic special effects rendering in film and television production or the colorful virtual scene presentation in large 3D games, they all rely on powerful graphics processing capabilities to achieve visual shock effects.
[0003] Behind the application of these cutting-edge technologies, the GPU (Graphics Processing Unit), as a key computing resource, has become increasingly prominent. The GPU was initially designed for graphics rendering, but with its powerful parallel computing capabilities, it has demonstrated excellent performance advantages in fields such as big data analysis and artificial intelligence acceleration, becoming the core driving force for the development of these technologies. However, currently, the performance and resource utilization efficiency of the GPU have become the key factors restricting the performance of the entire system.
[0004] Traditional GPU port allocation and management methods have many drawbacks. In the face of complex and changing computing requirements, it often lacks dynamism. For example, when multiple applications are running simultaneously and one artificial intelligence model training task suddenly requires a large amount of GPU resources to accelerate operations, the traditional method is difficult to quickly adjust resource allocation in a short time to meet this demand in a timely manner, resulting in slow task operation and low efficiency. At the same time, it also lacks predictability and cannot anticipate in advance the changes in GPU resource requirements of different applications at different times. For example, at night, big data analysis tasks may break out intensively, with a significant increase in the demand for GPU resources. However, due to the inability to anticipate in advance, the traditional management method still allocates resources according to the conventional mode, making the problem of uneven resource allocation more serious. Some applications may occupy a large amount of GPU resources for a long time but actually have low utilization rates, while some tasks in urgent need of resources do not receive sufficient support, thereby causing performance bottlenecks, and the overall GPU utilization rate is also at a low level, seriously hindering the improvement of system performance and the efficient development of business. Summary of the Invention
[0005] In view of the fact that the traditional GPU port allocation and management methods often lack dynamics and predictability, and are unable to effectively cope with complex and changing computing requirements, resulting in problems such as uneven resource allocation, performance bottlenecks, and low utilization rates, the present invention provides a GPU port switching method and device, which realize intelligent scheduling and optimization of GPU port resources through dynamic load balancing and intelligent prediction technology, and improve the overall performance and resource utilization rate of the system.
[0006] In a first aspect, the present invention provides a GPU port switching method, and the technical solution adopted to solve the above technical problems is as follows:
[0007] A GPU port switching method includes the following steps:
[0008] S1. Regularly collect the real-time performance data of each GPU port, preprocess it, and store it in the database;
[0009] S2. Select a prediction model, train and optimize the prediction model using historical performance data, and use the optimized prediction model to generate the GPU port load prediction results within a future set time period based on the regularly obtained real-time performance data;
[0010] S3. Based on the prediction results and real-time performance data, formulate and output the decision-making strategy for GPU port switching. The decision-making strategy includes the source port that needs to perform GPU port switching and the target port to which it will be switched;
[0011] S4. Generate a migration task list based on the decision-making strategy, and execute the migration tasks in sequence according to the migration task list to realize the switching of GPU ports;
[0012] S5. Monitor the execution results of the migration tasks, and update the configuration information of the GPU ports based on the execution results;
[0013] S6. Perform policy feedback according to the execution results of the migration tasks, adjust the parameters of the prediction model and the decision-making strategy for GPU port switching according to the feedback information. At the same time, store the current performance data and its corresponding prediction results and decision-making strategies in the knowledge base.
[0014] Optionally, the specific steps of step S1 include:
[0015] S1.1. Regularly collect the real-time performance data of each GPU port through the API provided by the GPU driver or a dedicated hardware monitoring tool. The performance data includes CPU occupancy rate, memory bandwidth utilization rate, GPU temperature, error count, and application-specific performance metrics;
[0016] S1.2. Clean and preprocess the collected performance data to remove noise and outliers, ensuring the accuracy and reliability of subsequent analysis;
[0017] S1.3. Store the cleaned and preprocessed performance data in a database or in-memory database for quick access and query of historical data.
[0018] Optionally, the specific steps involved in S2 include:
[0019] S2.1. Select a prediction model according to the application scenario and data characteristics;
[0020] S2.2. Extract features useful for prediction from the collected historical performance data, and perform feature scaling and feature encoding on the extracted features;
[0021] S2.3. Thoroughly clean the historical performance data that has undergone feature scaling and feature encoding to remove existing noise and outliers;
[0022] S2.4. Use the thoroughly cleaned historical performance data to train the prediction model. The prediction model adjusts its own parameters by learning the patterns and rules in the data to adapt to the characteristics of the data;
[0023] S2.5. Adopt a cross-validation method to evaluate the performance of the trained prediction model;
[0024] S2.6. According to the results of the performance evaluation, adjust and optimize the parameters of the prediction model, and output the prediction model with optimized parameters;
[0025] S2.7. Based on the real-time performance data obtained regularly, the prediction model generates prediction results for the GPU port load within a set future time period.
[0026] Optionally, the specific steps involved in S3 include:
[0027] S3.1. Develop a decision-making strategy for GPU port switching based on the current prediction results and the real-time performance data of each GPU port;
[0028] S3.2. During the process of developing the decision-making strategy for GPU port switching, consider the resource conflicts and dependencies that may be caused by port switching, and resolve the conflicts through methods such as priority sorting, task rearrangement, or resource reservation;
[0029] S3.3. Output the developed decision-making strategy for GPU port switching.
[0030] Optionally, execute step S5. After updating the configuration information of the GPU port based on the execution results, perform a consistency check to ensure that all relevant components have correctly recognized and used the new GPU port configuration;
[0031] Meanwhile, record the key information and operation logs during the GPU port switching process for subsequent analysis and troubleshooting.
[0032] In a second aspect, the present invention provides a GPU port switching device, and the technical solution adopted to solve the above technical problems is as follows:
[0033] A GPU port switching device includes:
[0034] An acquisition and processing module, configured to periodically acquire the real-time performance data of each GPU port, perform preprocessing, and store it in a database;
[0035] A model processing module, configured to select a prediction model and use historical performance data to train and optimize the prediction model;
[0036] A prediction model, configured to generate a GPU port load prediction result within a future set time period based on the periodically acquired real-time performance data;
[0037] A policy formulation module, configured to formulate and output a decision-making policy for GPU port switching based on the prediction result and real-time performance data. The decision-making policy includes the source port that needs to perform GPU port switching and the target port to be switched to;
[0038] A decision execution module, configured to generate a migration task list based on the decision-making policy and sequentially execute the migration tasks according to the migration task list to implement the switching of the GPU port;
[0039] A configuration management module, configured to monitor the execution result of the migration task and update the configuration information of the GPU port based on the execution result;
[0040] An adaptive optimization module, configured to perform policy feedback according to the execution result of the migration task and adjust the parameters of the prediction model and the decision-making policy for GPU port switching according to the feedback information;
[0041] A data storage module, configured to store the current performance data and its corresponding prediction result and decision-making policy in a knowledge base.
[0042] Optionally, the involved acquisition and processing module includes:
[0043] A periodic acquisition unit, configured to periodically acquire the real-time performance data of each GPU port through an API provided by the GPU driver or a dedicated hardware monitoring tool. The performance data includes CPU occupancy, memory bandwidth utilization, GPU temperature, error count, and application-specific performance metrics;
[0044] A preprocessing unit, configured to clean and preprocess the acquired performance data, remove noise and outliers, and ensure the accuracy and reliability of subsequent analysis;
[0045] A storage unit for storing the cleaned and preprocessed performance data in a database or in-memory database for quick access and query of historical data.
[0046] Optionally, the involved model processing module includes:
[0047] A model selection unit for selecting a prediction model according to the application scenario and data characteristics;
[0048] A feature processing unit for extracting useful features for prediction from the collected historical performance data, and performing feature scaling and feature encoding on the extracted features;
[0049] A comprehensive cleaning unit for comprehensively cleaning the historical performance data that has undergone feature scaling and feature encoding to remove existing noise and outliers;
[0050] A model training unit for training a prediction model using the comprehensively cleaned historical performance data. The prediction model adjusts its own parameters by learning the patterns and rules in the data to adapt to the characteristics of the data;
[0051] A performance evaluation unit for evaluating the performance of the trained prediction model using the cross-validation method;
[0052] A model optimization unit for adjusting and optimizing the parameters of the prediction model according to the results of the performance evaluation, and outputting the prediction model with optimized parameters;
[0053] A prediction model for generating GPU port load prediction results within a set future time period based on the regularly obtained real-time performance data.
[0054] Optionally, the involved policy formulation module includes:
[0055] A policy formulation unit for formulating a decision-making policy for GPU port switching according to the current prediction results and the real-time performance data of each GPU port;
[0056] A policy adjustment unit for considering the resource conflicts and dependencies that may be caused by port switching during the process of formulating a decision-making policy for GPU port switching, and resolving conflicts through methods such as priority sorting, task rearrangement, or resource reservation;
[0057] A policy output unit for outputting the decision-making policy for GPU port switching.
[0058] Optionally, the involved configuration management module has a built-in consistency check unit for performing a consistency check after the execution result updates the configuration information of the GPU port to ensure that all relevant components have correctly recognized and used the new GPU port configuration;
[0059] The configuration management module is built with an information storage unit, which is used to record key information and operation logs during the GPU port switching process for subsequent analysis and troubleshooting.
[0060] A GPU port switching method and device according to the present invention have the following beneficial effects compared with the prior art:
[0061] The present invention can achieve intelligent scheduling and optimization of GPU port resources through dynamic load balancing and intelligent prediction technology, improving the overall system performance and resource utilization rate; it is applied to high-performance GPU clusters, cloud computing servers, and edge computing devices, and the core components are GPU scheduling controllers and resource management software. Description of the Drawings
[0062] Att Figure 1 is a flowchart of the method according to Embodiment 1 of the present invention;
[0063] Att Figure 2 is a block diagram of module connections according to Embodiment 2 of the present invention. Detailed Embodiments
[0064] To make the technical solutions, technical problems to be solved, and technical effects of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in conjunction with specific embodiments.
[0065] Embodiment 1:
[0066] Referring to Att Figure 1 , this embodiment proposes a GPU port switching method, which includes the following steps:
[0067] S1. Regularly collect real-time performance data of each GPU port, preprocess it, and store it in a database.
[0068] This process specifically includes:
[0069] S1.1. Regularly (such as every hour, every day, etc.) collect real-time performance data of each GPU port through the API provided by the GPU driver or a dedicated hardware monitoring tool. The performance data includes CPU occupancy rate, memory bandwidth utilization rate, GPU temperature, error count, and application-specific performance metrics (such as frame rate, rendering time, etc.).
[0070] S1.2. Clean and preprocess the collected performance data to remove noise and outliers to ensure the accuracy and reliability of subsequent analysis.
[0071] S1.3. Store the cleaned and preprocessed performance data in a database or in-memory database for quick access and query of historical data.
[0072] S2. Select a prediction model, train and optimize the prediction model using historical performance data, and use the optimized prediction model to generate GPU port load prediction results within a future set time period based on regularly obtained real-time performance data.
[0073] This process specifically includes:
[0074] S2.1. Select a prediction model according to the application scenario and data characteristics.
[0075] For example, select a time series prediction model as the prediction model according to the application scenario and data characteristics. Specifically, select: (1) ARIMA (Autoregressive Integrated Moving Average Model), which is suitable for time series data with stationarity characteristics. If the GPU port load data shows a relatively stable fluctuation pattern within a certain period of time, without obvious trends or seasonal variations, the ARIMA model may be a good choice. It makes predictions by analyzing the autocorrelation and moving average of time series data. (2) LSTM (Long Short-Term Memory Network), which is good at dealing with time series data with long-term dependencies. The GPU port load may be affected by the cumulative effects of various factors over a long period in the past, and LSTM can effectively capture these complex time-dependent information. For example, if a large-scale data processing task in the past led to a continuously high GPU load in the subsequent period, the LSTM model can learn this long-distance dependency and use it for prediction. (3) Regression model, which is more applicable when there is a linear or transformable linear relationship between data. For example, if a linear correlation is found between the GPU port load and factors such as the number of tasks and data transfer volume, a linear regression model can be used for prediction. By establishing a linear equation between the load and these influencing factors, the future load can be predicted based on the known factor values.
[0076] For example, select an ensemble learning model as the prediction model according to the application scenario and data characteristics. Specifically, select: (1) Random Forest, which consists of multiple decision trees and improves the accuracy and stability of prediction by integrating the prediction results of numerous decision trees (such as voting method or averaging method). It can handle complex non-linear relationships in data and has good robustness to noise and outliers. For some irregular fluctuations or anomalies that may exist in the GPU port load data, the Random Forest model can give relatively reliable predictions by integrating the judgments of multiple decision trees. (2) Gradient Boosting Tree, which iteratively trains weak learners (such as decision trees) to gradually reduce the prediction error and improve the performance of the overall model. It can gradually optimize for different features and patterns in the data and is suitable for complex data distributions. If the GPU port load data shows complex distribution characteristics, the Gradient Boosting Tree can improve the prediction ability for load changes by continuously fitting the residuals.
[0077] S2.2. Extract useful features for prediction from the collected historical performance data, and perform feature scaling and feature encoding on the extracted features.
[0078] The extracted features may include historical load trends, current load levels, and task type distributions. Among them: ① Historical load trend: By calculating the slope of the load over a period of time, it is judged whether the load is rising, falling, or stable, which helps to understand the long-term change law of GPU load and provides a basis for predicting future loads. ② Current load level: Obtain the load value of the current GPU port directly from the collected real-time data, which reflects the current working state and has a direct impact on predicting the load change in the short term in the future. ③ Task type distribution: Check the types of tasks currently running on the GPU, and count the proportion of different types of tasks. Different types of tasks have quite different demands for GPU resources. For example, graphics rendering tasks and data analysis tasks have different demands for GPU computing resources, video memory, etc. Understanding the task type distribution can more accurately estimate the load of the GPU port.
[0079] Subsequently, perform scaling processing on the extracted numerical features (such as historical load values, current load levels, etc.). Since the value ranges of different features may vary greatly, for example, the historical load value may be between 0 and 100, while the number of tasks may be between 1 and 1000. By scaling (such as standardization or normalization), the values of all features are unified to a relatively fixed range, which helps to improve the efficiency and accuracy of model training and prevent some features from having too much impact on model training due to their overly large value ranges.
[0080] For categorical features (such as task types), encoding processing is required. Because most machine learning models cannot directly process categorical data in text form, common encoding methods include one-hot encoding, etc., to convert categorical features into numerical vectors so that the model can understand and process them.
[0081] S2.3. Thoroughly clean the historical performance data that has undergone feature scaling and feature encoding, and remove existing noise and outliers.
[0082] S2.4. Use the thoroughly cleaned historical performance data to train a prediction model. The prediction model learns the patterns and laws in the data and adjusts its own parameters to adapt to the characteristics of the data. For example, for an LSTM model, the weights and biases in the neural network will be adjusted so that the model can better fit the change trend of historical load data.
[0083] S2.5. Use the cross-validation method to evaluate the performance of the trained prediction model. Cross-validation is to divide the data into multiple subsets, take one of the subsets as the test set and the other subsets as the training set, train and test the model multiple times, and then comprehensively evaluate the performance of the model on different test sets. Common evaluation indicators include mean square error (MSE), mean absolute error (MAE), etc. These indicators can measure the degree of deviation between the model's predicted value and the actual value, so as to judge the accuracy of the model.
[0084] S2.6. According to the results of the performance evaluation, the parameters of the prediction model are adjusted and optimized, and the prediction model with optimized parameters is output. For example, for the random forest model, the number of decision trees, maximum depth and other parameters can be adjusted; by constantly trying different parameter combinations, the parameter configuration that makes the prediction model have the highest prediction accuracy can be found, and the prediction ability of the model can be predicted to be improved.
[0085] S2.7. The prediction model generates GPU port load prediction results within a future set time period (such as the next 10 minutes, 30 minutes or 1 hour) based on the real-time performance data obtained regularly.
[0086] These regularly generated prediction results can help operation and maintenance personnel understand the changing trend of GPU load in advance, arrange resources reasonably, and respond to possible high load situations in advance.
[0087] S3. Based on the prediction results and real-time performance data, a decision strategy for GPU port switching is formulated and outputted. The decision strategy includes a source port that needs to be switched to the GPU port and a target port to be switched to.
[0088] This process specifically includes:
[0089] S3.1. Develop a decision strategy for GPU port switching based on the current prediction results and the real-time performance data of each GPU port. Specifically, based on the current actual load level of each GPU port, the type and number of tasks being run, the available resources of the port, and other information, consider using simple rules based on thresholds (such as triggering switching when the load of a port exceeds a set threshold) or complex optimization algorithms based on cost-benefit analysis to develop a decision strategy for GPU port switching.
[0090] Simple rule based on thresholds: One or more load thresholds are set in advance. Taking a single threshold as an example, assume that the GPU port load threshold is set at 80%. When the load of a certain port is monitored in real time and reaches or exceeds this 80% threshold, the port switching mechanism is immediately triggered. This rule is intuitive and easy to understand, with high execution efficiency, and is more applicable in scenarios with high real-time requirements, relatively stable and regular system load changes. For example, in some graphics processing tasks in daily office environments, the task types and load requirements are relatively fixed. Using this simple threshold rule can quickly and effectively balance the GPU port load.
[0091] Complex optimization algorithm based on cost-benefit analysis: Multiple factors need to be comprehensively considered. Such as: in terms of cost, including the time loss caused by task interruption during port switching, the bandwidth resource cost consumed by data transmission between different ports, etc.; in terms of benefits, such as the improvement of task processing efficiency brought about by the overall performance improvement of the system after switching, the time cost saved by reducing task queuing waiting time, etc. By constructing a complex mathematical model and using algorithms such as linear programming and dynamic programming, among many possible port switching schemes, an optimal solution is found. This optimal solution can minimize costs and maximize benefits on the premise of meeting various system constraints (such as total resource limitations, task priority requirements, etc.). This strategy is applicable to scenarios with extremely high requirements for resource utilization efficiency, diverse task types and complex interrelationships, such as large-scale data centers and complex scientific research computing clusters.
[0092] S3.2. During the process of formulating the decision-making strategy for GPU port switching, consider the resource conflicts and dependencies that may be caused by port switching, and resolve the conflicts through methods such as priority sorting, task rearrangement, or resource reservation.
[0093] Possible resource conflicts and dependencies are as follows: There are priority differences in the GPU resource requirements of different tasks. Some high-priority tasks may not be interrupted or migrated at will; there are data dependencies between some tasks, and the output of one task is the input of another task. If not handled properly during port switching, it may lead to data transmission delays or errors, affecting the normal execution of tasks; in addition, shared resources in the system (such as memory, network bandwidth, etc.) may also have competition conflicts during port switching.
[0094] Possible methods for this may include: ① Priority sorting: Tasks in the system are prioritized according to importance or urgency. When performing port switching, the resource requirements and running continuity of high-priority tasks are guaranteed first. For example, in a system with both real-time video rendering and ordinary data processing tasks, the real-time video rendering task has extremely high requirements for timeliness and should have a higher priority than the ordinary data processing task. When the port load is unbalanced and needs to be switched, it is first ensured that the real-time video rendering task is not affected by the port switching, or it is migrated to a more suitable port to ensure its smooth operation. ② Task rearrangement: According to the dependency relationship and resource requirements between tasks, the execution order and running ports of tasks are rearranged. For example, if task A depends on the output result of task B, and task B is currently on a port with high load and needs to be switched, then before switching task B, it is necessary to first evaluate the execution progress and status of task A. If task A has not started execution, after switching task B to a suitable port, task A can be started to ensure the correct transfer of data dependencies. Through reasonable task rearrangement, task execution errors caused by port switching can be effectively avoided. ③ Resource reservation: To ensure the smooth execution of critical tasks or tasks about to perform port switching, a certain amount of system resources are reserved in advance. For example, before expecting to perform a switching operation on a certain port, sufficient GPU resources such as video memory and computing cores, as well as relevant system resources such as memory and network bandwidth, are reserved in advance for the tasks to be migrated to this port. This can prevent tasks from not being able to start normally or running slowly due to insufficient resources after port switching.
[0095] S3.3. Output the formulated decision strategy for GPU port switching. The decision strategy includes the source port that needs to perform GPU port switching (i.e., the GPU port that needs to be switched due to high load or other problems), and the target port to be switched to (i.e., the GPU port that is identified as having low load, sufficient resources, and being able to meet the requirements of the migration task).
[0096] S4. Generate a migration task list based on the decision strategy, and execute the migration tasks in sequence according to the migration task list to achieve the switching of GPU ports.
[0097] The migration task list details the task list that needs to be migrated from the source port to the target port, including multiple migration tasks. Each migration task has its specific attributes, such as task type, data volume size, execution progress, etc. By generating the migration task list, tasks can be migrated in sequence to ensure the orderliness and accuracy of the migration process. For example, the migration task list includes task A (graphics rendering task, 30% completed), task B (data analysis task, not started yet), etc.
[0098] It should be added that after the migration task is completed, the resources and data that are no longer needed on the source port are cleared, and the occupied port resources are released.
[0099] S5. Monitor the execution result of the migration task, and update the configuration information of the GPU port based on the execution result.
[0100] After updating the configuration information of the GPU port based on the execution result, perform a consistency check to ensure that all relevant components have correctly recognized and used the new GPU port configuration;
[0101] Meanwhile, record the key information and operation logs during the GPU port switching process for subsequent analysis and troubleshooting.
[0102] S6. Perform policy feedback according to the execution result of the migration task, adjust the parameters of the prediction model and the decision-making strategy for GPU port switching according to the feedback information. Meanwhile, store the current performance data, its corresponding prediction results and decision-making strategies (including the decision-making strategy output in step S3 and the decision-making strategy adjusted in this step) in the knowledge base.
[0103] Embodiment 2:
[0104] Refer to the appendix Figure 2 , this embodiment proposes a GPU port switching device, which includes:
[0105] An acquisition and processing module, which is used to regularly acquire the real-time performance data of each GPU port, perform preprocessing and store it in the database;
[0106] A model processing module, which is used to select a prediction model and train and optimize the prediction model using historical performance data;
[0107] A prediction model, which is used to generate a prediction result of the GPU port load within a future set time period based on the regularly acquired real-time performance data;
[0108] A policy formulation module, which is used to formulate and output a decision-making strategy for GPU port switching based on the prediction result and real-time performance data. The decision-making strategy includes the source port that needs to perform GPU port switching and the target port to which it will be switched;
[0109] A decision execution module, which is used to generate a migration task list based on the decision-making strategy and sequentially execute the migration tasks according to the migration task list to achieve the switching of the GPU port;
[0110] A configuration management module, which is used to monitor the execution result of the migration task and update the configuration information of the GPU port based on the execution result;
[0111] An adaptive optimization module for providing policy feedback based on the execution results of the migration task and adjusting the parameters of the prediction model and the decision-making strategy for GPU port switching according to the feedback information;
[0112] A data storage module for storing the current performance data, its corresponding prediction results, and decision-making strategies in the knowledge base.
[0113] In this embodiment, the acquisition and processing module includes:
[0114] A regular acquisition unit for regularly acquiring the real-time performance data of each GPU port through the API provided by the GPU driver or a dedicated hardware monitoring tool, where the performance data includes CPU occupancy, memory bandwidth utilization, GPU temperature, error count, and application-specific performance metrics;
[0115] A preprocessing unit for cleaning and preprocessing the acquired performance data to remove noise and outliers and ensure the accuracy and reliability of subsequent analysis;
[0116] A storage unit for storing the cleaned and preprocessed performance data in a database or in-memory database for quick access and query of historical data.
[0117] In this embodiment, the model processing module includes:
[0118] A model selection unit for selecting a prediction model according to the application scenario and data characteristics;
[0119] A feature processing unit for extracting useful features for prediction from the acquired historical performance data, and performing feature scaling and feature encoding on the extracted features;
[0120] A comprehensive cleaning unit for comprehensively cleaning the historical performance data that has undergone feature scaling and feature encoding to remove existing noise and outliers;
[0121] A model training unit for training the prediction model using the comprehensively cleaned historical performance data. The prediction model adjusts its own parameters by learning the patterns and rules in the data to adapt to the characteristics of the data;
[0122] A performance evaluation unit for evaluating the performance of the trained prediction model using the cross-validation method;
[0123] A model optimization unit for adjusting and optimizing the parameters of the prediction model according to the results of the performance evaluation and outputting the prediction model with optimized parameters;
[0124] A prediction model for generating a prediction result of the GPU port load within a future set time period based on the regularly obtained real-time performance data.
[0125] In this embodiment, the involved policy-making module includes:
[0126] A policy-making unit, configured to formulate a decision-making policy for GPU port switching according to the current prediction result and the real-time performance data of each port of the GPU;
[0127] A policy adjustment unit, configured to consider the resource conflicts and dependencies that may be caused by port switching during the process of formulating the decision-making policy for GPU port switching, and resolve the conflicts through methods such as priority sorting, task rearrangement, or resource reservation;
[0128] A policy output unit, configured to output the decision-making policy for GPU port switching.
[0129] In this embodiment, the involved configuration management module is built-in with a consistency check unit, which is configured to perform a consistency check after the execution result updates the configuration information of the GPU port, so as to ensure that all relevant components have correctly recognized and used the new GPU port configuration. The configuration management module is built-in with an information storage unit, which is configured to record the key information and operation logs during the GPU port switching process for subsequent analysis and troubleshooting.
[0130] In summary, by adopting a method and device for GPU port switching according to the present invention, intelligent scheduling and optimization of GPU port resources can be achieved through dynamic load balancing and intelligent prediction technologies, improving the overall performance and resource utilization rate of the system.
[0131] The above specific application examples have elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art of this technology without departing from the principle of the present invention shall fall within the scope of patent protection of the present invention.
Claims
1. A GPU port switching method, characterized in that, It includes the following steps: S1. Regularly collect the real-time performance data of each GPU port, preprocess it, and store it in the database; S2. Select a prediction model, train and optimize the prediction model using historical performance data, and use the optimized prediction model to generate the GPU port load prediction results within a future set time period based on the regularly obtained real-time performance data; S3. Based on the prediction results and real-time performance data, formulate and output the decision-making strategy for GPU port switching. The decision-making strategy includes the source port that needs to perform GPU port switching and the target port to which it will be switched; S4. Generate a migration task list based on the decision-making strategy, and sequentially execute the migration tasks according to the migration task list to achieve GPU port switching; S5. Monitor the execution results of the migration tasks and update the configuration information of the GPU ports based on the execution results; S6. Conduct policy feedback based on the execution results of the migration tasks, adjust the parameters of the prediction model and the decision-making strategy for GPU port switching according to the feedback information. At the same time, store the current performance data, its corresponding prediction results, and decision-making strategies in the knowledge base.
2. The GPU port switching method according to claim 1, wherein The specific steps of step S1 include: S1.
1. Regularly collect the real-time performance data of each GPU port through the API provided by the GPU driver or a dedicated hardware monitoring tool. The performance data includes CPU occupancy rate, memory bandwidth utilization rate, GPU temperature, error count, and application-specific performance metrics; S1.
2. Clean and preprocess the collected performance data to remove noise and outliers to ensure the accuracy and reliability of subsequent analysis; S1.
3. Store the cleaned and preprocessed performance data in a database or in-memory database for quick access and query of historical data.
3. A GPU port switching method according to claim 1, characterized in that, The specific steps of step S2 include: S2.
1. Select a prediction model according to the application scenario and data characteristics; S2.
2. Extract the features useful for prediction from the collected historical performance data, and perform feature scaling and feature encoding on the extracted features; S2.
3. Thoroughly clean the historical performance data after feature scaling and feature encoding to remove existing noise and outliers; S2.
4. Use the thoroughly cleaned historical performance data to train the prediction model. The prediction model adjusts its own parameters by learning the patterns and rules in the data to adapt to the characteristics of the data; S2.
5. Adopt the cross-validation method to evaluate the performance of the trained prediction model; S2.
6. Adjust and optimize the parameters of the prediction model according to the results of the performance evaluation, and output the prediction model with optimized parameters; S2.
7. The prediction model generates the GPU port load prediction results within a future set time period based on the regularly obtained real-time performance data.
4. A GPU port switching method according to claim 1, wherein The specific steps of step S3 include: S3.
1. Formulate the decision-making strategy for GPU port switching according to the current prediction results and the real-time performance data of each GPU port; S3.
2. During the process of formulating the decision-making strategy for GPU port switching, consider the resource conflicts and dependencies that may be caused by port switching, and resolve the conflicts through methods such as priority sorting, task rearrangement, or resource reservation; S3.
3. Output the formulated decision-making strategy for GPU port switching.
5. A GPU port switching method according to claim 1, characterized in that, Execute step S5. After updating the configuration information of the GPU port based on the execution result, perform a consistency check to ensure that all relevant components have correctly recognized and used the new GPU port configuration; Meanwhile, record the key information and operation logs during the GPU port switching process for subsequent analysis and troubleshooting.
6. A GPU port switching device, characterized in that, It includes: The acquisition and processing module is used to regularly acquire the real-time performance data of each GPU port, preprocess it, and store it in the database; The model processing module is used to select a prediction model and train and optimize the prediction model using historical performance data; The prediction model is used to generate the GPU port load prediction result within a set future time period based on the regularly obtained real-time performance data; The policy formulation module is used to formulate and output the decision-making policy for GPU port switching based on the prediction result and real-time performance data. The decision-making policy includes the source port that needs to perform GPU port switching and the target port to be switched to; The decision execution module is used to generate a migration task list based on the decision-making policy and execute the migration tasks in sequence according to the migration task list to achieve the switching of the GPU port; The configuration management module is used to monitor the execution result of the migration task and update the configuration information of the GPU port based on the execution result; The adaptive optimization module is used to perform policy feedback according to the execution result of the migration task and adjust the parameters of the prediction model and the decision-making policy for GPU port switching according to the feedback information; The data storage module is used to store the current performance data and its corresponding prediction result and decision-making policy in the knowledge base.
7. A GPU port switching method according to claim 6, wherein The acquisition and processing module includes: The regular acquisition unit is used to regularly acquire the real-time performance data of each GPU port through the API provided by the GPU driver or a dedicated hardware monitoring tool. The performance data includes CPU occupancy rate, memory bandwidth utilization rate, GPU temperature, error count, and application-specific performance metrics; The preprocessing unit is used to clean and preprocess the acquired performance data, remove noise and outliers, and ensure the accuracy and reliability of subsequent analysis; The storage unit is used to store the cleaned and preprocessed performance data in the database or in-memory database for quick access and query of historical data.
8. A GPU port switching method according to claim 6, wherein The model processing module includes: The model selection unit is used to select a prediction model according to the application scenario and data characteristics; The feature processing unit is used to extract useful features for prediction from the acquired historical performance data, and perform feature scaling and feature encoding on the extracted features; The comprehensive cleaning unit is used to comprehensively clean the historical performance data that has undergone feature scaling and feature encoding, and remove the existing noise and outliers; The model training unit is used to train the prediction model using the comprehensively cleaned historical performance data. The prediction model adjusts its own parameters by learning the patterns and rules in the data to adapt to the characteristics of the data; The performance evaluation unit is used to evaluate the performance of the trained prediction model using the cross-validation method; The model optimization unit is used to adjust and optimize the parameters of the prediction model according to the result of the performance evaluation, and output the prediction model with optimized parameters; A prediction model for generating GPU port load prediction results within a set future time period based on real-time performance data obtained regularly.
9. The GPU port switching method according to claim 6, wherein The policy-making module includes: A policy-making unit for formulating a decision-making policy for GPU port switching according to the current prediction results and the real-time performance data of each GPU port; A policy adjustment unit for considering resource conflicts and dependencies that may be caused by port switching during the process of formulating a decision-making policy for GPU port switching, and resolving conflicts through methods such as priority sorting, task rearrangement, or resource reservation; A policy output unit for outputting the decision-making policy for GPU port switching.
10. A GPU port switching method according to claim 6, characterized in that, The configuration management module is built with a consistency check unit for performing a consistency check after the execution result updates the configuration information of the GPU port to ensure that all relevant components have correctly recognized and used the new GPU port configuration; The configuration management module is built with an information storage unit for recording key information and operation logs during the GPU port switching process for subsequent analysis and troubleshooting.