Method and System for Optimizing and Improving Graphics Card Performance Based on Big Data
The method optimizes GPU performance by using big data analytics to predict and manage resource demands, addressing dynamic changes in application behavior and improving task execution efficiency.
Patent Information
- Application Number
- CN202510519447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the prior art, the allocation of graphics card resources lacks adaptability to dynamic changes in application behavior, resulting in problems such as rendering pipeline waiting, task allocation failure, task cascade timeout or frame loss.
Through big data analysis of graphics card performance, build a GPU activity mode classification index, generate a timing profile of typical resource requirements, dynamically adjust resource allocation strategies, combine task preemption priority evaluation table, set preemption trigger conditions, and optimize graphics card performance.
Adaptive adjustment of graphics card resource allocation is realized, reducing the risk of video memory bandwidth competition, avoiding compilation delays and video memory fragmentation, ensuring that high-priority tasks seize resources at the right time, and reducing rendering frame rate drop or computing pipeline blockage.
Smart Images

Figure CN120045335B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of program control, and particularly to a method and system for optimizing and improving the performance of a graphics card based on big data. Background Art
[0002] The technical field of program control focuses on the automated management and dynamic regulation of the operation process of computing devices through algorithms, protocols, and system-level policies. Its core goal is to optimize hardware performance and task execution efficiency through real-time monitoring, feedback decision-making, and resource scheduling.
[0003] Traditional program control technology relies on preset fixed rules or empirical thresholds for resource allocation and task scheduling, lacking adaptability to the dynamic changes of application behavior. For example, when deploying deep learning training and real-time rendering tasks in a hybrid manner, the fixed priority scheduling mode may cause the rendering pipeline to wait due to unallocated texture mapping units, or fail to allocate real-time tasks in time due to the failure to recover the video memory of training tasks in a timely manner. Moreover, when a high-priority task is inserted suddenly, it is prone to excessive preemption or response lag, resulting in cascading task timeouts or frame losses. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a method and system for optimizing and improving the performance of a graphics card based on big data.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. A method for optimizing and improving the performance of a graphics card based on big data includes the following steps:
[0006] Collect the shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application program, perform data cleaning, and establish a feature set of application program startup and multi-task GPU activities;
[0007] Based on the feature set of application program startup and multi-task GPU activities, group the application program startup behavior and task running status through numerical clustering to obtain the category labels to which each application or task belongs, establish a classification index of GPU activity patterns. Based on the classification index of GPU activity patterns, aggregate the corresponding feature set data for each category label, calculate the average value of the peak demand of GPU compilation units within the category and the video memory bandwidth, and construct a typical GPU resource demand time series profile;
[0008] According to the received application startup request, a typical GPU resource requirement timing profile of the matching application is retrieved, a shader identifier list of the startup call and a predicted initial video memory allocation mode parameter value are extracted, a target application startup resource requirement list is generated, the shader identifier in the target application startup resource requirement list is converted into a pre-compiled instruction, the video memory allocation mode parameter is converted into a video memory pre-allocation instruction, and a GPU pre-processing instruction sequence is established;
[0009] Monitor the task set currently running on the GPU, retrieve the typical GPU resource demand timing profile and the associated context switching overhead corresponding to each task, generate a task preemption priority and cost evaluation table, set the preemption trigger conditions and time points based on the task preemption priority and cost evaluation table, and obtain a dynamic GPU task preemption control parameter set.
[0010] Preferably, the steps of obtaining the application startup and multi-task GPU activity feature set are:
[0011] Collect shader compile time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, calculate the 25th percentile Q1 and 75th percentile Q3 of the shader compile time data, remove extreme values lower than Q1-1.5×(Q3-Q1) or higher than Q3+1.5×(Q3-Q1), and generate cleaned shader compile time data sets, initial video memory allocation mode parameter sets, and pipeline state setting sequence log sets;
[0012] Based on the cleaned shader compilation time data set, traverse the compilation durations recorded at all time points, select the maximum value as the compilation peak duration, parse the video memory block size field from the initial video memory allocation mode parameter set, extract the fixed allocation unit value of the video memory block size, count the number of state switching instructions per second from the pipeline state setting sequence log set as the state switching frequency, calculate the average of the absolute values of the differences between adjacent state switching timestamps as the switching delay, and count the total number of shader instruction executions per second as the computational intensity;
[0013] The compilation peak duration, video memory block size, state switching frequency, computational intensity, and switching delay are aligned by timestamp and merged into a multidimensional vector at the same time point. The time series is divided into time windows of 30 seconds to establish a feature set of application startup and multi-tasking GPU activities.
[0014] Preferably, the steps of obtaining the GPU activity mode classification index are:
[0015] Based on the application startup and multi-task GPU activity feature set, three feature dimensions, namely the compilation peak duration, video memory block size, and state transition frequency, are selected as the clustering inputs. Z-score normalization is performed on each data point to generate a normalized feature subset.
[0016] According to the normalized feature subset, k clustering centers are initialized, where the value of k is determined by the elbow method. The Euclidean distance from each data point to each clustering center is calculated, and the data points are assigned to the nearest neighbor clustering center. The coordinates of the clustering centers are iteratively updated to the mean of the data points they belong to until the change in the center coordinates between two consecutive iterations is less than the threshold, generating an optimized set of clustering centers.
[0017] Based on the optimized set of clustering centers, a cluster label is assigned to the time window of each application or task, a mapping table is established, and the mean of the peak demand of the same type of compilation units and the temporal pattern of the video memory bandwidth are aggregated by label to generate a GPU activity pattern classification index.
[0018] Preferably, the steps for obtaining the typical GPU resource demand time series profile are as follows:
[0019] Based on the GPU activity pattern classification index, all time windows under each category label are traversed, the peak demand of the GPU compilation units within each time window is extracted, and the mean of the peak demand of the compilation units in all time windows is calculated to generate a set of means of the peak demand of the compilation units within the category.
[0020] According to the set of means of the peak demand of the compilation units within the category, the temporal data of the video memory bandwidth corresponding to the category label is extracted, and the fluctuation intensity of the video memory bandwidth is calculated. The formula is:
[0021] ;
[0022] Among them, is the fluctuation intensity of the video memory bandwidth, is the video memory bandwidth value at time point is the bandwidth mean, is the total number of seconds in the time window; is the total number of seconds in the time window;
[0023] Based on the fluctuation intensity of the video memory bandwidth, the call timestamps of the texture mapping unit and the rasterization unit are parsed from the pipeline state setting sequence log set, the call frequency of each unit within the time window is counted, the median and variance of the call frequency are calculated, and the mean of the peak demand of the compilation units, the fluctuation intensity of the video memory bandwidth, the median and variance of the call frequency are concatenated in the order of the time window to form a multi-dimensional vector, constructing a typical GPU resource demand time series profile.
[0024] Preferably, the steps for obtaining the resource demand list for the target application startup are as follows:
[0025] Receive an application startup request, parse the application unique identifier and version number in the request header, traverse all category tags in the GPU activity pattern classification index, compare the application identifier with the features under each category tag one by one, screen out the category tag with the highest feature matching degree, and extract the typical GPU resource demand timing profile corresponding to the category tag;
[0026] Based on the typical GPU resource demand timing profile data block, extract all shader identifiers, video memory block sizes, and virtual address offset base addresses from the typical GPU resource demand timing profile data block to generate an intermediate data set;
[0027] According to the video memory block sizes in the intermediate data set, sort the shader identifiers in descending order, and encapsulate the sorted shader identifiers, corresponding video memory block sizes, and virtual address offset base addresses into a target application startup resource demand list.
[0028] Preferably, the step of obtaining the GPU preprocessing instruction sequence is as follows:
[0029] Parse the target application startup resource demand list, traverse the key-value pairs of each entry in the target application startup resource demand list, and extract the shader identifier, video memory block size, and virtual address offset base address;
[0030] Based on the shader identifier, replace each shader identifier with an instruction string according to the precompiled instruction template, and based on the video memory block size and virtual address offset base address, replace each video memory block size and virtual address offset base address with an instruction string according to the video memory preallocation instruction template to generate a GPU preprocessing instruction sequence.
[0031] Preferably, the step of obtaining the task preemption priority and cost evaluation table is as follows:
[0032] Monitor the set of tasks currently running on the GPU, obtain the process identifier of each task and the associated typical GPU resource demand timing profile through the operating system kernel interface, extract the time consumption of the last three context switches of each task from the system scheduler log, calculate the average value of the time consumption of the last three context switches as the context switch overhead, and generate a data set;
[0033] Based on the data set, calculate the task preemption priority score, and the formula is:
[0034] ;
[0035] Among them, is the task preemption priority score, is the task the The time consumption of the next context switch is the video memory bandwidth occupancy rate of the task ; is the foreground window activation status, where activation is 1 and non-activation is 0;
[0036] According to the task preemption priority score, all tasks are sorted from high to low, and an entry is generated for each task. The entry includes the process identifier, video memory bandwidth occupancy rate, average switching time consumption, activation status, and task preemption priority score, to construct a task preemption priority and cost evaluation table.
[0037] Preferably, the obtaining step of the dynamic GPU task preemption control parameter set is as follows:
[0038] Based on the task preemption priority and cost evaluation table, tasks with task preemption priority scores in the top 20% are selected as the high-priority task set, the average video memory bandwidth occupancy rate and response latency requirements of the high-priority task set are extracted, and at the same time, the 30% tasks with the highest average switching time consumption among the remaining tasks are extracted as the low-priority task set, to generate a high-priority response requirement data set and a low-priority switching cost data set;
[0039] According to the high-priority response requirement data set and the low-priority switching cost data set, calculate the preemption trigger threshold, and the formula is:
[0040] ;
[0041] wherein, is the preemption trigger threshold, is the average video memory bandwidth occupancy rate of high-priority tasks, is the maximum allowable response latency of high-priority tasks, is the average switching time consumption of low-priority tasks, is the current available video memory ratio of the GPU;
[0042] Based on the preemption trigger threshold, set the trigger condition as follows: when the predicted response latency of high-priority tasks exceeds the maximum allowable response latency of high-priority tasks, and the preemption trigger threshold is greater than 100, calculate the preemption time point as the start time of the next vertical blanking period , and the formula is , is the current frame number, to obtain the preemption control parameter set.
[0043] The present invention provides a graphics card performance optimization and improvement system, including:
[0044] A data acquisition module, which records shader compilation time data in the application startup phase, measures video memory allocation mode parameters, traces the pipeline state setting sequence, eliminates outliers, aligns multi-source data according to time stamps, and constructs a graphics card behavior feature set;
[0045] The feature clustering module extracts the timing feature vectors from the graphics card behavior feature set, calculates the distances between the vectors, sets the clustering centers, groups the application behaviors, assigns category identifiers, establishes a mapping relation table, statistically analyzes the time distribution, extracts the GPU compilation unit requirements, calculates the bandwidth utilization rate, performs timing decomposition, forms a resource change curve, marks the peaks and valleys, and stores them as the GPU resource demand profile library;
[0046] The resource demand analysis module queries the GPU resource demand profile library, matches the application categories, extracts the resource profile data, analyzes the shader call sequence, extracts the identifier list, parses the memory allocation peak value, deduces the pre-allocated space, calculates the bandwidth threshold, marks the resource competition points, and outputs the application startup resource planning table;
[0047] The preprocessing instruction generation module reads the shader identifiers in the application startup resource planning table, queries the shader library, verifies the compilation status, filters the pre-compiled subsets, constructs the instruction sequence, parses the video memory allocation parameters, converts them into reserved instructions, sets the block size, formulates the memory scheme, combines the priority sorting, and packages them as the startup optimization instruction sequence;
[0048] The task scheduling control module scans the GPU task set, identifies the resource occupancy, queries the GPU resource demand profile library to obtain the prediction, collects the switching records, calculates the switching time consumption, compares the task priorities, calculates the preemption cost, sets the competition threshold, determines the preemption timing, and forms a dynamic resource scheduling scheme.
[0049] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0050] By collecting shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting logs during the startup phase of an application and constructing a multi-task GPU activity feature set, the present invention can quantify the dynamic demands of different application startup behaviors on hardware resources and achieve accurate prediction of resource occupancy patterns. Based on numerical clustering, the task running states are grouped and a classification index is established. After aggregating the characteristic data of similar tasks, a typical GPU resource demand time series profile is generated, enabling subsequent resource pre-allocation strategies to be adaptively adjusted according to historical behavior patterns and reducing the risk of video memory bandwidth contention in high-fluctuation scenarios. According to real-time startup requests, the matching time series profiles are retrieved, shader identifiers and video memory allocation parameters are extracted to generate pre-compilation instructions and a video memory pre-allocation instruction sequence, and the static allocation of key resources is completed in advance to avoid compilation delays and video memory fragmentation caused by dynamic allocation during runtime. Combining the task preemption priority and cost evaluation table, by setting dynamic trigger conditions and vertical blanking period time points, it is ensured that high-priority tasks can preempt resources at the time when the video memory release is sufficient and the context switching cost is the lowest, reducing the sudden drop in rendering frame rate or computational pipeline blockage caused by preemption. Through the collaborative optimization of data-driven feature modeling and preemptive scheduling, the traditional static resource allocation mode is upgraded to a dynamic regulation mechanism based on historical load prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Please refer to Figure 1 , the present invention provides a technical solution for optimizing and improving the performance of a graphics card based on big data, including the following steps:
[0054] Collect shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, perform data cleaning, and establish an application startup and multi-task GPU activity feature set;
[0055] Based on the application startup and multi-task GPU activity feature set, group the application startup behaviors and task running states through numerical clustering, obtain the category labels to which each application or task belongs, establish a GPU activity mode classification index, and based on the GPU activity mode classification index, aggregate the corresponding feature set data for each category label, calculate the average value of the peak demand of GPU compilation units within the category and the video memory bandwidth, and construct a typical GPU resource demand time series profile;
[0056] According to the received application startup request, retrieve the typical GPU resource demand time series profile of the matching application, extract the list of shader identifiers for the startup call and the predicted initial video memory allocation mode parameter values, generate a target application startup resource demand list, convert the shader identifiers in the target application startup resource demand list into pre-compiled instructions, convert the video memory allocation mode parameters into video memory pre-allocation instructions, and establish a GPU preprocessing instruction sequence;
[0057] Monitor the set of tasks running on the current GPU, retrieve the typical GPU resource demand time series profile and associated context switching overhead corresponding to each task, generate a task preemption priority and cost evaluation table, and based on the task preemption priority and cost evaluation table, set the preemption trigger conditions and time points to obtain a dynamic GPU task preemption control parameter set.
[0058] The steps for obtaining the application startup and multi-task GPU activity feature set are as follows:
[0059] Collect the shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, calculate the 25th percentile Q1 and 75th percentile Q3 of the shader compilation time data, eliminate the extreme values below Q1 - 1.5×(Q3 - Q1) or above Q3 + 1.5×(Q3 - Q1), and generate a cleaned shader compilation time data set, initial video memory allocation mode parameter set, and pipeline state setting sequence log set;
[0060] Based on the cleaned shader compilation time data set, traverse the compilation durations recorded at all time points, select the maximum value as the compilation peak duration, parse the video memory block size field from the initial video memory allocation mode parameter set, extract the fixed allocation unit value of the video memory block size, count the number of state switching instructions per second from the pipeline state setting sequence log set as the state switching frequency, calculate the average value of the absolute differences of adjacent state switching timestamps as the switching delay, and count the total number of shader instructions executed per second as the computational intensity;
[0061] Align the compilation peak duration, video memory block size, state switching frequency, computational intensity, and switching delay by timestamp, merge them into a multi-dimensional vector at the same time point, divide the time series with a continuous 30-second time window, and establish an application startup and multi-task GPU activity feature set.
[0062] Specifically, based on the collected target application, such as the software named "Graphics Editor v2.1", the shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase are processed. First, the shader compilation time data is processed. For example, the sequence of 15 startup compilation times (unit: milliseconds) collected is [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60, 150]. This sequence is sorted to obtain [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60, 150]. The 25th percentile Q1 is calculated, and its position is , and the corresponding value is 30 ms. The 75th percentile Q3 is calculated, and its position is , and the corresponding value is 50 ms. The interquartile range IQR = Q3 - Q1 = 50 - 30 = 20 ms. The lower bound for outlier judgment is determined as Q1 - 1.5 IQR = 30 - 1.5 20 = 0 ms, and the upper bound is Q3 + 1.5 IQR = 50 + 1.5 20 = 80 ms. Values in the original data that are lower than 0 ms or higher than 80 ms are regarded as extreme values. In this example, 150 ms is identified as an extreme value and excluded. A similar data cleaning process is performed on the initial video memory allocation mode parameter set and the pipeline state setting sequence log set. For example, if there are obviously unreasonable large or zero value allocation requests in the video memory allocation records, or if there are timestamp errors or instruction out-of-order in the state setting logs, they are also excluded according to statistical rules or logical rules. Finally, a cleaned data set is obtained, generating a cleaned shader compilation time data set [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60], the corresponding initial video memory allocation mode parameter set, and the pipeline state setting sequence log set.
[0063] Using the generated cleaned shader compilation time data set [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60] (unit: ms), traverse and find the maximum value in it to determine the compilation peak duration is 60 ms. Then, parse the recorded video memory block size field from the cleaned initial video memory allocation mode parameter set. For example, the record is AllocInfo(BlockID = 001, Size = 16777216, Type = Texture). Extract the Size value. It is observed that there are fixed-size units in multiple allocation requests, such as repeated 16 MB allocations. Select this common and representative value as the fixed allocation unit value for the video memory block size , that is Bytes. Then, analyze the pipeline state setting sequence log set after cleaning, and count the number of pipeline state switching instructions (such as SetPipelineStateObject, IASetPrimitiveTopology, etc.) that occur within a specific time window, for example, within the 5th second. If the statistical result is 85 times, then the state switching frequency for this second is 85Hz. Further calculate the absolute value of the timestamp difference between adjacent state switching instructions within this second. For example, if the recorded timestamp sequence is [5.012s, 5.020s, 5.025s,..., 5.980s], calculate the adjacent differences |5.020 - 5.012| = 0.008s, |5.025 - 5.020| = 0.005s,... After summing up all the differences and dividing by the number of differences (i.e., the number of switches minus 1), the switching delay is obtained , for example, the calculation result is 0.011s (11ms). Finally, count the total number of shader instructions executed within the same time window (the 5th second), for example, obtained through GPU performance counters or driver logs. If the total number is 2.5 GigaInstructions, then the computational intensity is 2.5 GigaInstructions / s. Align these eigenvalue features obtained or associated at the 5th second according to the timestamp, and compile the peak duration Using the globally calculated 60ms, and using the values at this time point for others, merge them into a multi-dimensional vector V(t = 5s) = [60ms, 16777216Bytes, 85Hz, 2.5GI / s, 11ms]. Perform this operation for each second during the application startup phase (for example, lasting 60 seconds), generating 60 such multi-dimensional vectors. Then, divide them into time windows of 30 seconds continuously. For example, window 1 contains the vector sequence from the 1st to the 30th second, window 2 contains the vector sequence from the 2nd to the 31st second (using a sliding window), or window 1 contains 1 - 30 seconds, window 2 contains 31 - 60 seconds (non-overlapping window). Here, a 30-second non-overlapping window is used, then two time series segments are generated, and these time series segments together constitute the application startup and multi-task GPU activity feature set.
[0064] The steps to obtain the GPU activity mode classification index are as follows:
[0065] Based on the application startup and multi-task GPU activity feature set, select three feature dimensions: compile peak duration, video memory block size, and state switching frequency as the clustering input, perform Z-score normalization on each data point, and generate a normalized feature subset;
[0066] Initialize k cluster centers according to the standardized feature subset. The value of k is determined by the elbow method. Calculate the Euclidean distance from each data point to each cluster center, and assign the data point to the nearest cluster center. Iteratively update the cluster center coordinates to the mean of the data points belonging to it until the change in the center coordinates between two consecutive iterations is less than the threshold, generating an optimized set of cluster centers;
[0067] Based on the optimized set of cluster centers, assign the belonging cluster label to the time window of each application or task, establish a mapping table, aggregate the mean peak demand of the same category of compilation units and the temporal pattern of video memory bandwidth according to the label, and generate a GPU activity pattern classification index.
[0068] Specifically, based on the established application startup and multi-task GPU activity feature set, which contains data of multiple time windows, and each window is a multi-dimensional vector sequence. For example, two 30-second time window feature sequences are generated during the startup process of "Graph Editor v2.1". At the same time, the startup or running characteristics of other applications or tasks (such as "Video Player v1.0", "Background Rendering Task") may also be analyzed, forming a larger feature set, from which the compilation peak duration (ms), video memory block size (Bytes), status switching frequency (Hz) are selected as the input for clustering analysis. For example, there are 100 data points from different applications / tasks / time windows, and each point is a three-dimensional vector , , . For example, point P1 = [60, 16777216, 85], P2 = [40, 8388608, 50], P3 = [100, 33554432, 120]. Calculate the mean and standard deviation of these three feature dimensions for all 100 data points respectively and standard deviation ,for example of , , of , , of , ,Perform Z-score standardization on each data point, and the calculation formula is ,for example, after standardization, P1 is , this operation is performed on all 100 points to generate a standardized feature subset. Next, the elbow method is used to determine the number of clusters k. Calculate the sum of squared errors (SSE) within the clusters when k = 2, 3, 4, 5,.... Plot the relationship between the k value and the SSE, and observe the "elbow" of the graph, that is, the turning point where the rate of decrease of the SSE changes from fast to slow. For example, if this point appears when k = 4, determine the optimal number of clusters to be 4. Initialize 4 cluster centers. For example, 4 standardized data points can be randomly selected as the initial centers C1, C2, C3, C4, and enter the iterative process. Calculate the Euclidean distance from each standardized data point (such as ) to the current 4 cluster centers (C1, C2, C3, C4). For example:
[0069] , and assign to the cluster center with the closest distance. For example, the closest one is C2, then belongs to cluster 2. Perform this assignment on all 100 data points. After the assignment is completed, recalculate the position of each cluster center. The new center coordinates are the mean of the coordinates of all data points it contains in each dimension. For example, for example, cluster 2 now contains , then the new coordinates of C2 are ([[]] , , ) Update all 4 cluster centers, compare the change amounts of the old and new center coordinates. For example, calculate the Euclidean distance between the old and new coordinates of each center, or the sum of the change amounts of all center coordinates. Set a threshold, which is set according to the accuracy requirements of the clustering result. For example, if the numerical range after feature standardization is roughly between -3 and +3, the threshold can be set to 0.001. If the sum of the change amounts of all center coordinates is less than 0.001, it is considered that the cluster centers are stable and the iteration stops. If it is greater than the threshold, repeat the assignment and update steps until the stop condition is met. After the iteration stops, obtain the optimized set of 4 cluster centers {C1', C2', C3', C4'}. Based on the final data point assignment result, assign the cluster label (1, 2, 3, or 4) to each time window of each application or task, and establish a mapping table. For example, record <application name: "Graphics Editor v2.1", time window: 1, cluster label: 2>, <application name: "Video Player v1.0", time window: 1, cluster label: 1>, etc. Then, aggregate the data according to the cluster labels, calculate the mean of the peak demand of the compilation units for all time windows under each cluster category (such as label 2) (for example, this data is in the feature set or can be obtained by associating with the original log), and analyze the timing pattern of the video memory bandwidth corresponding to these time windows (such as calculating statistical quantities such as bandwidth mean, peak value, variance, or extracting a certain timing feature). Associate the cluster label with the corresponding mean of the peak demand of the compilation units and the timing pattern of the video memory bandwidth to generate a GPU activity pattern classification index.
[0070] The steps to obtain the typical GPU resource demand timing profile are as follows:
[0071] Based on the GPU activity pattern classification index, traverse all time windows under each category label, extract the peak demand of the GPU compilation units within each time window, calculate the mean of the peak demand of the compilation units for all time windows, and generate a set of means of the peak demand of the compilation units within the category;
[0072] According to the set of means of the peak demand of the compilation units within the category, extract the timing data of the video memory bandwidth corresponding to the category label, and calculate the video memory bandwidth fluctuation intensity. The formula is:
[0073] ;
[0074] where is the video memory bandwidth fluctuation intensity, is the video memory bandwidth value at time point is the bandwidth mean, is the bandwidth mean, is the total number of seconds of the time window;
[0075] Based on the fluctuation intensity of the video memory bandwidth, parse the call timestamps of the texture mapping unit and the rasterization unit from the pipeline state setting sequence log set, count the call frequencies of each unit within the time window, calculate the median and variance of the call frequencies, and concatenate the mean of the peak demand of the compilation unit, the fluctuation intensity of the video memory bandwidth, the median and variance of the call frequencies in the order of the time window to form a multi-dimensional vector, and construct a typical GPU resource demand time series profile.
[0076] Specifically, based on the generated GPU activity pattern classification index, which maps the time window of an application or task to different activity pattern categories (cluster labels) and associates some features of each category. For example, category 2 is associated with a mean peak demand of 150 units for the compilation unit and a specific video memory bandwidth time series pattern. Now, it is necessary to construct a more detailed resource demand profile for each category. Traverse one of the category labels, such as all the time window instances under label 2 (for example, 20 time windows from different application startup processes are classified into this category), extract the peak demand of the GPU compilation unit recorded in each time window (this requires this data to be collected earlier and included in the feature set or be queryable by association), for example, the peak compilation demands of these 20 windows are [140, 155, 150, 160, 145,..., 152], calculate the mean of these 20 values to obtain the mean peak demand of the compilation unit for this category (label 2). units, generate a set containing all categories and their corresponding mean peak demands of the compilation unit. Next, according to category label 2, extract the video memory bandwidth time series data of all 20 associated time windows. , each window contains the bandwidth values at 30 time points (seconds). For example, the bandwidth data (GB / s) of one window is [50, 55, 45, 60, 50,..., 58], calculate the mean bandwidth of this window. GB / s, and then calculate the fluctuation intensity of the video memory bandwidth of this window. , the formula is , where is the fluctuation intensity of the video memory bandwidth, is the video memory bandwidth value at time point , is the mean bandwidth of this window (53 GB / s), is the total number of seconds of the time window (30 s), and the specific calculation is , for example, it is calculated that GB / s. Calculate the value for all 20 windows belonging to category 2, and then calculate the mean of these 20 values to obtain the average bandwidth fluctuation intensity of category 2. GB / s (simplify the calculation here, and actually calculate the average of P for each window). Then, based on the 20 time windows associated with Category 2, parse the timestamps of instructions or events related to Texture Mapping Unit (TMU) and Rasterization Unit (ROP) calls from the corresponding original pipeline status setting sequence log set, and count the call frequencies of each unit within each 30 - second time window. For example, for one window, the TMU call count sequence (per second) is [10k, 12k, 9k,..., 11k], and the ROP call count sequence is [5k, 6k, 4k,..., 5.5k]. Calculate the median and variance of these two sequences. For example, the median of TMU call frequency is 10.5k calls per second, and the variance is 4M( ), the median of ROP call frequency is 5.2k calls per second, and the variance is 1.5M( ). Calculate these statistics for all 20 windows of Category 2, and then take the average to obtain the median of TMU call frequency for Category 2 , variance , as well as the median of ROP call frequency , variance . Finally, concatenate the average of the peak demand of compilation units (150 units), the average of the fluctuation intensity of video memory bandwidth (6.5 GB / s), the median of TMU call frequency (10.5k / s), the variance of TMU call frequency (4M), the median of ROP call frequency (5.2k / s), and the variance of ROP call frequency (1.5M) calculated for this category (label 2) in a fixed order to form a multi - dimensional vector [150, 6.5, 10500, 4000000, 5200, 1500000]. This vector represents the typical GPU resource demand time - series profile of Category 2 (here the profile is aggregated into a single vector, and the original meaning refers to the time - series profile, and it may be necessary to retain the time - series information within the window or a more complex representation). Perform the same operation for all category labels (1, 3, 4) to construct the typical GPU resource demand time - series profiles for all categories.
[0077] Formula Detailed description:
[0078] Formula Used to calculate the fluctuation intensity of video memory bandwidth.
[0079] Parameter description:
[0080] : Represents the fluctuation intensity of video memory bandwidth within a time window.
[0081] : Represents the total duration of the time window, which is set to 30 seconds in this example.
[0082] : Represents discrete time points within the time window, counting from 1 to .
[0083] :Indicates the time window from the 1st second to the Sum all time points in seconds.
[0084] : Indicates at a point in time Monitored video memory bandwidth value. This value is collected in real time when the application is running through GPU performance monitoring tools (such as NVIDIA's nvml library interface or AMD's rocmsmi tool). For example, at the 5th second, GB / s.
[0085] : Indicates that in the entire time window The average value of the internal memory bandwidth, calculated as . This value represents the average bandwidth level during the window period.
[0086] : Indicates at a point in time Bandwidth value Average bandwidth of window This measures the degree to which the bandwidth value at that point in time deviates from the average level.
[0087] : Indicates dividing the sum by the total number of seconds in the time window to get the average absolute deviation.
[0088] Operation logic and purpose:
[0089] This formula calculates the average of the absolute deviations between the memory bandwidth value and its average value at each time point in the time window. First, calculate the average bandwidth of the entire window. Then, for each second in the window, calculate the bandwidth of that second With average bandwidth The absolute value of the difference is calculated by summing up the absolute difference of all seconds and finally dividing it by the total number of seconds in the window. This result Quantifies the average amplitude of bandwidth fluctuations around its mean value within the time window. A larger value indicates more drastic and unstable bandwidth fluctuations; a smaller value indicates more stable bandwidth usage.
[0090] Calculation example:
[0091] For example, a Bandwidth data collected in a time window of seconds (GB / s) are [50, 55, 45, 60, 50] respectively.
[0092] Calculate the average bandwidth :
[0093] GB / s.
[0094] Calculate the absolute deviation per second:
[0095] t = 1: |50 - 52| = |-2| = 2;
[0096] t = 2: |55 - 52| = |3| = 3;
[0097] t = 3: |45 - 52| = |-7| = 7;
[0098] t = 4: |60 - 52| = |8| = 8;
[0099] t = 5: |50 - 52| = |-2| = 2;
[0100] Calculate the fluctuation intensity :
[0101] GB / s.
[0102] Benefits of the formula: By calculating , scenarios with similar total bandwidth consumption but different behavioral patterns can be distinguished. For example, one task running continuously at 50 GB / s and another task switching frequently between 25 GB / s and 75 GB / s may have similar average bandwidths, but the value of the latter will be significantly higher, revealing a stronger impact on bandwidth resources, which is valuable for resource scheduling and conflict prediction.
[0103] The calculated GB / s indicates that within this 5 - second window, the video memory bandwidth deviates from the mean value of 52 GB / s by approximately 4.4 GB / s per second on average. This value will be used as part of the typical GPU resource demand time - series profile to characterize the bandwidth usage stability characteristics of the corresponding activity pattern category. A higher value may mean that tasks of this category are more likely to cause bandwidth contention.
[0104] The steps to obtain the target application startup resource demand list are as follows:
[0105] Receive the application startup request, parse the application unique identifier and version number in the request header, traverse all category labels in the GPU activity pattern classification index, compare the application identifier with the features under each category label one by one, filter out the category label with the highest feature matching degree, and extract the typical GPU resource demand time - series profile corresponding to the category label;
[0106] Based on the typical GPU resource demand timing profile data block, extract all shader identifiers, video memory block sizes, and virtual address offset base addresses from the typical GPU resource demand timing profile data block to generate an intermediate data set;
[0107] According to the video memory block sizes in the intermediate data set, sort the shader identifiers in descending order from largest to smallest, and encapsulate the sorted shader identifiers together with the corresponding video memory block sizes and virtual address offset base addresses into a target application startup resource demand list.
[0108] Specifically, when a request to start the application "Graphic Editor v2.1" with version number "v2.1" is received, first parse the request to obtain the unique identifier of the application "Graphic Editor v2.1". Then, traverse all the category tags {1, 2, 3, 4} and their associated features recorded in the established GPU activity pattern classification index. It is necessary to compare the known startup features of "Graphic Editor v2.1" (for example, the typical compilation peak duration, video memory allocation size, status switching frequency, etc. obtained by pre-analysis when the application starts, such as [60ms, 16MB, 85Hz]) with the features represented by each category tag (usually the cluster center or the average feature of the category, for example, the central feature of category 2 is close to [58ms, 17MB, 80Hz], and category 1 is [30ms, 8MB, 40Hz]), and calculate "Graphic Editor v2.The similarity or distance between the "1" feature and the features of each category, for example, using the Euclidean distance, calculates that its feature distance from Category 2 is the smallest, determines that the matching degree is the highest, so the category label 2 is selected. Then, according to the matched category label 2, extract its corresponding typical GPU resource demand time series profile data block constructed, which contains information such as the mean value of the peak demand of the compilation unit, the bandwidth fluctuation intensity, and the TMU / ROP call statistics. At the same time, it is also necessary to extract the specific resource items required in the typical startup scenario of this category from the more detailed data associated with Category 2 (possibly stored in the index or retrieved through the index to find the original data), including all relevant shader identifiers (such as "Shader_UI_Main", "Shader_Render_Core", "Shader_PostFX"), their corresponding video memory block sizes (such as 8MB, 16MB, 4MB), and the recommended virtual address offset base address at allocation (such as 0x10000000, 0x20000000, 0x30000000). Pool these extracted {shader identifier, video memory block size, virtual address offset base address} entries to generate an intermediate data set, for example, containing a list: [{ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}]. Sort the entries in the list in descending order according to the video memory block size Size in the intermediate data set. After sorting, get: [{ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}]. Package this sorted list containing shader identifiers, corresponding video memory block sizes, and virtual address offset base addresses to form a resource demand list required for the startup of the target application "Graphics Editor v2.1".
[0109] The steps to obtain the GPU preprocessing instruction sequence are as follows:
[0110] Parse the startup resource demand list of the target application, traverse the key-value pairs of each entry in the startup resource demand list of the target application, and extract the shader identifier, video memory block size, and virtual address offset base address;
[0111] Based on the shader identifier, each shader identifier is replaced with an instruction string according to the pre-compilation instruction template. Based on the video memory block size and the virtual address offset base address, each video memory block size and virtual address offset base address are replaced with an instruction string according to the video memory pre-allocation instruction template to generate a GPU preprocessing instruction sequence.
[0112] Specifically, parse the generated target application startup resource requirement list, which is an ordered list. For example: [{ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}]. Traverse each entry in this list. Each entry is a data structure containing key-value pairs. For the first entry {ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, extract the shader identifier "Shader_Render_Core", the video memory block size 16MB (i.e., 16777216 Bytes), and the virtual address offset base address 0x20000000. Based on the extracted shader identifier "Shader_Render_Core", use the predefined pre-compilation instruction template. For example, the template is CMD_PRECOMPILE_SHADERName= <shaderid>, will <shaderid>Replace it with "Shader_Render_Core" to generate the instruction string "CMD_PRECOMPILE_SHADERName=Shader_Render_Core". Similarly, based on the extracted video memory block size of 16777216 Bytes and the virtual address offset base address of 0x20000000, use the predefined video memory preallocation instruction template, such as the template CMD_PREALLOCATE_MEMORYSize= <sizeinbytes>Address= <virtualaddress>, will <sizeinbytes>Replace with 16777216, the <virtualaddress>Replace it with 0x20000000 to generate the instruction string "CMD_PREALLOCATE_MEMORY Size=16777216 Address=0x20000000". Perform the same operation on the second entry in the list {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000} to generate the instructions "CMD_PRECOMPILE_SHADER Name=Shader_UI_Main" and "CMD_PREALLOCATE_MEMORY Size=8388608 Address=0x10000000". Perform the same operation on the third entry {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000} to generate the instructions "CMD_PRECOMPILE_SHADER Name=Shader_PostFX" and "CMD_PREALLOCATE_MEMORY Size=4194304 Address=0x30000000". Combine all the instruction strings generated in the list order (or other predefined logic, such as allocate first then compile) to form the final GPU preprocessing instruction sequence.
[0113] The steps to obtain the task preemption priority and cost evaluation table are as follows:
[0114] Monitor the set of tasks currently running on the GPU. Through the operating system kernel interface, obtain the process identifier of each task and the associated typical GPU resource demand timing profile. Extract the time consumption of the last three context switches of each task from the system scheduler log. Calculate the average value of the time consumption of the last three context switches as the context switch overhead to generate a data set;
[0115] Based on the data set, calculate the task preemption priority score. The formula is:
[0116] ;
[0117] Among them, is the task preemption priority score, is the task the time consumption of the nth context switch, is the video memory bandwidth occupancy rate of the task , is the foreground window activation state, 1 for activated and 0 for not activated;
[0118] Sort all tasks according to the task preemption priority score from high to low, generate entries for each task, where the entries include the process identifier, video memory bandwidth occupancy rate, average switching time, activation status, and task preemption priority score, and construct a task preemption priority and cost evaluation table.
[0119] Specifically, monitor the set of tasks running on the current GPU. For example, if three tasks are found: Task A (game rendering), Task B (video decoding), and Task C (scientific computing), obtain their process identifiers (PIDs), which are PID1234, PID5678, and PID9012 respectively, by calling the operating system kernel interface (such as Linux's / proc file system or Windows' ProcessExplorer API), and query the typical GPU resource demand time series profiles associated with these PIDs (for example, those already generated and stored for these tasks or their categories in the previous steps). Then, find the last three context switch records related to these three PIDs from the system scheduler logs (such as GPU driver logs or operating system scheduler logs), and extract the time taken for each switch. For example, the last three switch times for Task A (PID1234) are [4.5ms, 5.0ms, 4.8ms], for Task B (PID5678) are [12.0ms, 11.5ms, 12.5ms], and for Task C (PID9012) are [2.0ms, 2.2ms, 2.1ms]. Calculate the average of the last three context switch times for each task as its context switch overhead. For Task A: ms, for Task B: ms, for Task C: ms. At the same time, obtain the video memory bandwidth occupancy rate of each current task through a GPU monitoring tool. For example, Task A is 70%, Task B is 15%, and Task C is 30%, and query whether the window associated with each task is the foreground active window through the window manager API. For example, Task A is the current active game window ( ), Task B and C are background tasks ( ). Integrate the collected and calculated data to form a data set containing {PID, , , }, and based on this set, calculate the preemption priority score for each task. Use the formula where is the switching time of the kth time, is the calculated previously, is a percentage value. is 0 or 1.
[0120] Task A (PID1234): ,
[0121] Task B (PID5678): ,
[0122] Task C (PID9012): ,
[0123] According to the calculated task preemption priority scores , sort all tasks from high to low: Task A (76.32) > Task B (30.0) > Task C (8.4), generate an entry for each task, including the process identifier, video memory bandwidth occupancy rate, average switching time, activation status, and the calculated task preemption priority score, and construct a task preemption priority and cost evaluation table.
[0124] Table 1: Task Preemption Priority and Cost Evaluation Table
[0125] Process identifier Video memory bandwidth occupancy rate (%) Average switching time (ms) Activation status Task preemption priority score 1234 70 4.77 1 76.32 5678 15 12.00 0 30.00 9012 30 2.10 0 8.40
[0126] As shown in Table 1, this table shows the evaluation results of the tasks running on the current GPU, sorted in descending order according to the task preemption priority scores. The higher the score, the more important the task or the higher the preemption cost, and its execution should be ensured first.
[0127] Formula Detailed description:
[0128] Formula is used to calculate the task preemption priority score.
[0129] Parameter description:
[0130] : The task preemption priority score of task , which is a comprehensive score. The higher the value, the less likely the task should be preempted (i.e., the higher the priority or the greater the preemption cost).
[0131] : Represents a specific GPU task.
[0132] : Represents the index of the number of context switches, taking the last three here (k = 1, 2, 3).
[0133] : Represents the sum of the context switch times for the last three times.
[0134] : Task The context switch time recorded for the th time. This data is obtained from the system scheduler or GPU driver logs. For example, for task A, ms, ms,
[0135] : Calculate the average time of the last three context switches of the task, that is, . This value reflects the direct time cost of switching this task.
[0136] : The current video memory bandwidth occupancy rate of the task, given as a percentage (e.g., 70%). This data is obtained by real-time monitoring of GPU performance counters.
[0137] : This is a weight factor based on the video memory bandwidth occupancy rate. The higher the occupancy rate, the larger this factor, thus increasing the score. Dividing by 10 is to adjust the magnitude of its impact, and adding 1 ensures that the factor is at least 1.
[0138] : The window activation status of the task. If the window associated with the task is the foreground window for the current user interaction, then ; if it is a background window, then . This status is obtained by querying the operating system window manager.
[0139] : This is a weight factor based on the activation status. For foreground tasks ( ), this factor is 2; for background tasks ( ), this factor is 1. This directly doubles the score of foreground tasks.
[0140] Operation logic and purpose: This formula evaluates the preemption priority of a task (or, the "cost" of being preempted) by combining three aspects of factors:
[0141] Switching cost ( ): The longer the switching time, the higher the base score.
[0142] Resource occupancy ( ): The higher the video memory bandwidth occupancy rate, the larger the coefficient multiplied, and the higher the score. This reflects that high-bandwidth tasks may be more important or have a greater impact on interrupts.
[0143] User interaction ( ):The foreground task directly obtains double weight, reflecting the priority guarantee for user experience.
[0144] Multiply the three to get the final score . Its purpose is to quantitatively evaluate the "value" of each current task not being preempted, providing a basis for subsequent preemption decisions. Tasks with high scores are objects that the system tends to protect and not preempt easily.
[0145] Example:
[0146] Task A (PID1234): ms, , .
[0147] ;
[0148] ;
[0149] ;
[0150] ;
[0151] ;
[0152] Benefits of the formula: By comprehensively considering the switching overhead of the task itself, the occupancy of critical resources (video memory bandwidth), and the interaction status with the user, the formula provides a relatively comprehensive task importance evaluation index, which is better than the judgment based on a single factor (such as priority or resource occupancy), and helps to make more intelligent and balanced preemption decisions.
[0153] Numerical result correlation: The calculated , , These score values are directly used to construct a task preemption priority and cost evaluation table (as shown in Table 1). In subsequent decisions, these scores will be used to identify which tasks are high-priority (high scores, not easily preempted) and low-priority (low scores, can be considered as preemption objects).
[0154] The steps to obtain the dynamic GPU task preemption control parameter set are as follows:
[0155] Based on the task preemption priority and cost evaluation table, screen out the tasks with the top 20% of task preemption priority scores as the high-priority task set, extract the average video memory bandwidth occupancy rate and response delay requirements of the high-priority task set, and at the same time extract the top 30% of the tasks with the highest average switching time among the remaining tasks as the low-priority task set, generating a high-priority response requirement data set and a low-priority switching cost data set;
[0156] According to the high-priority response requirement data set and the low-priority switching cost data set, calculate the preemption trigger threshold. The formula is:
[0157] ;
[0158] where is the preemption trigger threshold, is the average video memory bandwidth occupancy rate of high-priority tasks, is the maximum allowable response delay of high-priority tasks, is the average switching time of low-priority tasks, is the proportion of available video memory of the current GPU;
[0159] Based on the preemption trigger threshold, set the trigger condition as follows: when the predicted response delay of high-priority tasks exceeds the maximum allowable response delay of high-priority tasks and the preemption trigger threshold is greater than 100, calculate the preemption time point as the start time of the next vertical blanking period , and the formula is , is the current frame number, and the preemption control parameter set is obtained.
[0160] Specifically, based on the constructed task preemption priority and cost evaluation table (see Table 1), first screen the task preemption priority scores The top 20% of the ranked tasks are used as the high-priority task set. Currently, there are 3 tasks, and 20% is 0.6. Rounding up to 1, so the task A with the highest score (PID1234, ) is selected as the high-priority task set {TaskA}, and the average video memory bandwidth occupancy rate of the tasks in this set is extracted . Since there is only one task, the average value is its occupancy rate of 70%. At the same time, obtain its response delay requirement , which is usually determined by the application type. For example, for the game task A, the response delay is required to not exceed 1 frame time, such as 16.67 ms (corresponding to a 60 Hz refresh rate). Set to 16.67 ms = 0.01667 s to form the high-priority response requirement data set { , }. Then, from the remaining tasks (task B, task C), select the highest 30% according to their average switching times( ) [12.0 ms, 2.1 ms]. 30% of 2 tasks is 0.6, rounding up to 1. The task with the highest switching time is task B (12.0 ms). Therefore, select task B (PID5678) as the representative low-priority task (for cost evaluation), and extract its average switching time ms = 0.012 s to form the low-priority switching cost data set { }, calculate the preemption trigger threshold based on these two data sets , using the formula , where is the average bandwidth occupancy rate of high-priority tasks (using the percentage value 70), is the maximum allowable response delay of high-priority tasks (0.01667s), is the average switching time of low-priority tasks (0.012s), is the proportion of the currently available GPU video memory, which needs to be obtained by real-time monitoring. For example, if the currently available video memory is 30% of the total video memory, then , substitute the values for calculation:
[0161] ,
[0162] Next, set the preemption trigger condition, which is: when the predicted response delay value of the high-priority task (Task A) (predicted by the performance model or inferred from recent historical data, for example, predicted to be 18ms = 0.018s) exceeds its maximum allowable response delay (0.01667s), and the calculated preemption trigger threshold (53.26) is greater than a preset reference value, for example, 100, the judgment condition: ( ) and ( ), that is, (0.018s > 0.01667s) and (53.26 > 100). The first condition is true, and the second condition is false. Therefore, the overall trigger condition is false, and preemption is not triggered. [Adjust the example to trigger preemption] For example, the available video memory is extremely low, , and the switching cost of low-priority tasks is very low, ms = 0.002s, then . At this time, the condition becomes (0.018s > 0.01667s) and (130.46 > 100). Both conditions are true, preemption is triggered, and calculate the time point for preemption execution , which is set as the start time of the next vertical blanking period. If the current monitor refresh rate is 60Hz, the V-blank period is approximately 16.67ms, and the dynamic GPU task preemption control parameter set is obtained.
[0163] Formula Detailed description:
[0164] Formula is used to calculate the dynamic preemption trigger threshold.
[0165] Parameter description:
[0166] : Preemption trigger threshold, a unitless value used to determine whether preemption should be performed.
[0167] : The average video memory bandwidth occupancy rate of the high-priority task set, substituted for calculation with a percentage value (e.g., 70). This value reflects the demand intensity of high-priority tasks for bandwidth resources. It is obtained by screening out the tasks with the highest scores (such as the top 20%) and calculating the average of their current bandwidth occupancy.
[0168] : The maximum allowable response latency of high-priority tasks. This is a quality of service (QoS) requirement, usually set according to the application type (such as interactive applications, games). For example, for a 60FPS game, it can be set to 1 / 60 0.01667s.
[0169] : The average context switch time of the low-priority task set (e.g., the 30% of tasks with the highest switching costs). This value represents the time cost required to preempt these low-priority tasks. It is obtained by screening out tasks with higher switching costs and calculating their average switching time.
[0170] : The proportion of currently available video memory on the GPU, a value between 0 and 1 (e.g., 0.3 represents 30% available). This value is obtained by querying the GPU status in real time and reflects the current tightness of video memory resources.
[0171] : The square root of the available video memory proportion. This factor adjusts the sensitivity of the threshold to video memory availability.
[0172] Operation logic and purpose:
[0173] This formula aims to establish a dynamic preemption decision threshold that weighs several key factors:
[0174] The urgency of high-priority tasks ( ): The product of high bandwidth occupancy rate and strict latency requirements (small ) gives a measure reflecting the potential performance risk of high-priority tasks. The larger this value, the more likely high-priority tasks are to exceed the deadline due to resource shortages, so it tends to increase the threshold , making preemption more likely to occur.
[0175] Preemption cost ( ): The average switching cost of low-priority tasks appears in the denominator. The higher the cost, the lower the threshold , making preemption less likely to occur because the preemption itself is costly.
[0176] System resource status( ): Proportion of available video memory Affects the threshold in the form of a square root. When there is more available video memory( is larger), is also larger, which will increase the threshold , meaning that the system resources are relatively abundant and the need for preemption is not so urgent. Conversely, when there is very little available video memory( is close to 0), is also close to 0, which will significantly reduce the threshold , making preemption more likely to be triggered even if other factors remain unchanged, reflecting that more active scheduling intervention is needed when resources are scarce.
[0177] Generally speaking, The higher the value, the stronger the reason the system believes is needed to trigger preemption (for example, the delay prediction of high-priority tasks far exceeds their requirements).
[0178] Numerical example:
[0179] , , , .
[0180] ;
[0181] ;
[0182] ;
[0183] ;
[0184] ;
[0185] Benefits of the formula: The formula dynamically adapts to the system state (high-priority task requirements, preemption cost, memory pressure), making the preemption decision not based on fixed rules but adjusted according to the current specific situation, which helps to achieve a better balance between ensuring the performance of high-priority tasks and avoiding unnecessary preemption overhead.
[0186] Numerical result association: The calculated . This value will be compared with a fixed reference value (for example, 100). In this example, . Combining the condition that the predicted delay of high-priority tasks exceeds the limit ( ), jointly determines whether to trigger preemption. The The value itself is not directly part of the preemption control parameter set, but it is a key intermediate calculation result for generating the parameter set (especially the trigger state). A high value combined with a delay overrun constitutes a sufficient condition for executing preemption.
[0187] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as they do not depart from the technical solution content of the present invention, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.< / virtualaddress> < / sizeinbytes> < / virtualaddress> < / sizeinbytes> < / shaderid> < / shaderid>
Claims
1. A method for optimizing and improving the performance of a graphics card based on big data, characterized in that, Including the following steps: Collect the shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, perform data cleaning, and establish the application startup and multi-task GPU activity feature set; Based on the application startup and multi-task GPU activity feature set, group the application startup behavior and task running status through numerical clustering, obtain the category labels to which each application or task belongs, establish a GPU activity mode classification index, and based on the GPU activity mode classification index, aggregate the corresponding feature set data for each category label, calculate the average value of the peak demand of GPU compilation units within the category and the video memory bandwidth, and construct a typical GPU resource demand time series profile; According to the received application startup request, retrieve the typical GPU resource demand time series profile of the matching application, extract the list of shader identifiers and the initial video memory allocation mode parameter values for the startup call, generate a target application startup resource demand list, convert the shader identifiers in the target application startup resource demand list into pre-compiled instructions, convert the video memory allocation mode parameters into video memory pre-allocation instructions, and establish a GPU preprocessing instruction sequence; Monitor the set of tasks running on the current GPU, retrieve the typical GPU resource demand time series profile and associated context switching overhead corresponding to each task, generate a task preemption priority and cost evaluation table, and based on the task preemption priority and cost evaluation table, set the preemption trigger conditions and time points to obtain a dynamic GPU task preemption control parameter set; The steps for obtaining the typical GPU resource demand time series profile are as follows: Based on the GPU activity mode classification index, traverse all time windows under each category label, extract the peak demand of GPU compilation units within each time window, calculate the average value of the peak demand of compilation units in all time windows, and generate a set of average values of peak demand of compilation units within the category; According to the set of average values of peak demand of compilation units within the category, extract the video memory bandwidth time series data corresponding to the category label, calculate the video memory bandwidth fluctuation intensity, and the formula is: ; Among them, is the fluctuation intensity of the video memory bandwidth, is the time point of the video memory bandwidth value, is the bandwidth average value, is the total number of seconds of the time window; Based on the video memory bandwidth fluctuation intensity, calculate the average value of the video memory bandwidth fluctuation intensity in all time windows corresponding to the category label, parse the call timestamps of the texture mapping unit and rasterization unit from the pipeline state setting sequence log set, count the call frequency of each unit within the time window, calculate the median and variance of the call frequency, and concatenate the average value of the peak demand of compilation units, the video memory bandwidth fluctuation intensity average value, the call frequency median, and variance in sequence into a multi-dimensional vector to construct a typical GPU resource demand time series profile.
2. The method for optimizing and improving the performance of a graphics card based on big data according to claim 1, wherein The steps for obtaining the application startup and multi-task GPU activity feature set are as follows: Collect shader compile time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, calculate the 25th percentile Q1 and 75th percentile Q3 of the shader compile time data, remove extreme values lower than Q1-1.5×(Q3-Q1) or higher than Q3+1.5×(Q3-Q1), and generate cleaned shader compile time data sets, initial video memory allocation mode parameter sets, and pipeline state setting sequence log sets; Based on the cleaned shader compilation time data set, traverse the compilation durations recorded at all time points, select the maximum value as the compilation peak duration, parse the video memory block size field from the initial video memory allocation mode parameter set, extract the fixed allocation unit value of the video memory block size, count the number of pipeline state switching instructions per second from the pipeline state setting sequence log set as the state switching frequency, calculate the average of the absolute values of the differences between adjacent state switching timestamps as the switching delay, and count the total number of shader instruction executions per second as the computational intensity; The compilation peak duration, video memory block size, state switching frequency, computational intensity, and switching delay are aligned by timestamp and merged into a multidimensional vector at the same time point. The time series is divided into time windows of 30 seconds to establish a feature set of application startup and multi-tasking GPU activities.
3. The method for optimizing and improving the performance of a graphics card based on big data according to claim 1, characterized in that, The steps for obtaining the GPU activity mode classification index are as follows: Based on the application startup and multi-task GPU activity feature set, three feature dimensions, namely, compilation peak duration, video memory block size, and state switching frequency, are selected as clustering inputs, and each data point is Z-score standardized to generate a standardized feature subset; Initialize k cluster centers according to the standardized feature subset, where the k value is determined by the elbow rule, calculate the Euclidean distance from each data point to each cluster center, assign the data point to the nearest neighbor cluster center, and iteratively update the cluster center coordinates to the mean of the data points to which they belong, until the center coordinate change for two consecutive iterations is less than a threshold, thereby generating an optimized cluster center set; Based on the optimized cluster center set, a cluster label is assigned to each time window of each application or task, a mapping table is established, and the mean peak demand of compilation units and the timing mode of video memory bandwidth of the same category are aggregated by label to generate a classification index of GPU activity mode.
4. The method for optimizing and improving the performance of a graphics card based on big data according to claim 1, wherein The steps for obtaining the target application startup resource requirement list are as follows: Receive an application startup request, parse the application unique identifier and version number in the request header, traverse all category tags in the GPU activity mode classification index, compare the application identifier with the features under each category tag one by one, select the category tag with the highest feature matching degree, and extract the typical GPU resource demand timing profile corresponding to the category tag; Based on the typical GPU resource requirement timing profile data block, extract all shader identifiers, video memory block sizes and virtual address offset base addresses from the typical GPU resource requirement timing profile data block to generate an intermediate data set; According to the video memory block size in the intermediate data set, sort the shader identifiers in descending order, and encapsulate the sorted shader identifiers, the corresponding video memory block size, and the virtual address offset base address into the target application startup resource requirement list.
5. The method for optimizing and improving the performance of a graphics card based on big data according to claim 4, wherein The steps for obtaining the GPU preprocessing instruction sequence are as follows: Parse the target application startup resource requirement list, traverse the key-value pairs of each entry in the target application startup resource requirement list, and extract the shader identifier, the video memory block size, and the virtual address offset base address; Based on the shader identifier, replace each shader identifier with an instruction string according to the precompiled instruction template. Based on the video memory block size and the virtual address offset base address, replace each video memory block size and virtual address offset base address with an instruction string according to the video memory preallocation instruction template, and generate the GPU preprocessing instruction sequence.
6. The method for optimizing and improving the performance of a graphics card based on big data according to claim 1, wherein The steps for obtaining the task preemption priority and cost evaluation table are as follows: Monitor the set of tasks currently running on the GPU, obtain the process identifier of each task and the associated typical GPU resource requirement time series profile through the operating system kernel interface, extract the time consumption of the last three context switches of each task from the system scheduler log, calculate the average value of the time consumption of the last three context switches as the context switch overhead, and generate a data set; Based on the data set, calculate the task preemption priority score, and the formula is: ; Among them, is the task preemption priority score, is the task the time consumption of the is the task video memory bandwidth occupancy rate, is the foreground window activation status, activated as 1, not activated as 0; According to the task preemption priority score, sort all tasks from high to low, generate an entry for each task, and the entry includes the process identifier, the video memory bandwidth occupancy rate, the average switching time consumption, the activation state, and the task preemption priority score, and construct the task preemption priority and cost evaluation table.
7. The method for optimizing and improving the performance of a graphics card based on big data according to claim 1, wherein The steps for obtaining the dynamic GPU task preemption control parameter set are as follows: Based on the task preemption priority and cost evaluation table, screen out the tasks with the top 20% of the task preemption priority scores as the high-priority task set, extract the average video memory bandwidth occupancy rate and the response delay requirement of the high-priority task set, and at the same time extract the top 30% of the tasks with the highest average switching time consumption among the remaining tasks as the low-priority task set, and generate the high-priority response requirement data set and the low-priority switching cost data set; According to the high-priority response requirement data set and the low-priority switching cost data set, calculate the preemption trigger threshold, and the formula is: ; Among them, is the preemption trigger threshold, is the average value of the video memory bandwidth occupancy rate of high-priority tasks, is the maximum allowable response delay of high-priority tasks, is the average value of the switching time of low-priority tasks, is the proportion of the available video memory of the current GPU; Based on the preemption trigger threshold, the trigger condition is set as follows: when the predicted response delay of the high-priority task exceeds the maximum allowable response delay of the high-priority task and the preemption trigger threshold is greater than 100, the preemption time point is calculated as the start time of the next vertical blanking period , and the formula is , is the current frame number, and the preemption control parameter set is obtained.
8. The graphics card performance optimization and improvement system for the method of optimizing and improving graphics card performance based on big data according to any one of claims 1-7, characterized in that, Including: The data acquisition module records the shader compilation time data during the application startup phase, measures the video memory allocation mode parameters, traces the pipeline state setting sequence, eliminates outliers, aligns multi-source data according to the time stamp, and constructs the graphics card behavior feature set; The feature clustering module extracts the time series feature vectors from the graphics card behavior feature set, calculates the distance between the vectors, sets the clustering center, groups the application behaviors, assigns category identifiers, establishes a mapping relationship table, statistically analyzes the time distribution, extracts the GPU compilation unit demand, calculates the bandwidth utilization rate, performs time series decomposition, forms a resource change curve, marks the peaks and valleys, and stores it as the GPU resource requirement profile library; Resource requirement analysis module, query the GPU resource requirement profile library, match the application category, extract the resource profile data, analyze the shader call sequence, extract the identifier list, parse the peak memory allocation, deduce the pre-allocated space, calculate the bandwidth threshold, mark the resource competition points, and output the application startup resource planning table; Preprocessing instruction generation module, read the shader identifiers in the application startup resource planning table, query the shader library, verify the compilation status, filter the pre-compiled subset, construct the instruction sequence, parse the video memory allocation mode parameters, convert them into reservation instructions, set the block size, formulate the memory plan, combine the priority sorting, and encapsulate them into the startup optimization instruction sequence; Task scheduling control module, scan the GPU task set, identify the resource occupancy, query the GPU resource requirement profile library to obtain the prediction, collect the switching records, calculate the switching time, compare the task priorities, calculate the preemption cost, set the competition threshold, determine the preemption timing, and form a dynamic resource scheduling plan.
Citation Information
Patent Citations
Multi-threaded, self-scheduling processor
CN112088359A
Method and device for dynamically deploying GPU resources and computer equipment
CN112559191A