Method and system for optimizing and improving performance of graphics card based on big data
By collecting and analyzing the data in the application startup stage, establishing a GPU activity feature set and generating a resource requirement timing profile, the problem of the fixed priority scheduling mode in the existing technology is difficult to adapt to dynamic changes, and accurate prediction and dynamic regulation of GPU resources are achieved, and graphics card performance and task execution efficiency are improved.
Patent Information
- Application Number
- CN202510519447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the mixed deployment of deep learning training and real-time rendering tasks, the fixed priority scheduling mode is difficult to adapt to the dynamic changes in application behavior, resulting in rendering pipeline waiting, video memory allocation failure, task cascade timeout or frame loss.
By collecting shader compilation time data, initial memory allocation mode parameters and pipeline status settings sequence logs in the application startup stage, an application startup and multitasking GPU activity feature set is established. Numerical clustering is used to group the application startup behavior and task running status to generate a GPU activity mode classification index. Based on this index, the average value and memory bandwidth of the peak demand of GPU compilation units within the category are calculated to construct a timing profile of typical GPU resource requirements.
It realizes the quantitative and accurate prediction of the dynamic demand for hardware resources of different application startup behaviors, reduces the risk of video memory bandwidth competition in high-volatility scenarios, avoids the problems of compilation delay and video memory fragmentation, ensures that high-priority tasks seize resources at the right time, and reduces the sharp drop in rendering frame rate or computing pipeline blockage.
Smart Images

Figure CN120045335A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of program control, and particularly to a method and system for optimizing and improving the performance of a graphics card based on big data. Background Art
[0002] The technical field of program control focuses on the automated management and dynamic regulation of the operation process of computing devices through algorithms, protocols, and system-level policies. Its core goal is to optimize the hardware performance and task execution efficiency through real-time monitoring, feedback decision-making, and resource scheduling.
[0003] Traditional program control technologies rely on preset fixed rules or empirical thresholds for resource allocation and task scheduling, lacking adaptability to the dynamic changes of application behaviors. For example, when deploying deep learning training and real-time rendering tasks in a hybrid manner, the fixed priority scheduling mode may cause the rendering pipeline to wait due to the lack of pre-allocation of texture mapping units, or fail to allocate real-time tasks in time due to the failure to recycle the video memory of training tasks in time. Moreover, when a high-priority task is inserted suddenly, it is prone to excessive preemption or response lag, resulting in cascading task timeouts or frame losses. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose a method and system for optimizing and improving the performance of a graphics card based on big data.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. A method for optimizing and improving the performance of a graphics card based on big data includes the following steps: Collect shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application program, perform data cleaning, and establish a feature set of application program startup and multi-task GPU activities; Based on the feature set of application program startup and multi-task GPU activities, group the application program startup behaviors and task running states through numerical clustering, obtain the category labels to which each application or task belongs, establish a classification index of GPU activity patterns, and based on the classification index of GPU activity patterns, aggregate the corresponding feature set data for each category label, calculate the average value of the peak demand of GPU compilation units within the category and the video memory bandwidth, and construct a typical GPU resource demand time series profile; According to the received application program startup request, retrieve the typical GPU resource demand time series profile of the matching application, extract the list of shader identifiers for startup calls and the predicted initial video memory allocation mode parameter values, generate a target application startup resource demand list, convert the shader identifiers in the target application startup resource demand list into pre-compiled instructions, convert the video memory allocation mode parameters into video memory pre-allocation instructions, and establish a GPU preprocessing instruction sequence; Monitor the task set currently running on the GPU, retrieve the typical GPU resource demand timing profile and the associated context switching overhead corresponding to each task, generate a task preemption priority and cost evaluation table, set the preemption trigger conditions and time points based on the task preemption priority and cost evaluation table, and obtain a dynamic GPU task preemption control parameter set.
[0006] Preferably, the steps of obtaining the application startup and multi-task GPU activity feature set are: Collect shader compile time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, calculate the 25th percentile Q1 and 75th percentile Q3 of the shader compile time data, remove extreme values lower than Q1-1.5×(Q3-Q1) or higher than Q3+1.5×(Q3-Q1), and generate cleaned shader compile time data sets, initial video memory allocation mode parameter sets, and pipeline state setting sequence log sets; Based on the cleaned shader compilation time data set, traverse the compilation durations recorded at all time points, select the maximum value as the compilation peak duration, parse the video memory block size field from the initial video memory allocation mode parameter set, extract the fixed allocation unit value of the video memory block size, count the number of state switching instructions per second from the pipeline state setting sequence log set as the state switching frequency, calculate the average of the absolute values of the differences between adjacent state switching timestamps as the switching delay, and count the total number of shader instruction executions per second as the computational intensity; The compilation peak duration, video memory block size, state switching frequency, computational intensity, and switching delay are aligned by timestamp and merged into a multidimensional vector at the same time point. The time series is divided into time windows of 30 seconds to establish a feature set of application startup and multi-tasking GPU activities.
[0007] Preferably, the steps of obtaining the GPU activity mode classification index are: Based on the application startup and multi-task GPU activity feature set, three feature dimensions, namely, compilation peak duration, video memory block size, and state switching frequency, are selected as clustering inputs, and each data point is Z-score standardized to generate a standardized feature subset; Initialize k cluster centers according to the standardized feature subset, where the k value is determined by the elbow rule, calculate the Euclidean distance from each data point to each cluster center, assign the data point to the nearest neighbor cluster center, and iteratively update the cluster center coordinates to the mean of the data points to which they belong, until the center coordinate change for two consecutive iterations is less than a threshold, thereby generating an optimized cluster center set; Based on the optimized cluster center set, a cluster label is assigned to each time window of each application or task, a mapping table is established, and the mean peak demand of compilation units and the timing mode of video memory bandwidth of the same category are aggregated by label to generate a classification index of GPU activity mode.
[0008] Preferably, the steps for obtaining the typical GPU resource requirement timing profile are: Based on the GPU activity mode classification index, traverse all time windows under each category label, extract the peak demand of GPU compilation units in each time window, calculate the average of the peak demand of compilation units in all time windows, and generate a set of average peak demand of compilation units in the category; According to the mean value set of peak demand of compilation units in the category, the timing data of video memory bandwidth corresponding to the category label is extracted, and the fluctuation intensity of video memory bandwidth is calculated. The formula is: ; in, is the memory bandwidth fluctuation intensity, For time point The memory bandwidth value, is the bandwidth mean, is the total number of seconds in the time window; Based on the memory bandwidth fluctuation intensity, the call timestamps of the texture mapping unit and the rasterization unit are parsed from the pipeline state setting sequence log set, the call frequency of each unit in the time window is counted, the median and variance of the call frequency are calculated, and the mean of the compilation unit peak demand, the memory bandwidth fluctuation intensity, the median and variance of the call frequency are spliced into a multi-dimensional vector in the order of the time window to construct a typical GPU resource demand timing profile.
[0009] Preferably, the steps of obtaining the target application startup resource requirement list are: Receive an application startup request, parse the application unique identifier and version number in the request header, traverse all category tags in the GPU activity mode classification index, compare the application identifier with the features under each category tag one by one, select the category tag with the highest feature matching degree, and extract the typical GPU resource demand timing profile corresponding to the category tag; Based on the typical GPU resource requirement timing profile data block, extract all shader identifiers, video memory block sizes and virtual address offset base addresses from the typical GPU resource requirement timing profile data block to generate an intermediate data set; According to the video memory block size in the intermediate data set, the shader identifiers are arranged in descending order from large to small, and the sorted shader identifiers are packaged with the corresponding video memory block size and virtual address offset base address into a target application startup resource requirement list.
[0010] Preferably, the step of obtaining the GPU preprocessing instruction sequence is as follows: Parse the target application startup resource requirement list, traverse the key-value pairs of each entry in the target application startup resource requirement list, and extract the shader identifier, video memory block size, and virtual address offset base address; Based on the shader identifier, replace each shader identifier with an instruction string according to the precompiled instruction template. Based on the video memory block size and virtual address offset base address, replace each video memory block size and virtual address offset base address with an instruction string according to the video memory preallocation instruction template to generate a GPU preprocessing instruction sequence.
[0011] Preferably, the step of obtaining the task preemption priority and cost evaluation table is as follows: Monitor the task set currently running on the GPU, obtain the process identifier of each task and the associated typical GPU resource requirement timing profile through the operating system kernel interface, extract the time consumption of the last three context switches of each task from the system scheduler log, calculate the average value of the time consumption of the last three context switches as the context switch overhead, and generate a data set; Based on the data set, calculate the task preemption priority score. The formula is: ; Wherein, is the task preemption priority score, is the task the time consumption of the nth context switch, is the task video memory bandwidth occupancy rate, is the foreground window activation status, activated as 1, not activated as 0; According to the task preemption priority score, sort all tasks from high to low, generate an entry for each task. The entry includes the process identifier, video memory bandwidth occupancy rate, average switching time consumption, activation status, and task preemption priority score, and construct a task preemption priority and cost evaluation table.
[0012] Preferably, the step of obtaining the dynamic GPU task preemption control parameter set is as follows: Based on the task preemption priority and cost evaluation table, screen out the tasks with the top 20% of the task preemption priority scores as the high-priority task set, extract the average video memory bandwidth occupancy rate and response delay requirements of the high-priority task set, and at the same time extract the top 30% of the tasks with the highest average switching time consumption among the remaining tasks as the low-priority task set, and generate a high-priority response requirement data set and a low-priority switching cost data set; According to the high-priority response requirement data set and the low-priority switching cost data set, calculate the preemption trigger threshold. The formula is: ; Wherein, is the preemption trigger threshold, is the average value of the video memory bandwidth occupancy rate of high-priority tasks, is the maximum allowable response delay of high-priority tasks, is the average value of the switching time of low-priority tasks, is the proportion of available video memory of the current GPU; Based on the preemption trigger threshold, the trigger condition is set as follows: when the predicted value of the response delay of high-priority tasks exceeds the maximum allowable response delay of high-priority tasks, and the preemption trigger threshold is greater than 100, the preemption time point is calculated as the start time of the next vertical blanking period , and the formula is , is the current frame number, and the preemption control parameter set is obtained.
[0013] The present invention provides a graphics card performance optimization and improvement system, including: A data acquisition module, which records the shader compilation time data in the application startup phase, measures the video memory allocation mode parameters, tracks the pipeline state setting sequence, eliminates outliers, aligns multi-source data according to time stamps, and constructs a graphics card behavior feature set; A feature clustering module, which extracts time series feature vectors from the graphics card behavior feature set, calculates the distance between vectors, sets the clustering center, groups application behaviors, assigns category identifiers, establishes a mapping relationship table, statistically analyzes the time distribution, extracts the demand for GPU compilation units, calculates the bandwidth utilization rate, performs time series decomposition, forms a resource change curve, marks peaks and valleys, and stores it as a GPU resource demand profile library; A resource demand analysis module, which queries the GPU resource demand profile library, matches the application category, extracts resource profile data, analyzes the shader call sequence, extracts the identifier list, parses the memory allocation peak value, deduces the pre-allocated space, calculates the bandwidth threshold, marks the resource competition points, and outputs an application startup resource planning table; A preprocessing instruction generation module, which reads the shader identifiers in the application startup resource planning table, queries the shader library, verifies the compilation status, filters the pre-compiled subset, constructs an instruction sequence, parses the video memory allocation parameters, converts them into reservation instructions, sets the block size, formulates a memory scheme, combines priority sorting, and packages them into a startup optimization instruction sequence; A task scheduling and control module, which scans the GPU task set, identifies resource occupancy, queries the GPU resource demand profile library for prediction, collects switching records, calculates the switching time, compares task priorities, calculates the preemption cost, sets the competition threshold, determines the preemption timing, and forms a dynamic resource scheduling scheme.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: The present invention collects shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting logs during the startup phase of an application, and constructs a multi-task GPU activity feature set, which can quantify the dynamic requirements of different application startup behaviors for hardware resources and achieve accurate prediction of resource occupancy patterns. Based on numerical clustering, the task running states are grouped and a classification index is established. After aggregating the characteristic data of similar tasks, a typical GPU resource demand time series profile is generated, enabling subsequent resource pre-allocation strategies to be adaptively adjusted according to historical behavior patterns and reducing the risk of video memory bandwidth contention in high-fluctuation scenarios. According to real-time startup requests, a matching time series profile is retrieved, shader identifiers and video memory allocation parameters are extracted to generate pre-compilation instructions and a video memory pre-allocation instruction sequence, and the static allocation of key resources is completed in advance to avoid compilation delays and video memory fragmentation problems caused by dynamic allocation during runtime. Combining the task preemption priority and cost evaluation table, by setting dynamic trigger conditions and vertical blanking period time points, it is ensured that high-priority tasks preempt resources at the time when the video memory release is sufficient and the context switching cost is the lowest, reducing the sudden drop in rendering frame rate or computational pipeline blockage caused by preemption. Through the collaborative optimization of data-driven feature modeling and preemptive scheduling, the traditional static resource allocation mode is upgraded to a dynamic regulation mechanism based on historical load prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] Please refer to Figure 1 , the present invention provides a technical solution for optimizing and improving the performance of a graphics card based on big data, including the following steps: Collect shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, perform data cleaning, and establish an application startup and multi-task GPU activity feature set; Based on the application startup and multi-task GPU activity feature set, through numerical clustering, group the application startup behaviors and task running states, obtain the category labels to which each application or task belongs, establish a GPU activity mode classification index, and based on the GPU activity mode classification index, aggregate the corresponding feature set data for each category label, calculate the average value of the peak demand of GPU compilation units within the category and the video memory bandwidth, and construct a typical GPU resource demand time series profile; According to the received application startup request, a typical GPU resource requirement timing profile of the matching application is retrieved, a shader identifier list of the startup call and a predicted initial video memory allocation mode parameter value are extracted, a target application startup resource requirement list is generated, the shader identifier in the target application startup resource requirement list is converted into a pre-compiled instruction, the video memory allocation mode parameter is converted into a video memory pre-allocation instruction, and a GPU pre-processing instruction sequence is established; Monitor the task set currently running on the GPU, retrieve the typical GPU resource demand timing profile and associated context switching overhead corresponding to each task, generate a task preemption priority and cost evaluation table, set the preemption trigger conditions and time points based on the task preemption priority and cost evaluation table, and obtain a dynamic GPU task preemption control parameter set.
[0018] The steps for obtaining the application startup and multi-tasking GPU activity feature set are: Collect shader compile time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, calculate the 25th percentile Q1 and 75th percentile Q3 of the shader compile time data, remove extreme values lower than Q1-1.5×(Q3-Q1) or higher than Q3+1.5×(Q3-Q1), and generate cleaned shader compile time data sets, initial video memory allocation mode parameter sets, and pipeline state setting sequence log sets; Based on the cleaned shader compilation time dataset, the compilation duration recorded at all time points is traversed, the maximum value is selected as the compilation peak duration, the video memory block size field is parsed from the initial video memory allocation mode parameter set, and the fixed allocation unit value of the video memory block size is extracted. The number of state switching instructions per second is counted from the pipeline state setting sequence log set as the state switching frequency, the average of the absolute value of the difference between adjacent state switching timestamps is calculated as the switching delay, and the total number of shader instruction executions per second is counted as the computational intensity; The compilation peak duration, video memory block size, state switching frequency, computational intensity, and switching delay are aligned by timestamp and merged into a multidimensional vector at the same time point. The time series is divided into time windows of 30 seconds to establish a feature set of application startup and multi-tasking GPU activities.
[0019] Specifically, based on the collected data of the target application, such as the software named "Graphic Editor v2.1", including shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase, first process the shader compilation time data. For example, the sequence of 15 startup compilation times (unit: milliseconds) collected is [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60, 150]. Sort this sequence to get [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60, 150]. Calculate the 25th percentile Q1, whose position is , and the corresponding value is 30ms. Calculate the 75th percentile Q3, whose position is , and the corresponding value is 50ms. Calculate the interquartile range IQR = Q3 - Q1 = 50 - 30 = 20ms. Determine the lower bound for outlier judgment as Q1 - 1.5 IQR = 30 - 1.5 × 20 = 0ms, and the upper bound is Q3 + 1.5 × IQR = 50 + 1.5 × 20 = 80ms. Consider the values in the original data that are lower than 0ms or higher than 80ms as extreme values. In this example, 150ms is identified as an extreme value and removed. Perform a similar data cleaning process on the initial video memory allocation mode parameter set and the pipeline state setting sequence log set. For example, if there are obviously unreasonable large or zero-value allocation requests in the video memory allocation records, or if there are timestamp errors or instruction out-of-order in the state setting logs, they are also removed according to statistical rules or logical rules. Finally, obtain the cleaned data set, and generate the cleaned shader compilation time data set [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60], the corresponding initial video memory allocation mode parameter set, and the pipeline state setting sequence log set.
[0020] Using the generated cleaned shader compilation time data set [15, 25, 28, 30, 32, 35, 38, 40, 42, 45, 48, 50, 55, 60] (unit: ms), traverse and find the maximum value in it to determine the compilation peak duration as 60ms. Then, parse the recorded video memory block size field from the cleaned initial video memory allocation mode parameter set. For example, the record is AllocInfo(BlockID = 001, Size = 16777216, Type = Texture). Extract the Size value. Observe that there are fixed-size units in multiple allocation requests, such as repeated 16MB allocations. Select this common and representative value as the fixed allocation unit value of the video memory block size. , namely Bytes. Then, analyze the pipeline state setting sequence log set after cleaning, and count the number of pipeline state switching instructions (such as SetPipelineStateObject, IASetPrimitiveTopology, etc.) that occur within a specific time window, for example, within the 5th second. If the statistical result is 85 times, then the state switching frequency for this second is 85Hz. Further calculate the absolute value of the timestamp difference between adjacent state switching instructions within this second. For example, if the recorded timestamp sequence is [5.012s, 5.020s, 5.025s,..., 5.980s], calculate the adjacent differences |5.020 - 5.012| = 0.008s, |5.025 - 5.020| = 0.005s,... After summing up all the differences and dividing by the number of differences (i.e., the number of switches minus 1), the switching delay is obtained , for example, the calculation result is 0.011s (11ms). Finally, count the total number of shader instructions executed within the same time window (the 5th second), for example, obtained through GPU performance counters or driver logs. If the total number is 2.5 GigaInstructions, then the computational intensity is 2.5 GigaInstructions / s. Align these eigenvalue features obtained or associated at the 5th second according to the timestamp, and compile the peak duration Use the globally calculated 60ms, and use the values at this time point for the others, and merge them into a multi-dimensional vector V(t = 5s) = [60ms, 16777216 Bytes, 85Hz, 2.5GI / s, 11ms]. Perform this operation for each second during the application startup phase (for example, lasting 60 seconds), generating 60 such multi-dimensional vectors. Then, divide them into time windows of 30 seconds continuously. For example, window 1 contains the vector sequence from the 1st to the 30th second, window 2 contains the vector sequence from the 2nd to the 31st second (using a sliding window), or window 1 contains 1 - 30 seconds, window 2 contains 31 - 60 seconds (non-overlapping window). Here, a 30 - second non - overlapping window is used, then two time series segments are generated, and these time series segments together constitute the application startup and multi - task GPU activity feature set.
[0021] The steps to obtain the GPU activity mode classification index are as follows: Based on the application startup and multi - task GPU activity feature set, select three feature dimensions of compile peak duration, video memory block size, and state switching frequency as the clustering input, perform Z - score standardization on each data point, and generate a standardized feature subset; Initialize k cluster centers according to the standardized feature subset. The value of k is determined by the elbow method. Calculate the Euclidean distance from each data point to each cluster center, and assign the data point to the nearest neighbor cluster center. Iteratively update the coordinates of the cluster center to the mean of the data points it belongs to until the change in the center coordinates between two consecutive iterations is less than the threshold, generating an optimized set of cluster centers; Based on the optimized set of cluster centers, assign the corresponding cluster labels to the time windows of each application or task, establish a mapping table, aggregate the mean peak demand of compilation units and the temporal pattern of video memory bandwidth of the same category according to the labels, and generate a GPU activity pattern classification index.
[0022] Specifically, based on the established feature set of application startup and multi-task GPU activities, this feature set contains data of multiple time windows, and each window is a multi-dimensional vector sequence. For example, during the startup process of "Graphics Editor v2.1", two 30-second time window feature sequences are generated. At the same time, the startup or running characteristics of other applications or tasks (such as "Video Player v1.0", "Background Rendering Task") may also be analyzed to form a larger feature set, from which the compilation peak duration (ms), video memory block size (Bytes), and status switching frequency (Hz) are selected as the input for clustering analysis. For example, there are 100 data points from different applications / tasks / time windows, and each point is a three-dimensional vector , , . For example, point P1 = [60, 16777216, 85], P2 = [40, 8388608, 50], P3 = [100, 33554432, 120]. Calculate the mean and standard deviation for these three feature dimensions on all 100 data points respectively. For example of , , of , , of , . Perform Z-score standardization on each data point, and the calculation formula is . For example, after standardization, P1 is , This operation is performed on all 100 points to generate a standardized feature subset. Next, the elbow method is used to determine the number of clusters k. Calculate the sum of squared errors (SSE) within the clusters when k = 2, 3, 4, 5, etc. Plot the relationship between the k values and the SSE, and observe the "elbow" of the graph, that is, the turning point where the rate of decrease of the SSE changes from fast to slow. For example, if this point appears when k = 4, determine the optimal number of clusters to be 4. Initialize 4 cluster centers. For example, 4 standardized data points can be randomly selected as the initial centers C1, C2, C3, C4, and enter the iterative process. Calculate the Euclidean distance from each standardized data point (such as ) to the current 4 cluster centers (C1, C2, C3, C4). For example: , Assign to the cluster center with the closest distance. For example, if the closest one is C2, then belongs to cluster 2. Perform this assignment for all 100 data points. After the assignment is completed, recalculate the positions of each cluster center. The new center coordinates are the means of the respective dimensional coordinates of all the data points it contains. For example, for example, cluster 2 now contains , then the new coordinates of C2 are ([[]] , , ). Update all 4 cluster centers. Compare the change amounts of the old and new center coordinates. For example, calculate the Euclidean distance between the old and new coordinates of each center, or the sum of the change amounts of all center coordinates. Set a threshold, which is set according to the accuracy requirements of the clustering result. For example, if the numerical range after feature standardization is roughly between -3 and +3, the threshold can be set to 0.001. If the sum of the change amounts of all center coordinates is less than 0.001, it is considered that the cluster centers have stabilized and the iteration stops. If it is greater than the threshold, repeat the assignment and update steps until the stop condition is met. After the iteration stops, obtain an optimized set of 4 cluster centers {C1', C2', C3', C4'}. Based on the final data point assignment results, assign the cluster label (1, 2, 3, or 4) to each time window of each application or task. Establish a mapping table, for example, record <application name: "Graph Editor v2.1", time window: 1, cluster label: 2>, <application name: "Video Player v1.0", time window: 1, cluster label: 1>, etc. Then, aggregate the data according to the cluster labels, calculate the mean of the peak demand of the compilation units for all time windows under each cluster category (such as label 2) (for example, this data is in the feature set or can be obtained by associating with the original log), and analyze the temporal pattern of the video memory bandwidth corresponding to these time windows (for example, calculate statistical quantities such as the bandwidth mean, peak, variance, etc. or extract certain temporal features). Associate the cluster labels with the corresponding mean of the peak demand of the compilation units and the temporal pattern of the video memory bandwidth to generate a GPU activity pattern classification index.
[0023] The steps to obtain the typical GPU resource demand time series profile are as follows: Based on the GPU activity pattern classification index, traverse all time windows under each category label, extract the peak demand of GPU compilation units within each time window, calculate the average value of the peak demand of compilation units in all time windows, and generate a set of average values of peak demand of compilation units within the category; According to the set of average values of peak demand of compilation units within the category, extract the memory bandwidth time series data corresponding to the category label, and calculate the memory bandwidth fluctuation intensity. The formula is: ; where, is the memory bandwidth fluctuation intensity, is the memory bandwidth value at time point , is the bandwidth average value, is the total number of seconds in the time window; Based on the memory bandwidth fluctuation intensity, parse the call timestamps of the texture mapping unit and the rasterization unit from the pipeline state setting sequence log set, count the call frequency of each unit within the time window, calculate the median and variance of the call frequency, and splice the average value of the peak demand of compilation units, the memory bandwidth fluctuation intensity, the median and variance of the call frequency into a multi-dimensional vector in the order of time windows to construct the typical GPU resource demand time series profile.
[0024] Specifically, based on the generated GPU activity pattern classification index, this index maps the time windows of an application or task to different activity pattern categories (cluster labels) and associates some features of each category. For example, category 2 is associated with an average peak demand of 150 compilation units and a specific memory bandwidth time series pattern. Now, it is necessary to construct a more detailed resource demand profile for each category. Traverse all time window instances under one category label, such as label 2 (for example, 20 time windows from different application startup processes are classified into this category), and extract the peak demand of GPU compilation units recorded in each time window (this requires the data to be collected early and included in the feature set or can be retrieved through associated queries). For example, the peak compilation demands of these 20 windows are [140, 155, 150, 160, 145,..., 152] respectively. Calculate the average value of these 20 values to obtain the average value of the peak demand of compilation units for this category (label 2) units, and generate a set containing all categories and their corresponding average values of peak demand of compilation units. Next, according to category label 2, extract the memory bandwidth time series data of all 20 associated time windows , each window contains bandwidth values at 30 time points (seconds). For example, the bandwidth data (GB / s) of one window is [50, 55, 45, 60, 50,..., 58]. Calculate the average bandwidth of this window GB / s, and then calculate the video memory bandwidth fluctuation intensity of this window . The formula is , where is the video memory bandwidth fluctuation intensity, is the video memory bandwidth value at time point , is the average bandwidth (53 GB / s) of this window, is the total number of seconds in the time window (30 s), and the specific calculation is . For example, the calculated result is GB / s. Calculate the value for all 20 windows belonging to category 2, and then calculate the average of these 20 values to obtain the average bandwidth fluctuation intensity of category 2 GB / s (simplified calculation here, actually the P of each window should be calculated and then averaged). Next, based on the 20 time windows associated with category 2, parse the timestamps of the instructions or events related to the texture mapping unit (TMU) and rasterization unit (ROP) calls from the corresponding original pipeline state setting sequence log set, and count the call frequencies of each unit in each 30 - second time window. For example, for one window, the TMU call count sequence (per second) is [10k, 12k, 9k,..., 11k], and the ROP call count sequence is [5k, 6k, 4k,..., 5.5k]. Calculate the median and variance of these two sequences. For example, the median of the TMU call frequency is 10.5k times per second, and the variance is 4M( ), the median of the ROP call frequency is 5.2k times per second, and the variance is 1.5M( ). Calculate these statistics for all 20 windows of category 2, and then take the average to obtain the median and variance of the TMU call frequency of category 2, as well as the median and variance , finally, the average value of the peak demand for compilation units (150 units), the average value of the fluctuation intensity of the video memory bandwidth (6.5 GB / s), the median of the TMU call frequency (10.5 k / s), the variance of the TMU call frequency (4 M), the median of the ROP call frequency (5.2 k / s), and the variance of the ROP call frequency (1.5 M) calculated for this category (label 2) are concatenated in a fixed order to form a multi-dimensional vector [150, 6.5, 10500, 4000000, 5200, 1500000]. This vector represents the typical GPU resource demand time series profile for category 2 (here the profile is aggregated into a single vector, and the original text means the time series profile, and it may be necessary to retain the time series information or more complex representation within the window). The same operation is performed on all category labels (1, 3, 4) to construct the typical GPU resource demand time series profiles for all categories.
[0025] Formula Detailed description: Formula Used to calculate the fluctuation intensity of the video memory bandwidth.
[0026] Parameter description: : Represents the fluctuation intensity of the video memory bandwidth within a time window.
[0027] : Represents the total duration of the time window, which is set to 30 seconds in this example.
[0028] : Represents the discrete time points within the time window, counting from 1 to .
[0029] : Represents the sum of all time points from the 1st second to the th second within the time window.
[0030] : Represents at time point The video memory bandwidth value monitored. This value is obtained by real-time collection through GPU performance monitoring tools (such as NVIDIA's nvml library interface or AMD's rocmsmi tool) during the application's runtime. For example, at the 5th second, GB / s is collected.
[0031] : Represents within the entire time window The average value of the video memory bandwidth, calculated as . This value represents the average bandwidth level within the window period.
[0032] : Represents at time point The bandwidth value Average bandwidth of window This measures the degree to which the bandwidth value at that point in time deviates from the average level.
[0033] : Indicates dividing the sum by the total number of seconds in the time window to get the average absolute deviation.
[0034] Operation logic and purpose: This formula calculates the average of the absolute deviations between the memory bandwidth value and its average value at each time point in the time window. First, calculate the average bandwidth of the entire window. Then, for each second in the window, calculate the bandwidth of that second With average bandwidth The absolute value of the difference is calculated by summing up the absolute difference of all seconds and finally dividing it by the total number of seconds in the window. This result Quantifies the average amplitude of bandwidth fluctuations around its mean value within the time window. A larger value indicates more drastic and unstable bandwidth fluctuations; a smaller value indicates more stable bandwidth usage.
[0035] Calculation example: For example, a Bandwidth data collected in a time window of seconds (GB / s) are [50, 55, 45, 60, 50] respectively.
[0036] Calculating Average Bandwidth : GB / s.
[0037] Calculate the absolute deviation per second: t=1:|50-52|=|-2|=2; t=2:|55-52|=|3|=3; t=3:|45-52|=|-7|=7; t=4:|60-52|=|8|=8; t=5:|50-52|=|-2|=2; Calculating Wave Strength : GB / s.
[0038] The formula is useful: by calculating , it is possible to distinguish scenarios with similar total bandwidth consumption but different behavior patterns. For example, a task that runs continuously at 50GB / s and another task that frequently switches between 25GB / s and 75GB / s may have similar average bandwidths, but the latter may have The value will be significantly higher, revealing a stronger impact on bandwidth resources, which is valuable for resource scheduling and conflict prediction.
[0039] Calculated GB / s indicates that within this 5-second window, the video memory bandwidth deviates from the mean by approximately 4.4 GB / s per second, about 52 GB / s. This value will be used as part of the typical GPU resource demand time series profile to characterize the bandwidth usage stability characteristics of the corresponding activity pattern category. A higher value may mean that tasks of this category are more likely to trigger bandwidth contention.
[0040] The steps to obtain the target application startup resource demand list are as follows: Receive an application startup request, parse the application unique identifier and version number in the request header, traverse all category labels in the GPU activity pattern classification index, compare the application identifier with the features under each category label one by one, filter out the category label with the highest feature matching degree, and extract the typical GPU resource demand time series profile corresponding to the category label; Based on the typical GPU resource demand time series profile data block, extract all shader identifiers, video memory block sizes, and virtual address offset base addresses from the typical GPU resource demand time series profile data block to generate an intermediate data set; According to the video memory block sizes in the intermediate data set, sort the shader identifiers in descending order, and encapsulate the sorted shader identifiers with the corresponding video memory block sizes and virtual address offset base addresses into the target application startup resource demand list.
[0041] Specifically, when a request to start the application "Graphical Editor v2.1" with version number "v2.1" is received, first parse the request to obtain the unique application identifier "Graphical Editor v2.1". Then, traverse all the category tags {1, 2, 3, 4} and their associated features recorded in the established GPU activity pattern classification index. It is necessary to compare the known startup features of "Graphical Editor v2.1" (for example, the typical compilation peak duration, video memory allocation size, state switching frequency, etc. obtained by pre-analysis when the application starts, such as [60ms, 16MB, 85Hz]) with the features represented by each category tag (usually the cluster center or the average feature of the category. For example, the central feature of category 2 is close to [58ms, 17MB, 80Hz], and category 1 is [30ms, 8MB, 40Hz]), and calculate "Graphical Editor v2.The similarity or distance between the "1" feature and the features of each category, for example, using the Euclidean distance, it is calculated that its feature distance from Category 2 is the smallest, and the matching degree is determined to be the highest. Therefore, Category Label 2 is selected. Then, according to the matched Category Label 2, the corresponding typical GPU resource demand time series profile data block constructed is extracted. This data block contains information such as the average value of the peak demand of the compilation unit, the bandwidth fluctuation intensity, the TMU / ROP call statistics, etc. At the same time, the specific resource items required in the typical startup scenario of this category also need to be extracted from the more detailed data associated with Category 2 (which may be stored in the index or the original data is retrieved through index backtracking), including all relevant shader identifiers (such as "Shader_UI_Main", "Shader_Render_Core", "Shader_PostFX"), their corresponding video memory block sizes (such as 8MB, 16MB, 4MB), and the recommended virtual address offset base addresses for allocation (such as 0x10000000, 0x20000000, 0x30000000). These extracted {shader identifier, video memory block size, virtual address offset base address} entries are aggregated to generate an intermediate data set, for example, containing a list: [{ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}]. According to the video memory block size Size in the intermediate data set, the entries in the list are sorted in descending order. After sorting, it is obtained: [{ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}]. This sorted list containing shader identifiers, corresponding video memory block sizes, and virtual address offset base addresses is encapsulated to form the resource demand list required for the startup of the target application "Graphics Editor v2.1".
[0042] The steps to obtain the GPU preprocessing instruction sequence are as follows: Parse the startup resource demand list of the target application, traverse the key-value pairs of each entry in the startup resource demand list of the target application, and extract the shader identifier, video memory block size, and virtual address offset base address; Based on the shader identifier, each shader identifier is replaced with an instruction string according to the pre-compilation instruction template. Based on the video memory block size and the virtual address offset base address, each video memory block size and virtual address offset base address are replaced with an instruction string according to the video memory pre-allocation instruction template, generating a GPU preprocessing instruction sequence.
[0043] Specifically, parse the generated target application startup resource requirement list. This list is an ordered list. For example: [{ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}]. Traverse each entry in this list. Each entry is a data structure containing key-value pairs. For the first entry {ID: "Shader_Render_Core", Size: 16MB, VA: 0x20000000}, extract the shader identifier "Shader_Render_Core", the video memory block size of 16MB (i.e., 16777216 Bytes), and the virtual address offset base address of 0x20000000. Based on the extracted shader identifier "Shader_Render_Core", use the predefined pre-compilation instruction template. For example, the template is CMD_PRECOMPILE_SHADERName= <shaderid>, will <shaderid>Replace it with "Shader_Render_Core" to generate the instruction string "CMD_PRECOMPILE_SHADERName=Shader_Render_Core". Similarly, based on the extracted video memory block size of 16777216 Bytes and the virtual address offset base address of 0x20000000, use the predefined video memory pre-allocation instruction template, such as the template CMD_PREALLOCATE_MEMORYSize= <sizeinbytes>Address= <virtualaddress>, will <sizeinbytes>Replace with 16777216, and <virtualaddress>Replace it with 0x20000000, generate the instruction string "CMD_PREALLOCATE_MEMORY Size=16777216 Address=0x20000000", perform the same operation on the second entry in the list {ID: "Shader_UI_Main", Size: 8MB, VA: 0x10000000}, generate the instructions "CMD_PRECOMPILE_SHADER Name=Shader_UI_Main" and "CMD_PREALLOCATE_MEMORY Size=8388608 Address=0x10000000", perform the same operation on the third entry {ID: "Shader_PostFX", Size: 4MB, VA: 0x30000000}, generate the instructions "CMD_PRECOMPILE_SHADER Name=Shader_PostFX" and "CMD_PREALLOCATE_MEMORY Size=4194304 Address=0x30000000", and combine all the instruction strings generated in the list order (or other predefined logic, such as allocate first and then compile) to form the final GPU preprocessing instruction sequence.
[0044] The steps to obtain the task preemption priority and cost evaluation table are as follows: Monitor the set of tasks currently running on the GPU, obtain the process identifier of each task and the associated typical GPU resource demand timing profile through the operating system kernel interface, extract the time consumption of the last three context switches of each task from the system scheduler log, calculate the average value of the time consumption of the last three context switches as the context switch overhead, and generate a data set; Based on the data set, calculate the task preemption priority score, and the formula is: ; Among them, is the task preemption priority score, is the task the time consumption of the context switch, is the task video memory bandwidth occupancy rate, is the foreground window activation status, 1 for activated and 0 for not activated; According to the task preemption priority score, sort all tasks from high to low, generate an entry for each task, and the entry includes the process identifier, video memory bandwidth occupancy rate, average switch time consumption, activation status, and task preemption priority score, and construct the task preemption priority and cost evaluation table.
[0045] Specifically, monitor the set of tasks running on the current GPU. For example, if it is found that there are three tasks: Task A (game rendering), Task B (video decoding), and Task C (scientific computing), obtain their process identifiers (PIDs), which are PID1234, PID5678, and PID9012 respectively, by calling the operating system kernel interface (such as the / proc file system in Linux or the Process Explorer API in Windows), and query the typical GPU resource demand time series profiles associated with these PIDs (for example, those generated and stored in the previous steps for these tasks or their categories). Then, find the last three context switch records related to these three PIDs from the system scheduler logs (such as GPU driver logs or operating system scheduler logs), and extract the time consumption of each switch. For example, the last three switch time consumptions of Task A (PID1234) are [4.5ms, 5.0ms, 4.8ms], those of Task B (PID5678) are [12.0ms, 11.5ms, 12.5ms], and those of Task C (PID9012) are [2.0ms, 2.2ms, 2.1ms]. Calculate the average of the last three context switch time consumptions of each task as its context switch overhead , Task A: ms, Task B: ms, Task C: ms. At the same time, obtain the video memory bandwidth occupancy rate of each current task through a GPU monitoring tool , for example, Task A is 70%, Task B is 15%, and Task C is 30%, and query whether the windows associated with each task are foreground active windows through the window manager API , for example, Task A is the current active game window( ), Task B and C are background tasks( ). Integrate the collected and calculated data to form a data set containing {PID, , , }. Based on this set, calculate the preemption priority score of each task , using the formula , where is the time consumption of the k-th switch, is the calculated previously, is a percentage value, is 0 or 1
[0046] Task A (PID1234): , Task B (PID5678): , Task C (PID9012): , According to the calculated task preemption priority score , sort all tasks from high to low: Task A (76.32) > Task B (30.0) > Task C (8.4). Generate an entry for each task, including the process identifier, video memory bandwidth occupancy rate, average switching time, activation status, and the calculated task preemption priority score, and construct a task preemption priority and cost evaluation table.
[0047] Table 1: Task Preemption Priority and Cost Evaluation Table Process identifier Video memory bandwidth occupancy rate (%) Average switching time (ms) Activation status Task preemption priority score 1234 70 4.77 1 76.32 5678 15 12.00 0 30.00 9012 30 2.10 0 8.40
[0048] As shown in Table 1, this table shows the evaluation results of the tasks running on the current GPU, sorted in descending order according to the task preemption priority score. The higher the score, the more important the task or the higher the preemption cost, and its operation should be guaranteed first.
[0049] Formula Detailed description: Formula is used to calculate the task preemption priority score.
[0050] Parameter description: : The task preemption priority score of task . This is a comprehensive score. The higher the value, the less likely the task should be preempted (i.e., the higher the priority or the greater the preemption cost).
[0051] : Represents a specific GPU task.
[0052] : Represents the index of the number of context switches. Here, the last three times are taken (k = 1, 2, 3).
[0053] : Represents the sum of the context switch times for the last three times.
[0054] : The th recorded context switch time of task . This data is obtained from the system scheduler or GPU driver logs. For example, for task A, ms, ms, ms.
[0055] : Calculate the average context switch time for the last three times of task , that is, This value reflects the direct time cost of switching this task.
[0056] : Task The current video memory bandwidth occupancy rate, given as a percentage (e.g., 70%). This data is obtained by real-time monitoring of GPU performance counters.
[0057] : This is a weight factor based on the video memory bandwidth occupancy rate. The higher the occupancy rate, the larger this factor, thus increasing the score. Dividing by 10 is to adjust the magnitude of its influence, and adding 1 ensures that the factor is at least 1.
[0058] : Task 's window activation status. If the window associated with the task is the foreground window for the current user interaction, then ; if it is a background window, then . This status is obtained by querying the operating system window manager.
[0059] : This is a weight factor based on the activation status. For foreground tasks ( ), this factor is 2; for background tasks ( ), this factor is 1. This doubles the score of foreground tasks directly.
[0060] Operation logic and purpose: This formula evaluates the preemption priority of a task (or the "cost" of being preempted) by combining three aspects of factors: Switching cost ( ): The longer the switching time, the higher the base score.
[0061] Resource occupancy ( ): The higher the video memory bandwidth occupancy rate, the larger the coefficient multiplied, and the higher the score. This reflects that high-bandwidth tasks may be more important or have a greater impact on interruptions.
[0062] User interaction ( ): Foreground tasks directly obtain double the weight, reflecting the priority guarantee for the user experience.
[0063] Multiplying the three gives the final score . Its purpose is to quantitatively evaluate the "value" of each current task not being preempted, providing a basis for subsequent preemption decisions. Tasks with high scores are objects that the system tends to protect and not preempt easily.
[0064] Example: Task A (PID1234): ms, , 。
[0065] ; ; ; ; ; Benefits of the formula: By comprehensively considering the switching overhead of the task itself, the occupancy of critical resources (video memory bandwidth), and the interaction status with the user, the formula provides a relatively comprehensive task importance evaluation index, which is superior to the judgment based on a single factor (such as priority or resource occupancy), and helps to make more intelligent and balanced preemption decisions.
[0066] Numerical result correlation: The calculated , , These scoring values are directly used to construct a task preemption priority and cost evaluation table (as shown in Table 1). In subsequent decisions, these scores will be used to identify which tasks are high-priority (high scores, not easily preempted) and low-priority (low scores, can be considered as preemption targets).
[0067] The steps to obtain the dynamic GPU task preemption control parameter set are as follows: Based on the task preemption priority and cost evaluation table, screen out the tasks with the top 20% of the task preemption priority scores as the high-priority task set, extract the average video memory bandwidth occupancy rate and the response delay requirement of the high-priority task set, and at the same time extract the top 30% of the tasks with the highest average switching time among the remaining tasks as the low-priority task set, generating a high-priority response requirement data set and a low-priority switching cost data set; According to the high-priority response requirement data set and the low-priority switching cost data set, calculate the preemption trigger threshold, and the formula is: ; where, is the preemption trigger threshold, is the average video memory bandwidth occupancy rate of high-priority tasks, is the maximum allowable response delay of high-priority tasks, is the average switching time of low-priority tasks, is the available video memory ratio of the current GPU; Based on the preemption trigger threshold, set the trigger condition as follows: when the predicted response delay of high-priority tasks exceeds the maximum allowable response delay of high-priority tasks, and the preemption trigger threshold is greater than 100, calculate the preemption time point as the start time of the next vertical blanking period , and the formula is , For the current frame number, obtain the preemption control parameter set.
[0068] Specifically, based on the constructed task preemption priority and cost evaluation table (see Table 1), first filter the task preemption priority scores The top 20% of the tasks are ranked as the high-priority task set. Currently, there are 3 tasks, and 20% is 0.6, which is rounded up to 1. Therefore, the task A with the highest score (PID1234, ) is selected as the high-priority task set {TaskA}, and the average value of the video memory bandwidth occupancy rate of the tasks in this set is extracted . Since there is only one task, the average value is its occupancy rate of 70%. At the same time, obtain its response latency requirement , which is usually determined by the application type. For example, for the game task A, the required response latency does not exceed 1 frame time, such as 16.67 ms (corresponding to a 60 Hz refresh rate). Set to 16.67 ms = 0.01667 s to form the high-priority response requirement data set { , }. Then, from the remaining tasks (Task B, Task C), according to their average switching time ) [12.0 ms, 2.1 ms], filter out the highest 30%. 30% of 2 tasks is 0.6, which is rounded up to 1. The task with the highest switching time is Task B (12.0 ms). Therefore, select Task B (PID5678) as the representative low-priority task (for cost evaluation), and extract its average switching time ms = 0.012 s to form the low-priority switching cost data set { }. According to these two data sets, calculate the preemption trigger threshold , using the formula , where is the average value of the bandwidth occupancy rate of high-priority tasks (using the percentage value 70), is the maximum allowable response latency of high-priority tasks (0.01667 s), is the average switching time of low-priority tasks (0.012 s), is the proportion of the currently available video memory of the GPU, which needs to be monitored and obtained in real time. For example, if the currently available video memory is 30% of the total video memory, then , substitute the values into the calculation: , Next, set the preemption trigger condition, and the condition is: when the predicted value of the response latency of the high-priority task (Task A) is monitored Exceeds its maximum allowable response latency (predicted by a performance model or inferred from recent historical data, e.g., predicted to be 18 ms = 0.018 s). (0.01667 s), and the calculated preemption trigger threshold (53.26) is greater than a preset reference value, e.g., 100. The judgment conditions are: ( ) and ( ), i.e., (0.018 s > 0.01667 s) and (53.26 > 100). The first condition is true, and the second condition is false. Therefore, the overall trigger condition is false, and preemption is not triggered. [Adjust the example to trigger preemption] For example, the available video memory is extremely low, , and the switching cost of low-priority tasks is very low, ms = 0.002 s. Then , at this time, the conditions become (0.018 s > 0.01667 s) and (130.46 > 100). Both conditions are true, preemption is triggered, and the time point for preemption execution is calculated , which is set to the start time of the next vertical blanking period. If the current monitor refresh rate is 60 Hz, the V-blank period is approximately 16.67 ms, and the dynamic GPU task preemption control parameter set is obtained.
[0069] Formula Detailed description: Formula Used to calculate the dynamic preemption trigger threshold.
[0070] Parameter description: : The preemption trigger threshold, a unitless value used to determine whether a preemption operation should be performed.
[0071] : The average video memory bandwidth occupancy rate of the high-priority task set, substituted for calculation as a percentage value (e.g., 70). This value reflects the demand intensity of high-priority tasks for bandwidth resources. It is obtained by screening out the tasks with the highest scores (e.g., the top 20%) and calculating the average of their current bandwidth occupancies.
[0072] : The maximum allowable response latency of high-priority tasks. This is a quality of service (QoS) requirement, usually set according to the application type (such as interactive applications, games). For example, for a 60 FPS game, it can be set to 1 / 60 0.01667 s.
[0073] : The average context switch latency of the low-priority task set (e.g., the 30% tasks with the highest switching costs). This value represents the time cost required to preempt these low-priority tasks. It is obtained by filtering out tasks with higher switching costs and calculating their average switching latency.
[0074] : The proportion of the currently available GPU video memory, which is a value between 0 and 1 (e.g., 0.3 indicates 30% available). This value is obtained by querying the GPU status in real time and reflects the tightness of the current video memory resources.
[0075] : The square root of the available video memory proportion. This factor adjusts the sensitivity of the threshold to video memory availability.
[0076] Operation logic and purpose: This formula aims to establish a dynamic preemption decision threshold, which weighs several key factors: The urgency of high-priority tasks ( ): The high bandwidth occupancy rate and strict latency requirements (small ) are multiplied to obtain a metric reflecting the potential performance risk of high-priority tasks. The larger this value, the more likely high-priority tasks are to miss their deadlines due to resource shortages, so it tends to increase the threshold , making preemption more likely to occur.
[0077] Preemption cost ( ): The average switching cost of low-priority tasks appears in the denominator. The higher the cost, the lower the threshold , making preemption less likely to occur because the preemption itself is costly.
[0078] System resource status ( ): The available video memory proportion affects the threshold in the form of a square root. When there is more available video memory ( is larger), is also larger, which will increase the threshold , meaning that the system resources are relatively abundant and the need for preemption is not so urgent. Conversely, when there is very little available video memory ( is close to 0), is also close to 0, which will significantly reduce the threshold , making preemption more likely to be triggered even if other factors remain unchanged, reflecting that more active scheduling intervention is needed when resources are tight.
[0079] Generally speaking, the higher the value, the stronger the reason the system deems is needed to trigger preemption (e.g., the latency prediction of high-priority tasks far exceeds their requirements).
[0080] Calculation example: , , , .
[0081] ; ; ; ; ; Benefits of the formula: The formula dynamically adapts to the system state (high-priority task requirements, preemption cost, memory pressure), enabling the preemption decision to be adjusted according to the current specific situation rather than based on fixed rules, which helps to achieve a better balance between ensuring the performance of high-priority tasks and avoiding unnecessary preemption overhead.
[0082] Numerical result correlation: The calculated . This value will be compared with a fixed reference value (e.g., 100). In this example, . Combining with the condition of the predicted delay of high-priority tasks exceeding the limit ( ), it jointly determines whether to trigger preemption. This value itself is not directly part of the preemption control parameter set, but it is a key intermediate calculation result for generating this parameter set (especially the trigger state). The high value combined with the delay exceeding the limit constitutes a sufficient condition for executing preemption.
[0083] The above is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any person skilled in the relevant art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.< / virtualaddress> < / sizeinbytes> < / virtualaddress> < / sizeinbytes> < / shaderid> < / shaderid>
Claims
1. A method for optimizing and improving graphics card performance based on big data, characterized in that: The following steps are involved: Collect shader compilation time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, perform data cleaning, and establish a feature set of application startup and multi-tasking GPU activities; Based on the application startup and multi-task GPU activity feature set, the application startup behavior and task running status are grouped by numerical clustering, the category label of each application or task is obtained, and a GPU activity mode classification index is established. Based on the GPU activity mode classification index, the corresponding feature set data is aggregated for each category label, the average value of the peak demand of the GPU compilation unit in the category and the video memory bandwidth are calculated, and a typical GPU resource demand timing profile is constructed; According to the received application startup request, a typical GPU resource requirement timing profile of the matching application is retrieved, a shader identifier list of the startup call and a predicted initial video memory allocation mode parameter value are extracted, a target application startup resource requirement list is generated, the shader identifier in the target application startup resource requirement list is converted into a pre-compiled instruction, the video memory allocation mode parameter is converted into a video memory pre-allocation instruction, and a GPU pre-processing instruction sequence is established; Monitor the task set currently running on the GPU, retrieve the typical GPU resource demand timing profile and the associated context switching overhead corresponding to each task, generate a task preemption priority and cost evaluation table, set the preemption trigger conditions and time points based on the task preemption priority and cost evaluation table, and obtain a dynamic GPU task preemption control parameter set.
2. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps of obtaining the application startup and multi-task GPU activity feature set are as follows: Collect shader compile time data, initial video memory allocation mode parameters, and pipeline state setting sequence logs during the startup phase of the target application, calculate the 25th percentile Q1 and 75th percentile Q3 of the shader compile time data, remove extreme values lower than Q1-1.5×(Q3-Q1) or higher than Q3+1.5×(Q3-Q1), and generate cleaned shader compile time data sets, initial video memory allocation mode parameter sets, and pipeline state setting sequence log sets; Based on the cleaned shader compilation time data set, traverse the compilation durations recorded at all time points, select the maximum value as the compilation peak duration, parse the video memory block size field from the initial video memory allocation mode parameter set, extract the fixed allocation unit value of the video memory block size, count the number of state switching instructions per second from the pipeline state setting sequence log set as the state switching frequency, calculate the average of the absolute values of the differences between adjacent state switching timestamps as the switching delay, and count the total number of shader instruction executions per second as the computational intensity; The compilation peak duration, video memory block size, state switching frequency, computational intensity, and switching delay are aligned by timestamp and merged into a multidimensional vector at the same time point. The time series is divided into time windows of 30 seconds to establish a feature set of application startup and multi-tasking GPU activities.
3. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps for obtaining the GPU activity mode classification index are as follows: Based on the application startup and multi-task GPU activity feature set, three feature dimensions, namely, compilation peak duration, video memory block size, and state switching frequency, are selected as clustering inputs, and each data point is Z-score standardized to generate a standardized feature subset; Initialize k cluster centers according to the standardized feature subset, where the k value is determined by the elbow rule, calculate the Euclidean distance from each data point to each cluster center, assign the data point to the nearest neighbor cluster center, and iteratively update the cluster center coordinates to the mean of the data points to which they belong, until the center coordinate change for two consecutive iterations is less than a threshold, thereby generating an optimized cluster center set; Based on the optimized cluster center set, a cluster label is assigned to each time window of each application or task, a mapping table is established, and the mean peak demand of compilation units and the timing mode of video memory bandwidth of the same category are aggregated by label to generate a classification index of GPU activity mode.
4. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps for obtaining the typical GPU resource requirement timing profile are as follows: Based on the GPU activity mode classification index, traverse all time windows under each category label, extract the peak demand of GPU compilation units in each time window, calculate the average of the peak demand of compilation units in all time windows, and generate a set of average peak demand of compilation units in the category; According to the mean value set of peak demand of compilation units in the category, the timing data of video memory bandwidth corresponding to the category label is extracted, and the fluctuation intensity of video memory bandwidth is calculated. The formula is: ; in, is the memory bandwidth fluctuation intensity, For time point The memory bandwidth value, is the bandwidth mean, is the total number of seconds in the time window; Based on the memory bandwidth fluctuation intensity, the call timestamps of the texture mapping unit and the rasterization unit are parsed from the pipeline state setting sequence log set, the call frequency of each unit in the time window is counted, the median and variance of the call frequency are calculated, and the mean of the compilation unit peak demand, the memory bandwidth fluctuation intensity, the median and variance of the call frequency are spliced into a multi-dimensional vector in the order of the time window to construct a typical GPU resource demand timing profile.
5. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps for obtaining the target application startup resource requirement list are as follows: Receive an application startup request, parse the application unique identifier and version number in the request header, traverse all category tags in the GPU activity mode classification index, compare the application identifier with the features under each category tag one by one, select the category tag with the highest feature matching degree, and extract the typical GPU resource demand timing profile corresponding to the category tag; Based on the typical GPU resource requirement timing profile data block, extract all shader identifiers, video memory block sizes and virtual address offset base addresses from the typical GPU resource requirement timing profile data block to generate an intermediate data set; According to the video memory block size in the intermediate data set, the shader identifiers are arranged in descending order from large to small, and the sorted shader identifiers are packaged with the corresponding video memory block size and virtual address offset base address into a target application startup resource requirement list.
6. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps of obtaining the GPU preprocessing instruction sequence are: Parsing the target application startup resource requirement list, traversing the key-value pair of each entry in the target application startup resource requirement list, and extracting the shader identifier, the video memory block size, and the virtual address offset base address; Based on the shader identifier, each shader identifier is replaced with an instruction string according to a precompiled instruction template; based on the video memory block size and the virtual address offset base address, each video memory block size and the virtual address offset base address is replaced with an instruction string according to a video memory pre-allocation instruction template to generate a GPU preprocessing instruction sequence.
7. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps for obtaining the task preemption priority and cost evaluation table are as follows: Monitor the task set currently running on the GPU, obtain the process identifier of each task and the associated typical GPU resource demand timing profile through the operating system kernel interface, extract the time consumption of the last three context switches of each task from the system scheduler log, calculate the average time consumption of the last three context switches as the context switch overhead, and generate a data set; Based on the data set, the task preemption priority score is calculated, and the formula is: ; in, Score the task preemption priority, For the task No. The context switch time is For the task The memory bandwidth usage, The activation status of the foreground window, 1 for activated and 0 for inactivated; According to the task preemption priority score, all tasks are sorted from high to low, and an entry is generated for each task. The entry includes a process identifier, a video memory bandwidth occupancy rate, an average switching time, an activation state and a task preemption priority score, and a task preemption priority and cost evaluation table is constructed.
8. The method for optimizing and improving graphics card performance based on big data according to claim 1, characterized in that: The steps for obtaining the dynamic GPU task preemption control parameter set are: Based on the task preemption priority and cost evaluation table, the top 20% of the tasks in the task preemption priority score are selected as a high-priority task set, the average memory bandwidth occupancy rate and response delay requirement of the high-priority task set are extracted, and the 30% of the remaining tasks with the highest average switching time are extracted as a low-priority task set to generate a high-priority response requirement data set and a low-priority switching cost data set; According to the high-priority response demand data set and the low-priority switching cost data set, the preemption trigger threshold is calculated as follows: ; in, is the preemption trigger threshold, is the average memory bandwidth usage of high-priority tasks, The maximum allowed response delay for high priority tasks, The average switching time for low-priority tasks, The current GPU available video memory ratio; Based on the preemption trigger threshold, the trigger condition is set as follows: when the predicted value of the high priority task response delay exceeds the maximum allowable response delay of the high priority task, and the preemption trigger threshold is greater than 100, the preemption time point is calculated as the start time of the next vertical blanking period. , the formula is , is the current frame number and gets the preemption control parameter set.
9. The graphics card performance optimization and improvement system according to any one of claims 1 to 8, characterized in that: include: The data collection module records the shader compilation time data at the application startup stage, measures the video memory allocation mode parameters, tracks the pipeline state setting sequence, removes outliers, aligns multi-source data by timestamp, and builds a graphics card behavior feature set; The feature clustering module extracts timing feature vectors from the graphics card behavior feature set, calculates the distance between vectors, sets cluster centers, groups application behaviors, assigns category identifiers, establishes a mapping relationship table, counts time distribution, extracts GPU compilation unit requirements, calculates bandwidth usage, performs timing decomposition, forms resource change curves, marks peaks and valleys, and stores them as a GPU resource demand profile library; Resource demand analysis module, which queries the GPU resource demand profile library, matches application categories, extracts resource profile data, analyzes shader call sequences, extracts identifier lists, parses memory allocation peaks, derives pre-allocated space, calculates bandwidth thresholds, marks resource contention points, and outputs application startup resource planning tables; The preprocessing instruction generation module reads the shader identifier in the application startup resource planning table, queries the shader library, verifies the compilation status, filters the precompiled subset, builds the instruction sequence, parses the video memory allocation parameters, converts them into reserved instructions, sets the block size, formulates the memory plan, combines the priority sorting, and encapsulates them into a startup optimization instruction sequence; The task scheduling control module scans the GPU task set, identifies resource occupancy, queries the GPU resource demand profile library to obtain predictions, collects switching records, calculates switching time, compares task priorities, calculates preemption costs, sets competition thresholds, determines preemption timing, and forms a dynamic resource scheduling plan.
Citation Information
Patent Citations
Multi-threaded, self-scheduling processor
CN112088359A
Method and device for dynamically deploying GPU resources and computer equipment
CN112559191A
Self-tuning thread dispatch policy
CN116302107A
Large model training method based on improved ZeRO-Offload technology
CN117992220A
Method and apparatus for software based preemption using two-level binning to improve forward progress of preempted workloads
US20220301095A1
Cited By
Multi-task switching processing method and system of intelligent terminal
CN121029359A
GPU resource processing method and device
CN121092328A