GPU kernel execution time prediction method

By using global state vectors and lightweight machine learning models, the problems of local feature misjudgment and low accuracy of static mean in GPU program execution time prediction are solved, achieving efficient and accurate GPU kernel execution time prediction, which is suitable for simulation and performance evaluation of various GPU programs.

CN121300905APending Publication Date: 2026-01-09UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511385132.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In existing technologies for predicting GPU program execution time, local feature analysis cannot accurately reflect global behavior, leading to misjudgments and delayed simulation switching. Furthermore, the accuracy of static mean estimation is low and cannot adapt to changes in dynamic factors, affecting simulation accuracy and efficiency.

Method used

By employing global state vector generation and a lightweight machine learning model, and by comparing the global state vector with the local state vector, the machine learning model dynamically adapts to changes in the execution time of basic blocks, thereby constructing a dynamic prediction model and improving the accuracy of stability detection and prediction.

Benefits of technology

It improves the accuracy of GPU program execution time prediction and simulation efficiency, avoids misjudgments and delayed switching, maintains low computational overhead, and is suitable for execution time modeling and system performance evaluation of various GPU programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121300905A_ABST
    Figure CN121300905A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU (Graphics Processing Unit) kernel execution time prediction method, which comprises the following steps of: 1, performing a round of quick simulation on a GPU kernel, counting the execution frequency of each basic block, and establishing a global state vector G; 2, performing periodic accurate simulation on the GPU kernel, and constructing a local state vector L; step 3, if the similarity of L and G in a plurality of consecutive rounds of sampling exceeds a set threshold value, determining that the program behavior of the GPU kernel tends to be stable; 4, training a lightweight machine learning model, learning a distribution rule of basic block execution time along with global state change, and constructing a dynamic prediction model; and step 5, predicting the execution time by using the trained model, and accumulating the execution time of all the basic blocks to obtain the total execution time of the GPU kernel. The lightweight machine learning model is introduced, the method can dynamically adapt to stage change and non-stationary features of basic block execution time, prediction errors are effectively reduced, and simulation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer system architecture, and particularly relates to a GPU kernel execution time prediction method based on a global state vector. BACKGROUND

[0002] With the wide application of GPUs in high-performance computing, artificial intelligence and scientific simulation, how to accurately and efficiently predict the execution time of GPU programs has become a core problem in architecture design and system optimization. Although the traditional cycle-accurate simulation method has extremely high modeling accuracy, its simulation efficiency is low and it is difficult to meet the needs of large-scale program or system design space exploration. In recent years, sampling simulation has gradually become mainstream. Its basic strategy is to only simulate representative fragments after the program behavior reaches a stable state, thereby significantly improving the simulation efficiency while ensuring accuracy. In this process, stable state detection and execution time prediction are two key links in the sampling simulation system.

[0003] The existing method uses local feature trend analysis based on a sliding window to determine whether to enter a stable state. When the change amplitude of a certain local behavior index (such as instructions per cycle IPC, basic block execution time, etc.) is lower than a set threshold, it is determined that the program has entered a stable phase, and the low-overhead functional simulation mode is switched to. At the same time, in the execution time modeling aspect, the mainstream method uses a static mean estimation strategy, that is, directly using the historical average execution time as the predicted value of the basic block.

[0004] However, the method has the following problems in actual application:

[0005] 1. The local feature is too one-sided: the current method relies on the change of the local behavior trend in the sliding window, and cannot accurately reflect the convergence of the global behavior of the program. GPU programs usually contain complex phase control flow, thread activity fluctuation and dynamic changes of resource occupation, etc. The local stable state is not equivalent to the global stable state, and this local-global inconsistency easily leads to:

[0006] 1) misjudging the phase stability as long-term stability, and then introducing invalid sampling;

[0007] 2) delaying the entry into the sampling mode, missing the best simulation switching opportunity, and affecting the overall simulation speedup ratio.

[0008] 2. Static mean estimation has low accuracy: the execution time of the basic block is affected by various dynamic factors (such as scheduling strategy, resource conflict, data correlation, etc.), and there is obvious non-stationarity and phase fluctuation. The static mean cannot describe these dynamic evolution processes, which easily leads to the accumulation of prediction errors and reduces the overall accuracy of the simulator.

[0009] Therefore, there is an urgent need for a GPU kernel execution time modeling method that can capture the evolution characteristics of global execution behavior and has dynamic prediction capability. SUMMARY

[0010] The present application aims to overcome the shortcomings of the prior art, and provides a GPU kernel execution time prediction method that introduces a lightweight machine learning model, can dynamically adapt to the phased changes and non-stationary characteristics of basic block execution time, effectively reduces the prediction error, and improves the simulation accuracy.

[0011] The purpose of the present application is achieved by the following technical scheme: a GPU kernel execution time prediction method, comprising the following steps:

[0012] Step 1, global state vector generation: before the GPU kernel is formally executed, a round of fast simulation is performed on the GPU kernel, and the execution frequency of each basic block is counted, and the execution frequency of all basic blocks is constructed into a global state vector G;

[0013] Step 2, local state vector construction: formally start executing the GPU kernel, and perform cycle-accurate simulation on the GPU kernel, and sample the execution frequency of each basic block in the current stage of the program being executed by the GPU kernel at a fixed cycle, and construct all basic block execution frequencies in each sampling cycle into a local state vector L;

[0014] Step 3, stability detection: compare the similarity between the local state vector L and the global state vector G, if the similarity of L and G in continuous several rounds of sampling exceeds the set threshold, it is determined that the program behavior of the GPU kernel has tended to be stable;

[0015] Step 4, prediction model training: after entering the stable state, the modeling process of the basic block is started; train the lightweight machine learning model to learn the distribution rule of the basic block execution time with the global state change, and construct a dynamic prediction model;

[0016] Step 5, execution time prediction: when the GPU kernel executes the program and encounters a basic block, the trained model is used to predict its execution time, and the execution times of all basic blocks are added to obtain the total execution time of the GPU kernel.

[0017] The specific method of step 4 is: for the jth basic block BB j , collect its execution time data for the last m times, and divide the collected execution time data in a sliding window manner to generate training samples; train the lightweight machine learning model using the training samples to learn the distribution rule of the basic block execution time with the global state change, and construct a dynamic prediction model for the basic block execution time.

[0018] The GPU kernel execution time prediction method has the advantages that the method breaks through the limitation of traditional methods which only rely on local behavior trend to judge the steady state, introduces a global state vector, comprehensively reflects the whole execution behavior characteristics of the program, significantly improves the accuracy of stability detection, avoids misjudgment and delays entering the sampling stage, and compared with a static mean strategy, the method introduces a lightweight machine learning model, can dynamically adapt to the phased changes and non-stationary characteristics of the basic block execution time, effectively reduces the prediction error, and improves the simulation accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A flowchart of the GPU kernel execution time prediction method of the application. DETAILED DESCRIPTION

[0020] The technical solutions of the application will be further described below with reference to the drawings.

[0021] As shown in the drawings, the GPU kernel execution time prediction method of the application comprises the following steps: Figure 1

[0022] Step 1, global state vector generation: before the GPU kernel is formally executed, a simulator is used to perform a round of fast simulation (functional simulation) on the GPU kernel, the execution frequency of each basic block is counted, and the execution frequency of all basic blocks is constructed into a global state vector G as a representation of the whole behavior characteristics of the program;

[0023] Step 2, local state vector construction: formally start executing the GPU kernel, perform cycle-accurate simulation on the GPU kernel, sample the execution frequency of each basic block in the current stage of the program being executed by the GPU kernel at a fixed cycle, and construct all the basic block execution frequencies in each sampling cycle into a local state vector L representing the execution characteristics of the program in the time window;

[0024] Step 3, stability detection: compare the similarity between the local state vector L and the global state vector G, if the similarity between L and G in continuous several rounds of sampling all exceeds a set threshold, it is determined that the program behavior of the GPU kernel has tended to be stable, and step 4 is executed, otherwise, return to step 2;

[0025] Step 4, prediction model training: after entering the stable state, the modeling process of the basic block is started; the specific method is: for the jth basic block BB j ​, collect its recent m times (usually 1024 or 2048) execution time data, and divide the collected execution time data in a sliding window (window size n) to generate training samples; train a lightweight machine learning model (such as decision tree, shallow MLP or sliding window average) using the training samples, learn the distribution rule of the basic block execution time with the global state change, and construct a low-overhead dynamic prediction model of the basic block execution time. For the jth basic block BB j , the lightweight machine learning model predicts the ith execution time through the historical execution time sequence [t i-n ,t i-n+1 ,…,t i-1 ]; the model form can be expressed as:

[0026]

[0027] Wherein is the regression function trained on the basic block BB j , the input is the recent n times execution time, and the output is the predicted value of the ith execution time.

[0028] Step 5, execution time prediction: after the model training is completed, the simulator switches to the functional simulation mode and no longer performs the cycle-accurate simulation; the subsequent GPU kernel execution program uses the trained model to predict the execution time of each basic block when it encounters a basic block, and the total execution time of the GPU kernel is obtained by accumulating the execution times of all basic blocks.

[0029] Those skilled in the art will realize that the embodiments described herein are for the purpose of helping the reader understand the principles of the present application and should be understood as not limiting the scope of protection of the present application to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the scope of protection of the present application.

Claims

1. A method for predicting GPU kernel execution time, characterized in that, Includes the following steps: Step 1, Global State Vector Generation: Before the GPU kernel officially starts executing, perform a quick simulation of the GPU kernel to count the execution frequency of each basic block, and construct the global state vector G from the execution frequencies of all basic blocks. Step 2, Local State Vector Construction: The GPU kernel is officially executed. Periodic precise simulation of the GPU kernel is performed. The execution frequency of each basic block of the program being executed by the GPU kernel is sampled at fixed intervals. The execution frequency of all basic blocks in each sampling period is used to construct a local state vector L. Step 3, Stability Detection: Compare the similarity between the local state vector L and the global state vector G. If the similarity between L and G exceeds the set threshold in several consecutive rounds of sampling, it is determined that the program behavior of the GPU kernel has become stable. Step 4, Predictive Model Training: After entering a stable state, start the modeling process of basic blocks; train a lightweight machine learning model to learn the distribution of the execution time of basic blocks as the global state changes, and build a dynamic prediction model. Step 5: Execution Time Prediction: Each time the GPU kernel execution program encounters a basic block, it uses a trained model to predict its execution time, and adds up the execution times of all basic blocks to obtain the total execution time of the GPU kernel.

2. The GPU kernel execution time prediction method according to claim 1, characterized in that, The specific method for step 4 is as follows: For the j-th basic block BB j Collect the time data of its most recent m executions, and segment the collected execution time data in a sliding window manner to generate training samples; A lightweight machine learning model is trained using training samples to learn the distribution pattern of basic block execution time as the global state changes, and a dynamic prediction model of basic block execution time is constructed.