Thread management optimization method and device based on GPU (Graphics Processing Unit) sharing
By using prediction models and shared memory optimization algorithms in the GPU shared environment, dynamically adjusting thread allocation and scheduling, the thread management efficiency problem in the GPU shared environment is solved, and efficient resource utilization and task execution are achieved.
Patent Information
- Application Number
- CN202510705997.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-15
AI Technical Summary
In the GPU sharing environment, how to optimize thread management to ensure task execution efficiency and solve the problems of resource scarcity and high costs.
By using the prediction model to determine thread block configuration and shared memory, dynamically adjust thread allocation and scheduling, monitor memory fragmentation in real time and optimize resource utilization, combining heuristics and feedback mechanisms, we ensure efficient execution of tasks in the GPU shared environment.
It improves the utilization rate of GPU resources and task execution efficiency, reduces resource conflicts and performance bottlenecks, and reduces the cost burden of individual users or applications.
Smart Images

Figure CN120492128A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing, and in particular to a method and device for optimizing thread management based on GPU sharing. Background Art
[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, GPUs (Graphics Processing Units) have transformed from traditional graphics rendering tools into essential core resources for high-performance computing and deep learning. The GPU's parallel processing capabilities and powerful floating-point computing power make it crucial for accelerating tasks such as large-scale data processing, image recognition, natural language processing, and complex model training. However, the scarcity and high cost of GPU resources have limited their widespread adoption. To more efficiently utilize GPU resources while reducing the cost burden for individual users or applications, GPU sharing technology has emerged.
[0003] GPU sharing aims to improve GPU utilization and overall computing efficiency by enabling multiple users or applications to simultaneously access and use the same GPU through appropriate resource allocation and management strategies. In this context, thread management optimization has become a key technology for achieving efficient GPU sharing. Threads are the fundamental unit of task scheduling and resource allocation in an operating system. In a GPU sharing environment, a sound thread management strategy ensures that multiple tasks or applications can run efficiently while sharing GPU resources, avoiding resource conflicts and performance bottlenecks.
[0004] Therefore, how to optimize thread management for GPU sharing to ensure the execution efficiency of tasks in a GPU sharing environment has become a research hotspot in this field. Summary of the Invention
[0005] This application provides a thread management optimization method and device based on GPU sharing, the purpose of which is to achieve GPU sharing thread management optimization to ensure the execution efficiency of tasks in a GPU sharing environment.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] A thread management optimization method based on GPU sharing, comprising:
[0008] When a target task is obtained, the attribute information of the target task and the current GPU state are input into a pre-trained target prediction model to obtain a prediction result output by the target prediction model; the target task is a computing task for requesting GPU resources; the target prediction model is obtained by training a machine learning model using training samples carrying resource configuration labels; the training samples include the attribute information of the sample task and the corresponding sample GPU state; the prediction result includes the thread block configuration and the corresponding performance parameters;
[0009] When the performance parameters meet the requirements, calling the thread block configuration and the shared memory to execute the target task, so as to improve the data processing efficiency of each thread during the execution of the target task; the shared memory is used to store temporary data in each thread;
[0010] When an abnormal load of any thread is detected, the load of the thread is distributed to other threads with a fragmentation rate that meets the requirements to ensure the execution efficiency of the target task; the fragmentation rate is determined based on the memory fragmentation of the thread.
[0011] Optionally, the method further includes:
[0012] In the process of calling the shared memory, using a shared memory optimization algorithm to reduce the memory access latency of each thread in the thread block configuration;
[0013] The use efficiency of the shared memory is monitored in real time, and the shared memory configuration corresponding to each of the threads is updated according to the use efficiency; the shared memory configuration is the configuration when the thread uses the shared memory.
[0014] Optionally, the method further includes:
[0015] During the process of calling the thread block configuration, a preset fragmentation monitoring module is used to monitor the memory fragmentation of each thread in the thread block configuration at different time points in real time; the fragmentation monitoring module is used to calculate the memory fragmentation using a bitmap marking method;
[0016] Determine the time series data of the target task based on the memory fragmentation of each thread at different time points and the execution status information of the task at different time points; the execution status information includes at least task progress and execution time;
[0017] Using a pre-built time series model, predicting the time series data to obtain a high-risk execution period of the target task;
[0018] Before the target task enters the high-risk execution period, target resources are pre-allocated to the target task; the target resources include continuous large-grained memory in GPU resources.
[0019] Optionally, the method further includes:
[0020] Determining a fragmentation index of the target task based on the memory fragmentation of each of the threads; the fragmentation index is used to quantify the degree of fragmentation of the used memory;
[0021] When the fragmentation index exceeds a preset threshold, suspending execution of the target task, and merging memory fragments of each of the threads to obtain corresponding memory blocks;
[0022] The memory block is allocated to the target task, and execution of the target task is started.
[0023] Optionally, the method further includes:
[0024] When there are multiple target tasks, decompose the GPU resources requested by the multiple target tasks into thread grids corresponding to the multiple target tasks; each thread grid includes a group of thread blocks, and each thread grid is provided with a priority tag; the priority tag is determined according to the importance of the target task;
[0025] Determining the thread block size and shared memory configuration corresponding to each thread grid based on the priority tag of each thread grid and the hardware attributes of the GPU resources; the thread block size is used to represent the size of the thread block;
[0026] For each of the thread blocks, a response order of each of the threads is determined based on the execution efficiency of each thread in the thread block and the cooperation capability index between threads, in combination with a preset thread synchronization mechanism.
[0027] Optionally, the method further includes:
[0028] During the execution of the plurality of target tasks, the performance indicators of the GPU resources and the load of the thread corresponding to each target task are monitored in real time;
[0029] Based on the performance indicators and the load of each thread, combined with a heuristic algorithm, an optimized scheduling strategy is obtained; the optimized scheduling strategy includes a new thread grid, a new thread block, and a new thread corresponding to each target task;
[0030] Execute each of the target tasks according to the optimized scheduling strategy.
[0031] Optionally, the method further includes:
[0032] Obtaining current performance parameters of the GPU resources during execution of the target task;
[0033] If the deviation between the current performance parameter and the performance parameter output by the target prediction model is greater than a calibration value, suspending execution of the target task and retraining the target prediction model;
[0034] A new thread block configuration corresponding to the target task is obtained by using the retrained target prediction model.
[0035] A thread management optimization device based on GPU sharing, comprising:
[0036] The prediction unit is configured to, when a target task is obtained, input the attribute information of the target task and the current GPU state into a pre-trained target prediction model to obtain a prediction result output by the target prediction model; the target task is a computing task for requesting GPU resources; the target prediction model is obtained by training a machine learning model using training samples carrying resource configuration labels; the training samples include the attribute information of the sample task and the corresponding sample GPU state; the prediction result includes the thread block configuration and the corresponding performance parameters;
[0037] A task execution unit is configured to call the thread block configuration and shared memory to execute the target task when the performance parameter meets the standard, so as to improve the data processing efficiency of each thread during the execution of the target task; the shared memory is used to store temporary data within each thread;
[0038] The exception handling unit is used to distribute the load of any thread to other threads with a fragmentation rate that meets the requirements when an exception is detected in the load of any thread, so as to ensure the execution efficiency of the target task; the fragmentation rate is determined based on the memory fragmentation of the thread.
[0039] A storage medium includes a stored program, wherein the program is executed by a processor to execute the GPU-sharing-based thread management optimization method.
[0040] An electronic device comprises: a processor, a memory and a bus; the processor and the memory are connected via the bus;
[0041] The memory is used to store programs, and the processor is used to run programs, wherein the program is executed by the processor to execute the GPU-sharing-based thread management optimization method when it is run.
[0042] The technical solution provided by the present application, when the target task is obtained, the attribute information of the target task and the current GPU state are input into the target prediction model to obtain the prediction result output by the target prediction model. When the performance parameters meet the standards, the thread block configuration and shared memory are called to execute the target task to improve the data processing efficiency of each thread during the execution of the target task. When the load of any thread is monitored to be abnormal, the load of any thread is distributed to other threads whose fragmentation rate meets the requirements to ensure the execution efficiency of the target task. The present application utilizes the target prediction model to determine the thread block configuration of the target task, calls the thread block configuration and shared memory, improves the data processing efficiency of the target task, and distributes the abnormal load to other threads for execution, effectively ensuring the execution efficiency of the target task in a GPU sharing environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 A flowchart of a thread management optimization method based on GPU sharing provided in an embodiment of the present application;
[0045] Figure 2 A flowchart of another GPU-sharing-based thread management optimization method provided in an embodiment of the present application;
[0046] Figure 3 A flowchart of another GPU-sharing-based thread management optimization method provided in an embodiment of the present application;
[0047] Figure 4 A flowchart of another GPU-sharing-based thread management optimization method provided in an embodiment of the present application;
[0048] Figure 5 A schematic diagram of the architecture of a thread management optimization device based on GPU sharing provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0050] In this application, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or apparatus comprising the element.
[0051] like Figure 1 As shown, it is a flow chart of a thread management optimization method based on GPU sharing provided in an embodiment of the present application, which includes the following steps.
[0052] S101: When a target task is obtained, attribute information of the target task and the current GPU state are input into a pre-trained target prediction model to obtain a prediction result output by the target prediction model.
[0053] The target task is a computing task that requests GPU resources, and the target prediction model is obtained by training a machine learning model using training samples with resource configuration labels. The training samples include the attribute information of the sample task and the corresponding sample GPU status. The prediction results include the thread block configuration and the corresponding performance parameters.
[0054] In some examples, the attribute information includes, but is not limited to, task type, etc.
[0055] In some examples, the resource configuration tag represents the thread block configuration and performance parameters corresponding to the training sample.
[0056] In some examples, the thread block configuration includes the size and number of thread blocks. A thread block is a grouping of multiple threads that can collaborate during execution.
[0057] In some examples, performance parameters include, but are not limited to, execution time, resource utilization, and the like.
[0058] In some examples, the types of machine learning models include, but are not limited to, decision trees, random forests, neural networks, or reinforcement learning.
[0059] In a possible implementation, a machine learning model is used to learn and predict the execution patterns and performance bottlenecks of different tasks in a GPU sharing environment, and then dynamically adjust the thread block allocation strategy to adapt to real-time changing computing needs. Historical GPU operation data (including attribute information of sample tasks and sample GPU status) is used as training samples to learn the complex relationship between thread block configuration and GPU performance.
[0060] It should be noted that by continuously optimizing the model parameters and training process, the model can be ensured to have high prediction accuracy and generalization ability. After the target prediction model training is completed, the target prediction model can be integrated into the GPU scheduling system. When the target task is submitted to the GPU scheduling system, the scheduling system first analyzes the task characteristics (i.e., attribute information) and calls the target prediction model to predict the optimal thread block configuration. Based on the prediction results of the target prediction model, the scheduling system dynamically adjusts the thread block allocation strategy to maximize the GPU computing resource utilization and the execution efficiency of the target task.
[0061] In some examples, in order to cope with dynamic changes in the GPU sharing environment (such as the addition of other tasks, the increase in GPU temperature, etc.), a feedback mechanism can also be introduced (that is, during the execution process, real-time performance data is continuously collected and compared with the prediction results of the target prediction model. If a large deviation is found between the actual performance and the prediction, the model parameters are adjusted or the model is retrained in a timely manner to improve the accuracy of the prediction).
[0062] Optionally, the feedback mechanism can be summarized as follows: obtaining the current performance parameters of the GPU resources during the execution of the target task; if the deviation between the current performance parameters and the performance parameters output by the target prediction model is greater than a calibration value, pausing the execution of the target task and retraining the target prediction model; using the retrained target prediction model to obtain a new thread block configuration corresponding to the target task.
[0063] In some examples, a preset performance monitoring module can be used to monitor the current performance parameters of GPU resources in real time during the execution of the target task.
[0064] It is understandable that the technical solution of using the target prediction model to dynamically allocate GPU thread blocks can fully utilize GPU resources and improve the execution efficiency and overall performance of the target task through intelligent prediction and dynamic adjustment of thread block configuration.
[0065] S102: When the performance parameters meet the requirements, the thread block configuration and the shared memory are called to execute the target task, so as to improve the data processing efficiency of each thread during the execution of the target task.
[0066] Shared memory is used to store temporary data within each thread. Generally speaking, shared memory allows multiple processes to directly access the same physical memory area (i.e., GPU memory, which can be considered as video memory), thereby reducing data transmission overhead.
[0067] Optionally, for the use of shared memory, you can also use a step-by-step optimization, the optimization method can be found in Figure 2 The steps are shown and the corresponding explanations.
[0068] S103: When it is detected that the load of any thread is abnormal, the load of any thread is distributed to other threads whose fragmentation rates meet the requirements, so as to ensure the execution efficiency of the target task.
[0069] The fragmentation rate is determined based on the memory fragmentation of the thread.
[0070] It's important to note that thread allocation and scheduling are dynamically adjusted based on the actual load of each thread to improve resource utilization efficiency. Specifically, when memory becomes fragmented due to frequent task starts and stops, the affected thread's load is automatically migrated to other less fragmented GPU resources to ensure the target task continues to run efficiently. At the same time, a fragmented resource merging algorithm is used to automatically organize fragmented GPU resources, improving overall resource utilization.
[0071] Optionally, when there are multiple target tasks, an efficient hierarchical scheduling strategy can be constructed to schedule tasks at different levels (such as threads, thread blocks, and grids) to improve GPU resource utilization. The specific implementation process of this efficient hierarchical scheduling strategy can be found in Figure 4 The steps are shown and the corresponding explanations.
[0072] The process shown in S101-S103 above uses the target prediction model to determine the thread block configuration of the target task. When the performance parameters meet the standards, the thread block configuration and shared memory are called to improve the data processing efficiency of the target task, and the abnormal load is distributed to other threads for execution, effectively ensuring the execution efficiency of the target task in the GPU sharing environment.
[0073] like Figure 2 As shown, it is a flow chart of another thread management optimization method based on GPU sharing provided in an embodiment of the present application, including the following steps.
[0074] S201: In the process of calling the shared memory, using a shared memory optimization algorithm to reduce the memory access latency of each thread in the thread block configuration.
[0075] Among them, shared memory optimization algorithms are used to improve multithreaded performance. They ensure efficient data storage in shared memory by precisely planning data layout. This typically involves data partitioning and padding strategies to reduce memory access conflicts and cross-block accesses, improving the continuity and parallelism of data access. Proper data partitioning allows frequently accessed data blocks to be placed in shared memory, reducing reliance on global memory or constant memory and significantly reducing memory access latency.
[0076] In some examples, shared memory optimization algorithms can fully utilize the broadcast and reduce operations of shared memory. These operations can synchronize data processing across multiple threads within a single clock cycle, significantly improving the efficiency and speed of data processing. For example, in image processing or matrix operations, the broadcast nature of shared memory can be leveraged to load common data into multiple threads at once, reducing repeated memory accesses. Simultaneously, reduction operations can quickly aggregate or summarize local data in shared memory, preparing for subsequent global operations.
[0077] In some examples, shared memory optimization algorithms also consider inter-thread coordination and synchronization. In a shared GPU environment, inter-thread synchronization is critical for ensuring data consistency and avoiding race conditions. By using specific locations in shared memory as synchronization points, inter-thread synchronization control can be achieved, ensuring that data access or modification by multiple threads does not conflict.
[0078] S202: Monitor the usage efficiency of the shared memory in real time, and update the shared memory configuration corresponding to each thread according to the usage efficiency.
[0079] The shared memory configuration refers to the configuration when the thread uses the shared memory.
[0080] In some examples, performance analysis tools can be used to monitor the usage efficiency of shared memory in a GPU shared environment, and the shared memory configuration corresponding to each thread can be adjusted based on the shared memory usage efficiency. Alternatively, automatic tuning or machine learning methods can be used to optimize the shared memory configuration corresponding to each thread, and the optimal shared memory configuration can be automatically selected based on the program characteristics and hardware conditions of the process.
[0081] It is understandable that the shared memory optimization algorithm is a comprehensive optimization process involving multiple aspects such as data layout, memory access mode, thread synchronization and automatic tuning. It can maximize the parallel computing capabilities and shared memory resources of the GPU and improve the execution efficiency of the target task.
[0082] S203: During the process of calling the thread block configuration, a preset fragmentation monitoring module is used to monitor the memory fragmentation of each thread in the thread block configuration at different time points in real time.
[0083] Among them, the fragmentation monitoring module is used to calculate the memory fragmentation using a bitmap marking method.
[0084] In some examples, a multi-dimensional data collection mechanism can be established to track GPU workload characteristics in real time (including the number of threads, type, execution time, resource utilization, and performance bottlenecks), and an additional fragmentation monitoring module can be added to use a bitmap marking method to accurately record the boundary distribution of allocated and free areas in the video memory, thereby obtaining real-time memory fragmentation for each process and providing data support for subsequent decision-making.
[0085] S204: Determine time series data of the target task based on the memory fragmentation of each thread at different time points and the execution state information of the task at different time points.
[0086] The execution status information includes at least task progress and execution time.
[0087] S205: Using a pre-built time series model, predict the time series data to obtain a high-risk execution period of the target task.
[0088] Among them, in the process of constructing the time series model, the time series analysis algorithm can be used to analyze the execution status information of historical tasks and the memory fragmentation generation rules to obtain a time series model, so that the time series model can comprehensively evaluate the task arrival rate, execution time distribution and resource demand pattern, and predict the execution period of high fragmentation risk tasks in advance.
[0089] S206: Before the target task enters the high-risk execution period, pre-allocate target resources to the target task.
[0090] The target resources include continuous large-grained memory in GPU resources.
[0091] It is understandable that when a high-risk execution period is detected, the target resources (such as a complete video memory block or a computing core cluster) are allocated to the target task first to avoid resource split execution exacerbating the fragmentation problem.
[0092] Optionally, in order to improve resource utilization, memory fragmentation can be utilized for threads. For the specific utilization process, see Figure 3 The steps are shown and the corresponding explanations.
[0093] The process shown in S201-S206 can take corresponding measures to improve the memory fragmentation of the threads, thereby ensuring that each thread corresponding to the target task can maintain a high level of data processing capability.
[0094] like Figure 3As shown, it is a flow chart of another thread management optimization method based on GPU sharing provided in an embodiment of the present application, including the following steps.
[0095] S301: Determine a fragmentation index of a target task based on memory fragmentation of each thread.
[0096] Among them, the fragmentation index is used to quantify the degree of fragmentation of the used memory.
[0097] In a possible implementation, the more memory fragments each thread has, the higher the fragmentation index of the target task.
[0098] S302: When the fragmentation index exceeds a preset threshold, suspending execution of the target task and merging memory fragments of each thread to obtain corresponding memory blocks.
[0099] When it is detected that the degree of fragmentation exceeds a threshold, the target task (or the thread with lower priority) is suspended and resource reorganization is triggered. The memory fragments of the threads with lower priority are merged and the computing memory blocks are replanned.
[0100] S303: Allocate the memory block to the target task and start executing the target task.
[0101] Specifically, fragmentation metrics are incorporated into the performance evaluation system, dynamically adjusting parameters such as thread block allocation granularity and resource migration thresholds to continuously improve the scheduling strategy's adaptability to complex workloads. To ensure system stability, parameters such as thread block allocation granularity, resource migration thresholds, and defragmentation trigger conditions are dynamically adjusted, forming a closed-loop optimization mechanism of monitoring, prediction, adjustment, and feedback. This mechanism focuses on evaluating the extent of resource utilization improvements and fault tolerance in extreme situations. This feedback mechanism continuously optimizes the scheduling logic, ultimately achieving efficient, low-fragmentation operation of GPU resources in mixed workload scenarios such as deep learning training and graphics rendering.
[0102] The above process shown in S301-S303 can fully utilize the memory fragments of the thread to participate in the execution of the target task, thereby effectively improving the effective utilization rate of GPU resources.
[0103] like Figure 4 As shown, it is a flow chart of another thread management optimization method based on GPU sharing provided in an embodiment of the present application, including the following steps.
[0104] S401: When there are multiple target tasks, decompose GPU resources requested by the multiple target tasks into thread grids corresponding to the multiple target tasks.
[0105] Each thread grid contains a set of thread blocks, and each thread grid is provided with a priority label, which is determined according to the importance of the target task.
[0106] S402: Determine the thread block size and shared memory configuration corresponding to each thread grid based on the priority tag of each thread grid and the hardware attributes of the GPU resources.
[0107] The thread block size is used to represent the size of the thread block.
[0108] It should be noted that at the grid level, the scheduling strategy focuses on the planning and allocation of overall target tasks. Large-scale tasks (multiple target tasks) are decomposed into multiple smaller thread grids. Each thread grid contains a set of thread blocks, which are prioritized based on task importance and urgency. Intelligent scheduling algorithms dynamically adjust the execution order and resource allocation of thread grids to minimize latency and maximize throughput. Furthermore, the scheduling strategy is optimized by considering dependencies and resource contention between target tasks to avoid deadlock and resource waste.
[0109] S403: For each thread block, based on the execution efficiency of each thread in the thread block and the cooperation capability index between threads, combined with a preset thread synchronization mechanism, determine the response order of each thread.
[0110] Based on the execution efficiency of each thread in a thread block and the inter-thread collaboration metrics, combined with a pre-set thread synchronization mechanism, the response order of each thread is determined, reducing inter-thread waiting time and resource conflicts. Furthermore, the GPU's vector processing capabilities and memory access optimization techniques, such as using shared memory to reduce global memory access, further accelerate thread execution.
[0111] In some examples, scheduling strategies focus on thread block organization and scheduling at the thread block level. Based on the GPU's hardware parallel capabilities, the number and size of thread blocks, as well as their distribution within shared memory, are dynamically adjusted to balance resource usage and maximize parallel efficiency. By optimizing data sharing and communication patterns between thread blocks, data transfer latency is reduced, improving overall performance. Furthermore, consideration is given to leveraging dedicated shared memory resources, such as registers and shared memory, to support more complex and efficient computational models.
[0112] S404: During the execution of the multiple target tasks, the performance indicators of the GPU resources and the load of the thread corresponding to each target task are monitored in real time.
[0113] S405: Based on the performance indicators and the load of each thread, combined with the heuristic algorithm, an optimized scheduling strategy is obtained.
[0114] The optimized scheduling strategy includes a new thread grid, a new thread block, and a new thread corresponding to each target task.
[0115] S406: Execute each target task according to the optimized scheduling strategy.
[0116] By monitoring GPU resource performance indicators in real time, key data is collected to analyze workload trends and scheduling effectiveness. Based on this data, machine learning or heuristic algorithms are used to continuously optimize scheduling strategies to adapt to dynamically changing workloads and ensure that the GPU is always operating efficiently.
[0117] The process shown in S401-S406 above can significantly improve GPU resource utilization through an efficient hierarchical scheduling strategy. By implementing refined scheduling decisions and dynamic adjustments between different levels, this strategy can fully utilize the parallel computing capabilities and resources of the GPU, providing an efficient and reliable execution environment for various complex computing tasks.
[0118] like Figure 5 , which is a schematic diagram of the architecture of a thread management optimization device based on GPU sharing provided in an embodiment of the present application, including the units shown below.
[0119] The configuration prediction unit 100, when obtaining a target task, inputs the attribute information of the target task and the current GPU state into a pre-trained target prediction model to obtain a prediction result output by the target prediction model; the target task is a computing task for requesting GPU resources; the target prediction model is obtained by training a machine learning model using training samples carrying resource configuration labels; the training samples include the attribute information of the sample task and the corresponding sample GPU state; the prediction result includes the thread block configuration and the corresponding performance parameters.
[0120] Optionally, the configuration prediction unit 100 is also used to: obtain the current performance parameters of the GPU resources during the execution of the target task; if the deviation between the current performance parameters and the performance parameters output by the target prediction model is greater than the calibration value, suspend the execution of the target task and retrain the target prediction model; use the retrained target prediction model to obtain a new thread block configuration corresponding to the target task.
[0121] The task execution unit 200 is used to call the thread block configuration and shared memory to execute the target task when the performance parameters meet the standards, so as to improve the data processing efficiency of each thread during the execution of the target task; the shared memory is used to store temporary data within each thread.
[0122] Optionally, the task execution unit 200 is also used to: use a shared memory optimization algorithm to reduce the memory access delay of each thread in the thread block configuration during the process of calling the shared memory; monitor the usage efficiency of the shared memory in real time, and update the shared memory configuration corresponding to each thread based on the usage efficiency; the shared memory configuration is the configuration when the thread uses the shared memory.
[0123] Optionally, the task execution unit 200 is also used to: use a preset fragmentation monitoring module to monitor the memory fragmentation of each thread in the thread block configuration at different time points in real time during the process of calling the thread block configuration; the fragmentation monitoring module is used to calculate the memory fragmentation using a bitmap marking method; based on the memory fragmentation of each thread at different time points, and the execution status information of the task at different time points, determine the time series data of the target task; the execution status information includes at least the task progress and the execution time; use a pre-built time series model to predict the time series data to obtain the high-risk execution period of the target task; before the target task enters the high-risk execution period, pre-allocate target resources to the target task; the target resources include continuous large-grained memory in the GPU resources.
[0124] Optionally, the task execution unit 200 is also used to: determine the fragmentation index of the target task based on the memory fragmentation of each thread; the fragmentation index is used to quantify the degree of fragmentation of the used memory; when the fragmentation index exceeds a preset threshold, suspend the execution of the target task, and merge the memory fragments of each thread to obtain the corresponding memory block; allocate the memory block to the target task, and start executing the target task.
[0125] Optionally, the task execution unit 200 is also used to: when there are multiple target tasks, decompose the GPU resources requested by the multiple target tasks into thread grids corresponding to the multiple target tasks; each thread grid contains a group of thread blocks, and each thread grid is provided with a priority label; the priority label is determined according to the importance of the target task; according to the priority label of each thread grid, combined with the hardware properties of the GPU resources, determine the thread block size and shared memory configuration corresponding to each thread grid; the thread block size is used to characterize the size of the thread block; for each thread block, based on the execution efficiency of each thread in the thread block and the cooperation capability index between threads, combined with the preset thread synchronization mechanism, determine the response order of each thread.
[0126] Optionally, the task execution unit 200 is also used to: monitor the performance indicators of GPU resources and the load of the thread corresponding to each target task in real time during the execution of multiple target tasks; obtain an optimized scheduling strategy based on the performance indicators and the load of each thread, combined with a heuristic algorithm; the optimized scheduling strategy includes a new thread grid, a new thread block and a new thread corresponding to each target task; and execute each target task according to the optimized scheduling strategy.
[0127] The exception handling unit 300 is used to distribute the load of any thread to other threads with a required fragmentation rate when detecting an exception in the load of any thread, so as to ensure the execution efficiency of the target task; the fragmentation rate is determined based on the memory fragmentation of the thread.
[0128] Each unit shown above uses the target prediction model to determine the thread block configuration of the target task. When the performance parameters meet the standards, the thread block configuration and shared memory are called to improve the data processing efficiency of the target task, and the abnormal load is distributed to other threads for execution, effectively ensuring the execution efficiency of the target task in the GPU sharing environment.
[0129] The present application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the GPU sharing-based thread management optimization method provided by the present application.
[0130] The present application also provides an electronic device comprising: a processor, a memory, and a bus. The processor and the memory are connected via the bus, the memory is used to store programs, and the processor is used to run the programs, wherein the GPU-sharing-based thread management optimization method provided by the present application is executed when the programs are run.
[0131] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination.
[0132] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the disclosure herein is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A thread management optimization method based on GPU sharing, characterized in that: include: When a target task is obtained, the attribute information of the target task and the current GPU state are input into a pre-trained target prediction model to obtain a prediction result output by the target prediction model; the target task is a computing task for requesting GPU resources; the target prediction model is obtained by training a machine learning model using training samples carrying resource configuration labels; the training samples include the attribute information of the sample task and the corresponding sample GPU state; the prediction result includes the thread block configuration and the corresponding performance parameters; When the performance parameters meet the standards, calling the thread block configuration and shared memory to execute the target task, so as to improve the data processing efficiency of each thread during the execution of the target task; The shared memory is used to store temporary data in each of the threads; When an abnormal load of any thread is detected, the load of the thread is distributed to other threads with a fragmentation rate that meets the requirements to ensure the execution efficiency of the target task; the fragmentation rate is determined based on the memory fragmentation of the thread.
2. The method according to claim 1, characterized in that The method further comprises: In the process of calling the shared memory, using a shared memory optimization algorithm to reduce the memory access latency of each thread in the thread block configuration; The use efficiency of the shared memory is monitored in real time, and the shared memory configuration corresponding to each of the threads is updated according to the use efficiency; the shared memory configuration is the configuration when the thread uses the shared memory.
3. The method according to claim 1, characterized in that The method further comprises: During the process of calling the thread block configuration, a preset fragmentation monitoring module is used to monitor the memory fragmentation of each thread in the thread block configuration at different time points in real time; the fragmentation monitoring module is used to calculate the memory fragmentation using a bitmap marking method; Determine the time series data of the target task based on the memory fragmentation of each thread at different time points and the execution status information of the task at different time points; the execution status information includes at least task progress and execution time; Using a pre-built time series model, predicting the time series data to obtain a high-risk execution period of the target task; Before the target task enters the high-risk execution period, target resources are pre-allocated to the target task; the target resources include continuous large-grained memory in GPU resources.
4. The method according to claim 3, characterized in that The method further comprises: Determining a fragmentation index of the target task based on the memory fragmentation of each of the threads; the fragmentation index is used to quantify the degree of fragmentation of the used memory; When the fragmentation index exceeds a preset threshold, suspending execution of the target task, and merging memory fragments of each of the threads to obtain corresponding memory blocks; The memory block is allocated to the target task, and execution of the target task is started.
5. The method according to claim 1, wherein The method further comprises: When there are multiple target tasks, decompose the GPU resources requested by the multiple target tasks into thread grids corresponding to the multiple target tasks; each thread grid includes a group of thread blocks, and each thread grid is provided with a priority tag; the priority tag is determined according to the importance of the target task; Determining the thread block size and shared memory configuration corresponding to each thread grid based on the priority tag of each thread grid and the hardware attributes of the GPU resources; the thread block size is used to represent the size of the thread block; For each of the thread blocks, a response order of each of the threads is determined based on the execution efficiency of each thread in the thread block and the cooperation capability index between threads, in combination with a preset thread synchronization mechanism.
6. The method according to claim 5, characterized in that The method further comprises: During the execution of the plurality of target tasks, the performance indicators of the GPU resources and the load of the thread corresponding to each target task are monitored in real time; Based on the performance indicators and the load of each thread, combined with a heuristic algorithm, an optimized scheduling strategy is obtained; the optimized scheduling strategy includes a new thread grid, a new thread block, and a new thread corresponding to each target task; Execute each of the target tasks according to the optimized scheduling strategy.
7. The method according to claim 1, characterized in that The method further comprises: Obtaining current performance parameters of the GPU resources during execution of the target task; If the deviation between the current performance parameter and the performance parameter output by the target prediction model is greater than a calibration value, suspending execution of the target task and retraining the target prediction model; A new thread block configuration corresponding to the target task is obtained by using the retrained target prediction model.
8. A thread management optimization device based on GPU sharing, characterized in that: include: The prediction unit is configured to, when a target task is obtained, input the attribute information of the target task and the current GPU state into a pre-trained target prediction model to obtain a prediction result output by the target prediction model; the target task is a computing task for requesting GPU resources; the target prediction model is obtained by training a machine learning model using training samples carrying resource configuration labels; the training samples include the attribute information of the sample task and the corresponding sample GPU state; the prediction result includes the thread block configuration and the corresponding performance parameters; A task execution unit, configured to call the thread block configuration and shared memory to execute the target task if the performance parameter meets the standard, so as to improve the data processing efficiency of each thread during the execution of the target task; The shared memory is used to store temporary data in each of the threads; The exception handling unit is used to distribute the load of any thread to other threads with a fragmentation rate that meets the requirements when an exception is detected in the load of any thread, so as to ensure the execution efficiency of the target task; the fragmentation rate is determined based on the memory fragmentation of the thread.
9. A storage medium, characterized in that: The storage medium includes a stored program, wherein the program is executed by a processor to execute the GPU sharing-based thread management optimization method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: processor, memory, and bus; The processor is connected to the memory via the bus; The memory is used to store programs, and the processor is used to run programs, wherein the program, when run by the processor, executes the GPU-sharing-based thread management optimization method according to any one of claims 1 to 7.
Citation Information
Cited By
Video memory allocation method and device and storage medium
CN121144052A
High-concurrency data processing method and device based on memory pool, equipment and medium
CN122240334A
Memory pool-based high-concurrency data processing method, device, equipment and medium
CN122240334B