Thread block configuration method and device, readable storage medium and program product

By dynamically configuring thread block sizes based on system performance metrics and bottleneck features, the method addresses the challenges of adapting to dynamic loads, reducing costs and complexity in thread block size configuration.

CN120315751AActive Publication Date: 2025-07-15LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510780874.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-15
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing thread block configuration methods cannot adapt to dynamic load changes, have high maintenance costs, high data acquisition costs, and high complexity in model deployment and migration.

Method used

By collecting system operation status indicators, obtaining bottleneck features, calculating feature scores and weights, combining system operation status and feature weight information for thread block configuration, and dynamically generate the optimal size.

Benefits of technology

It realizes adaptive dynamic load changes, saves maintenance and data acquisition costs, dynamically generates the optimal thread block size, and improves resource utilization and program execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315751A_ABST
    Figure CN120315751A_ABST
Patent Text Reader

Abstract

The invention discloses a thread block configuration method and device, a readable storage medium and a program product, and relates to the technical field of computers, the thread block configuration method comprises the following steps: collecting each system operation state index, obtaining various preset bottleneck features, respectively carrying out feature score calculation on each bottleneck feature according to each system operation state index, and obtaining a feature score of each bottleneck feature; according to the method, the system operation state indexes and the feature weight information are calculated, weight calculation is performed according to the calculated feature scores, thread block configuration is performed in combination with the system operation state indexes and the feature weight information, the influence of various bottleneck features on thread block configuration is fully considered, and a specific thread block configuration strategy is adopted according to the feature weight information of various bottleneck features. The technical problems that the dynamic load change cannot be adapted, the maintenance cost is high, the data acquisition cost is high, and the complexity of model deployment and migration is high are solved, and the technical effects that the dynamic load change can be adapted, the maintenance cost and the data acquisition cost are saved, and the optimal thread block size can be dynamically generated are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a thread block configuration method, device, readable storage medium, and program product. Background Art

[0002] In a parallel computing architecture, the configuration of the size of a thread block (block_size) directly affects the utilization rate of computing resources. Currently, the commonly used thread block configuration methods mainly include two types: one is that during runtime, the program statically selects the optimal size configuration by querying a pre-generated thread block size suggestion table; the other is to predict the optimal thread block size by calling a trained neural network model based on the current input data shape.

[0003] However, both of the above two thread block configuration methods have their respective drawbacks. First, the method of statically selecting the optimal size configuration by querying a pre-generated thread block size suggestion table cannot adapt to dynamic load changes, relies on manual parameter tuning, and increases the maintenance cost. Second, the method of predicting the optimal thread block size by calling a trained neural network model relies on a large amount of training data, has a high data acquisition cost, a high prediction latency, and a high hardware dependence, increasing the complexity of model deployment and migration. Summary of the Invention

[0004] This application provides a thread block configuration method, device, readable storage medium, and program product to at least solve the problems in the related art of being unable to adapt to dynamic load changes, having a high maintenance cost, a high data acquisition cost, and a high complexity of model deployment and migration.

[0005] This application provides a thread block configuration method, including: Collecting various system operation status indicators and obtaining preset various bottleneck features; Calculating feature scores for various bottleneck features respectively according to the various system operation status indicators to obtain each feature score; Calculating feature weights for each feature score respectively to obtain the feature weight information corresponding to various bottleneck features; Performing thread block configuration according to the various system operation status indicators and the various feature weight information.

[0006] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above thread block configuration methods when executing the computer program.

[0007] This application also provides a non-volatile computer-readable storage medium, in which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above thread block configuration methods.

[0008] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above thread block configuration methods.

[0009] With the present application, by collecting the operation status indicators of each system, obtaining various preset bottleneck features, calculating the feature scores for each type of bottleneck feature respectively according to the operation status indicators of each system, calculating the weights based on the calculated feature scores, and then combining the operation status indicators of each system and the feature weight information for thread block configuration, the influence of each type of bottleneck feature on thread block configuration is fully considered, and a specific thread block configuration strategy is adopted according to the feature weight information of each type of bottleneck feature. Therefore, technical problems such as inability to adapt to dynamic load changes, high maintenance costs, high data acquisition costs, and high complexity of model deployment and migration can be solved, and the technical effects of being able to adapt to dynamic load changes automatically, saving maintenance costs and data acquisition costs, and being able to dynamically generate the optimal thread block size can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is the flowchart of the implementation of a thread block configuration method provided by an embodiment of the present application; Figure 2 It is the flowchart of the implementation of another thread block configuration method provided by an embodiment of the present application; Figure 3 It is the structural block diagram of a thread block configuration provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0013] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0015] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the thread block configuration method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0016] An embodiment of this application provides a thread block configuration method. In combination with the execution process of the thread block configuration method, the method will be described in detail.

[0017] See Figure 1 , Figure 1 which is the implementation flowchart of a thread block configuration method provided by an embodiment of this application. The method may include the following steps.

[0018] S101: Collect each system operation status indicator and obtain various preset bottleneck characteristics.

[0019] During the operation of the system, each system operation status indicator is collected. For example, system operation status indicators such as the occupancy rate of the Streaming Multiprocessor (SM), the memory bandwidth utilization rate, and the hit rate of the Level 1 Cache (L1) can be obtained.

[0020] Various preset bottleneck characteristics can also be obtained. The various bottleneck characteristics may include compute-intensive characteristics, memory-intensive characteristics, and control-flow-intensive characteristics.

[0021] S102: Calculate the feature scores for each type of bottleneck characteristic respectively according to each system operation status indicator to obtain each feature score.

[0022] After collecting each system operation status indicator and obtaining various preset bottleneck characteristics, calculate the feature scores for each type of bottleneck characteristic respectively according to each system operation status indicator to obtain each feature score.

[0023] The sliding window weighting mechanism can be adopted to process time series index data. The window size is default set to 10 sampling periods to balance real-time performance and stability. For the index set Mi within the i-th window period, the feature score can be calculated through the following formula.

[0024] ; ; ; Among them, represents calculating the mean value of , represents the occupancy rate of the stream processor, represents the memory bandwidth utilization rate, represents the instruction issue efficiency, is the calculated score of the i-th sliding window, is the memory score of the i-th sliding window, is the control flow score of the i-th sliding window.

[0025] S103: Calculate the feature weights for each feature score respectively to obtain the feature weight information corresponding to each type of bottleneck feature.

[0026] After calculating each feature score, calculate the feature weights for each feature score respectively to obtain the feature weight information corresponding to each type of bottleneck feature. For example, the feature scores can be normalized first, and then the feature weight information can be calculated based on the normalized feature scores.

[0027] Continuing with the example in step S102, finally, it is processed through weighted ratio normalization to form a standardized feature vector: ; Among them, is the computing intensity, is the memory intensity, is the control flow intensity.

[0028] A circular buffer can be maintained to store the system operation state indicators of the most recent N sampling periods. Each time it is updated, first calculate the median value of each system operation state indicator within the window, and then generate a three-dimensional feature vector based on the predefined feature discrimination rules.

[0029] S104: Configure the thread blocks according to each system operation state indicator and each feature weight information.

[0030] After calculating the feature weight information corresponding to various bottleneck features respectively, thread block configuration is performed according to each system operation status index and each feature weight information. Through the feature weight information corresponding to various bottleneck features respectively, the current bottleneck scenario can be determined, and then the optimal thread block size at present can be dynamically generated by integrating the feature weight information corresponding to various bottleneck features respectively and resource constraints.

[0031] Through this application, by collecting each system operation status index, obtaining various preset bottleneck features, calculating the feature scores for various bottleneck features respectively according to each system operation status index, calculating weights according to the calculated feature scores, and then performing thread block configuration by combining each system operation status index and each feature weight information, the influence of various bottleneck features on thread block configuration is fully considered, and a specific thread block configuration strategy is adopted according to the feature weight information of various bottleneck features. Therefore, technical problems such as inability to adapt to dynamic load changes, high maintenance costs, high data acquisition costs, and high complexity of model deployment and migration can be solved, and the technical effects of being able to adapt to dynamic load changes automatically, saving maintenance costs and data acquisition costs, and being able to dynamically generate the optimal thread block size are achieved.

[0032] See Figure 2 , Figure 2 which is the implementation flowchart of another thread block configuration method provided by the embodiment of this application. This method may include the following steps.

[0033] S201: Obtain each system operation status index by performing asynchronous sampling of the index, and obtain various preset bottleneck features.

[0034] When collecting the system operation status index, obtain each system operation status index by performing asynchronous sampling of the index. For example, use the Compute Unified Device Architecture Profiling Tools Interface (CUPTI) to perform asynchronous sampling to obtain each system operation status index. By using the Compute Unified Device Architecture Profiling Tools Interface to perform asynchronous sampling on each system operation status index, it is ensured that the collection process of hardware performance data will not interfere with the main computing process.

[0035] When collecting the system operation status index, select performance indicators that are highly sensitive to thread block configuration to ensure that the monitoring results can accurately reflect the influence of thread block configuration on the utilization of hardware resources.

[0036] By deeply analyzing the characteristics of the currently executing Compute Unified Device Architecture (CUDA) computing tasks, these characteristics are crucial for subsequent adjustment of thread block configurations. Specifically, when the occupancy rate of the streaming processors remains high while the memory bandwidth utilization rate is low, it can be determined that the current task is computationally intensive. Conversely, if the memory bandwidth utilization rate approaches the peak while the occupancy rate of the streaming processors is insufficient, it indicates that the task is memory access-limited. For control flow-intensive tasks, it is manifested as a low instruction issue efficiency. By analyzing these characteristics, accurate task features can be provided, enabling more reasonable thread block configuration decisions based on the actual needs of the task, thereby achieving the optimal allocation and utilization of hardware resources and enhancing the execution efficiency and scalability of CUDA programs.

[0037] For example, computationally intensive characteristics: When the occupancy rate of the streaming processors continuously exceeds the dynamically adjusted threshold (such as 70%), and the memory bandwidth utilization rate is lower than the set lower limit in the current environment (such as 30%), the task is determined to have a high computational density. Memory-intensive characteristics: If the memory bandwidth utilization rate exceeds 60% and the occupancy rate of the streaming processors is less than 40%, and is accompanied by a decrease in the hit rate of the level 1 / texture (L1 / TEX) cache, it is identified as a task dominated by a memory bottleneck. Control flow-intensive characteristics: By quantifying the instruction issue efficiency, when this value exceeds 0.5 after normalization, it indicates that there is a control flow divergence in the task.

[0038] S202: Calculate the feature scores for each type of bottleneck feature based on the respective system operating state indicators to obtain the respective feature scores.

[0039] S203: Calculate the feature weights for each feature score to obtain the feature weight information corresponding to each type of bottleneck feature.

[0040] S204: Obtain the preset candidate thread block sizes.

[0041] Preset the candidate thread block sizes in advance. For example, they can be set to include 128, 160, 192, 224,..., 1024. Obtain the preset candidate thread block sizes.

[0042] In a specific embodiment of the present application, step S204 may include the following steps: Step 1: Obtain the system hardware resource constraints; Step 2: Determine the candidate thread block sizes based on the system hardware resource constraints.

[0043] For ease of description, the above two steps can be combined for explanation.

[0044] Thread blocks are subject to system hardware resource constraints. Obtain the system hardware resource constraints and determine candidate thread block sizes according to the system hardware resource constraints. By determining candidate thread block sizes according to the system hardware resource constraints, illegal setting of candidate thread block sizes is avoided, ensuring the legality of the setting of candidate thread block sizes.

[0045] S205: Select the current candidate thread block size from the candidate thread block sizes according to each system operation status index and each feature weight information.

[0046] After obtaining the preset candidate thread block sizes, select the current candidate thread block size from the candidate thread block sizes according to each system operation status index and each feature weight information.

[0047] S206: Obtain the system performance feedback value during the actual operation of the system based on the current candidate thread block size.

[0048] After selecting the current candidate thread block size, obtain the system performance feedback value during the actual operation of the system based on the current candidate thread block size. For example, the system execution time can be obtained.

[0049] S207: Perform adaptive trade-off reinforcement learning between exploration and exploitation according to the system performance feedback value to obtain the target thread block size, and configure the thread block of the system to the target thread block size.

[0050] After obtaining the system performance feedback value during the actual operation of the system based on the current candidate thread block size, perform adaptive trade-off reinforcement learning between exploration and exploitation according to the system performance feedback value to obtain the target thread block size, and configure the thread block of the system to the target thread block size, so as to safely and efficiently apply the target thread block size to the operation configuration of the CUDA kernel function. Thus, dynamic tuning of the thread block size is achieved. By adopting a closed-loop feedback mechanism, it can adapt to different hardware architectures and computing tasks and has high scalability.

[0051] The system deeply analyzes the characteristics of the currently executing CUDA computing task by combining the characteristics of the computing task (such as data access mode, computational density distribution, etc.). On this basis, the system comprehensively considers the hardware status and task characteristics, and adopts a multi-strategy fusion method combining a rule engine and reinforcement learning to dynamically generate the optimal thread block size, thereby achieving the optimal allocation and utilization of hardware resources, and significantly improving the resource utilization rate and program execution efficiency.

[0052] In a specific embodiment of the present application, after collecting each system operation status index, the method may further include the following steps: Step 1: Obtain the preset index weights corresponding to each preset index dimension; Step 2: Normalize each system operation status indicator to obtain each normalized indicator value; Step 3: Calculate the dimension scores corresponding to each preset indicator dimension according to each normalized indicator value; Step 4: Perform weighted summation calculation on each dimension score according to each preset indicator weight to obtain a comprehensive score; Step 5: Adjust the indicator sampling frequency according to the comprehensive score.

[0053] For convenience of description, the above five steps can be combined for explanation.

[0054] Obtain the preset indicator weights corresponding to each preset indicator dimension, normalize each system operation status indicator to obtain each normalized indicator value, calculate the dimension scores corresponding to each preset indicator dimension according to each normalized indicator value, perform weighted summation calculation on each dimension score according to each preset indicator weight to obtain a comprehensive score, and adjust the indicator sampling frequency according to the comprehensive score. During the process of collecting monitoring indicators, by performing adaptive sampling frequency control according to the comprehensive score of each system operation status indicator and performing asynchronous sampling, the accuracy and low overhead of data collection are ensured.

[0055] A multi-level monitoring system from the macro level (stream processor level) to the micro level (warp level) can be established to comprehensively cover all dimensions of hardware performance.

[0056] At the macro level of the stream processor level, the focus is mainly on the stream processor occupancy rate (active Blocks per SM / max Blocks per SM) and the active warp ratio (active Warps per SM / max Warps per SM). These two indicators can intuitively reflect the utilization of stream processor resources and provide an important basis for judging whether the thread block configuration is reasonable. For example, when the stream processor occupancy rate is low, it may mean that the thread block size is set too small, resulting in insufficient utilization of stream processor resources. And too high an active warp ratio may indicate fierce competition among warps within the thread block, affecting the execution efficiency.

[0057] At the memory level, the monitoring focus is on the memory bandwidth utilization rate and the L1 / TEX cache hit rate. These indicators are directly related to the access efficiency of the CUDA program to memory resources and are the key basis for judging whether the function is limited by the memory bandwidth during operation. For example, too low a memory bandwidth utilization rate may indicate that data transmission has become a performance bottleneck, while the L1 / TEX cache hit rate reflects the effect of data locality utilization and is of great significance for optimizing the memory access pattern.

[0058] At the thread bundle level, the monitoring metrics include the instruction issue efficiency. This metric can deeply reveal the execution details of warps within a thread block and provide guidance for optimizing parallel execution within the thread block. For example, low instruction issue efficiency may indicate resource competition or complex instruction dependencies within the thread block, affecting the parallel execution efficiency.

[0059] In a specific embodiment of the present application, adjusting the index sampling frequency according to the comprehensive score may include the following steps: Step 1: Obtain the preset upper limit and lower limit of the comprehensive score; Step 2: When the comprehensive score exceeds the upper limit of the comprehensive score, adjust the index sampling frequency to the first preset frequency; Step 3: When the comprehensive score is lower than the lower limit of the comprehensive score, adjust the index sampling frequency to the second preset frequency; Wherein, the first preset frequency is less than the second preset frequency.

[0060] For ease of description, the above three steps can be combined for explanation.

[0061] Obtain the preset upper limit and lower limit of the comprehensive score. When the comprehensive score exceeds the upper limit of the comprehensive score, adjust the index sampling frequency to the first preset frequency. When the comprehensive score is lower than the lower limit of the comprehensive score, adjust the index sampling frequency to the second preset frequency, and the first preset frequency is less than the second preset frequency. For example, the first preset frequency can be set to 30Hz and the second preset frequency can be set to 200Hz. By adjusting the index sampling frequency down to the first preset frequency when the comprehensive score exceeds the upper limit of the comprehensive score, over-sampling can be avoided. By adjusting the index sampling frequency up to the second preset frequency when the comprehensive score is lower than the lower limit of the comprehensive score, potential changes can be captured.

[0062] In a specific embodiment of the present application, after collecting the operating state metrics of each system, the method may further include the following steps: Perform smoothing processing on the operating state metrics of each system.

[0063] Since the hardware performance metrics may have short-term fluctuations, after collecting the operating state metrics of each system, it is necessary to perform smoothing processing on the operating state metrics of each system. For example, smoothing processing can be performed through moving average or exponentially weighted average to reduce noise interference.

[0064] In a specific embodiment of the present application, after collecting the operating state metrics of each system, the method may further include the following steps: Step 1: Obtain the metric thresholds corresponding to the operating state metrics of each system; Step 2: When it is determined that the system is in a bottleneck scenario based on the system operation status indicators and each indicator threshold, perform thread block size adjustment.

[0065] For ease of description, the above two steps can be combined for explanation.

[0066] Obtain the indicator thresholds corresponding to each system operation status indicator. When it is determined that the system is in a bottleneck scenario based on the system operation status indicators and each indicator threshold, perform thread block size adjustment to relieve memory pressure or increase the occupancy rate of the stream processor.

[0067] The bottleneck scenario can include a memory - limited scenario, a compute - limited scenario, a compute - limited and high - compute - density scenario, etc.

[0068] In a specific implementation manner of the present application, obtaining the indicator thresholds corresponding to each system operation status indicator may include the following steps: Step 1: Obtain the historical indicator sets corresponding to each system operation status indicator; Step 2: Calculate the mean value of each historical indicator set respectively to obtain the historical indicator mean values corresponding to each historical indicator set; Step 3: Calculate the variance of each historical indicator set respectively to obtain the historical indicator variances corresponding to each historical indicator set; Step 4: Calculate the indicator thresholds corresponding to each system operation status indicator according to each system operation status indicator, each historical indicator mean value, and each historical indicator variance.

[0069] For ease of description, the above four steps can be combined for explanation.

[0070] Obtain the historical indicator sets corresponding to each system operation status indicator, calculate the mean value of each historical indicator set respectively to obtain the historical indicator mean values corresponding to each historical indicator set, calculate the variance of each historical indicator set respectively to obtain the historical indicator variances corresponding to each historical indicator set, and calculate the indicator thresholds corresponding to each system operation status indicator according to each system operation status indicator, each historical indicator mean value, and each historical indicator variance. By dynamically adjusting the indicator threshold boundary according to the mean and variance of the stream processor occupancy rate and memory bandwidth within a past period of time, the system avoids inaccurate feature classification caused by improper static threshold setting.

[0071] In a specific implementation manner of the present application, when it is determined that the system is in a bottleneck scenario based on the system operation status indicators and each indicator threshold, performing thread block size adjustment may include the following steps: Step 1: When the memory bandwidth utilization rate among the system running state indicators exceeds the current memory bandwidth utilization rate threshold and the first-level cache hit rate is lower than the current first-level cache hit rate threshold, determine that the current bottleneck scenario is a memory-constrained scenario; Step 2: Downscale the thread block size according to the memory-constrained scenario.

[0072] For convenience of description, the above two steps can be combined for explanation.

[0073] When the memory bandwidth utilization rate among the system running state indicators exceeds the current memory bandwidth utilization rate threshold and the first-level cache hit rate is lower than the current first-level cache hit rate threshold, determine that the current bottleneck scenario is a memory-constrained scenario, and downscale the thread block size according to the memory-constrained scenario to relieve the memory pressure.

[0074] In a specific implementation manner of the present application, when it is determined that the system is in a bottleneck scenario according to the system running state indicators and each indicator threshold, adjusting the thread block size may include the following steps: Step 1: When the occupancy rate of the stream processors among the system running state indicators is lower than the current stream processor occupancy rate threshold and the warp efficiency exceeds the current memory bandwidth utilization rate threshold, determine that the current bottleneck scenario is a compute-constrained scenario; Step 2: Upscale the thread block size according to the compute-constrained scenario.

[0075] For convenience of description, the above two steps can be combined for explanation.

[0076] When the occupancy rate of the stream processors among the system running state indicators is lower than the current stream processor occupancy rate threshold and the warp efficiency exceeds the current memory bandwidth utilization rate threshold, determine that the current bottleneck scenario is a compute-constrained scenario, and upscale the thread block size according to the compute-constrained scenario to increase the occupancy rate of the stream processors.

[0077] In a specific implementation manner of the present application, when it is determined that the system is in a bottleneck scenario according to the system running state indicators and each indicator threshold, adjusting the thread block size may include the following steps: When it is determined that the current bottleneck scenario is a compute-constrained and high-compute-density scenario according to the system running state indicators and each indicator threshold, upscale the thread block size.

[0078] When it is determined that the current bottleneck scenario is a compute-constrained and high-compute-density scenario according to the system running state indicators and each indicator threshold, upscale the thread block size to increase the occupancy rate of the stream processors.

[0079] To avoid the impact of frequent sampling and excessive metrics on the running efficiency, during actual deployment, several main metrics such as SM occupancy rate, active warp ratio, memory bandwidth utilization, L1 / TEX cache hit rate, and instruction issue efficiency are continuously sampled. The following is the pseudo-code of the implementation idea to illustrate the basic process of data collection and processing and the implementation at the hardware interface layer.

[0080] 。

[0081] During the process of collecting the running state metrics of the monitoring system, adaptive sampling frequency control is adopted. The adjustment of the sampling frequency is based on the currently monitored metric data and is achieved through the weighted comprehensive scoring method to dynamically balance the sampling overhead and data accuracy.

[0082] For example, the weight of the stream processor occupancy rate metric is 0.4, the weight of the memory occupancy rate metric is 0.4, and the weight of the warp-level metric is 0.2. Then, the normalized values of each metric are multiplied by the weights and summed to obtain the comprehensive score. The sampling frequency is then dynamically adjusted according to the comprehensive score. Different types of tasks can also adjust the metric weights. The following example is used to illustrate the idea: When the comprehensive score exceeds the threshold (such as 0.8), the sampling frequency is reduced; When the comprehensive score is below the threshold (such as 0.2), the sampling frequency is increased.

[0083] 。

[0084] The pseudo-code for smoothing the collected system running state metrics is as follows.

[0085] 。

[0086] This module maintains a circular buffer to store the hardware metrics of the most recent N sampling periods. Each time it is updated, the median of each metric within the window is first calculated, and then a three-dimensional feature vector is generated based on the predefined feature discrimination rules. To avoid the interference of single-sampling period noise, the module adopts a double-verification mechanism: a short-term window (5 - 10 periods) is used for rapid response to feature changes, and long-term trend analysis (50 - 100 periods) is used to correct systematic biases. This hierarchical processing method not only ensures real-time performance but also guarantees the robustness of feature determination, providing a reliable input feature space for the subsequent configuration decision engine. The following is the pseudo-code for the core function of this module.

[0087] 。

[0088] To improve the sensitivity of bottleneck analysis, a dynamic weight mechanism is introduced to evaluate the computing and memory bottlenecks. The pseudo-code is as follows.

[0089] 。

[0090] Reinforcement learning policy input: The dynamic weight can be used as a component of the state, improving the convergence and generalization ability of reinforcement learning.

[0091] The reinforcement learning module is mainly responsible for further fine-tuning the thread block size recommendation policy when the empirical rules and performance models cannot cover or perform poorly. It searches for thread block configurations by considering factors such as bottleneck types, resource utilization, and instruction issue efficiency. The goal is to maximize the overall efficiency under the condition of meeting the hardware resource constraints. To improve the flexibility and scalability of decision-making, a caching mechanism, empirical rules, and reinforcement learning fine-tuning mechanism are introduced to achieve multi-strategy integration.

[0092] The reinforcement learning sub-module is defined as: State: The hardware state vector (Occupancy, mem_util, warp_eff, and the dynamically calculated weight above) under the current configuration; Action: The selection from the candidate block_size options (e.g., {128, 160, 192, 224,..., 1024}); Reward: The performance gain (decrease in execution time) of the current block configuration during actual operation; Policy: Adopt a policy to adaptively balance exploration and exploitation.

[0093] The reinforcement learning module trains using cached data and continuously fine-tunes the configuration recommendation policy.

[0094] 。

[0095] After obtaining the thread block configuration parameters, the parameters are safely applied to the startup process of the CUDA kernel function, and an execution performance monitoring mechanism is integrated. Its core objectives include: avoiding illegal configurations (such as exceeding shared memory and thread number limits); optimizing the kernel startup overhead using configuration caching; providing real-time performance data support during kernel runtime. The pseudocode is as follows.

[0096] 。

[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0098] An embodiment of the present application also provides a thread block configuration device.

[0099] See Figure 3 , Figure 3 , which is a structural block diagram of a thread block configuration provided by an embodiment of the present application. The thread block configuration device may include: An index and feature acquisition module 31, configured to collect various system operation state indexes and obtain various preset bottleneck features; A feature score acquisition module 32, configured to calculate feature scores for various bottleneck features respectively according to various system operation state indexes to obtain each feature score; A feature weight information acquisition module 33, configured to calculate feature weights for each feature score respectively to obtain feature weight information corresponding to various bottleneck features respectively; A thread block configuration module 34, configured to perform thread block configuration according to various system operation state indexes and each feature weight information.

[0100] Through the present application, by collecting various system operation state indexes, obtaining various preset bottleneck features, calculating feature scores for various bottleneck features respectively according to various system operation state indexes, calculating weights according to the obtained feature scores, and then combining various system operation state indexes and each feature weight information to perform thread block configuration, the influence of various bottleneck features on thread block configuration is fully considered, and a specific thread block configuration strategy is adopted according to the feature weight information of various bottleneck features. Therefore, technical problems such as inability to adapt to dynamic load changes, high maintenance costs, high data acquisition costs, and high complexity of model deployment and migration can be solved, and the technical effects of being able to adapt to dynamic load changes automatically, saving maintenance costs and data acquisition costs, and being able to dynamically generate the optimal thread block size can be achieved.

[0101] In a specific implementation manner of the present application, the index and feature acquisition module 31 is specifically a module that obtains various system operation state indexes by performing index asynchronous sampling.

[0102] In a specific implementation manner of the present application, the device may further include: An index weight acquisition module, configured to obtain preset index weights corresponding to each preset index dimension; An index normalization module, configured to perform normalization processing on various system operation state indexes to obtain each normalized index value; A dimension score calculation module, which is used to calculate the dimension scores corresponding to each preset index dimension according to the respective normalized index values; A comprehensive score obtaining module, which is used to perform weighted summation calculation on the dimension scores according to each preset index weight to obtain a comprehensive score; A sampling frequency adjustment module, which is used to adjust the index sampling frequency according to the comprehensive score.

[0103] In a specific embodiment of the present application, the sampling frequency adjustment module may include: A comprehensive score upper and lower limit obtaining sub-module, which is used to obtain a preset comprehensive score upper limit and a comprehensive score lower limit; A first preset frequency adjustment sub-module, which is used to adjust the index sampling frequency to a first preset frequency when the comprehensive score exceeds the comprehensive score upper limit; A second preset frequency adjustment sub-module, which is used to adjust the index sampling frequency to a second preset frequency when the comprehensive score is lower than the comprehensive score lower limit; Wherein, the first preset frequency is less than the second preset frequency.

[0104] In a specific embodiment of the present application, the device may further include: A smoothing processing module, which is used to perform smoothing processing on each system operation status index after collecting each system operation status index.

[0105] In a specific embodiment of the present application, the device may further include: An index threshold obtaining module, which is used to obtain the index thresholds corresponding to each system operation status index after collecting each system operation status index; A thread block size adjustment module, which is used to adjust the thread block size when it is determined that the system is in a bottleneck scenario according to each system operation status index and each index threshold.

[0106] In a specific embodiment of the present application, the index threshold obtaining module may include: A historical index set obtaining sub-module, which is used to obtain the historical index sets corresponding to each system operation status index; An average value calculation sub-module, which is used to calculate the average value of each historical index set respectively to obtain the historical index average values corresponding to each historical index set; A variance calculation sub-module, which is used to calculate the variance of each historical index set respectively to obtain the historical index variances corresponding to each historical index set; An index threshold calculation sub-module, which is used to calculate the index thresholds corresponding to each system operation status index according to each system operation status index, each historical index average value and each historical index variance.

[0107] In a specific implementation manner of the present application, the thread block size adjustment module may include: A memory - limited scenario determination sub - module, configured to determine that the current bottleneck scenario is a memory - limited scenario when the memory bandwidth utilization rate among the system operation state indicators exceeds the current memory bandwidth utilization rate threshold and the first - level cache hit rate is lower than the current first - level cache hit rate threshold; A thread block size reduction sub - module, configured to reduce the thread block size according to the memory - limited scenario.

[0108] In a specific implementation manner of the present application, the thread block size adjustment module may include: A computation - limited scenario determination sub - module, configured to determine that the current bottleneck scenario is a computation - limited scenario when the stream processor occupancy rate among the system operation state indicators is lower than the current stream processor occupancy rate threshold and the warp efficiency exceeds the current memory bandwidth utilization rate threshold; A thread block size increase sub - module, configured to increase the thread block size according to the computation - limited scenario.

[0109] In a specific implementation manner of the present application, the thread block size adjustment module is specifically a module that increases the thread block size when it is determined according to the system operation state indicators and each index threshold that the current bottleneck scenario is a computation - limited and high - computation - density scenario.

[0110] In a specific implementation manner of the present application, the thread block configuration module 34 may include: A candidate item acquisition sub - module, configured to acquire each preset thread block size candidate item; A candidate item selection sub - module, configured to select the current thread block size candidate item from each thread block size candidate item according to the system operation state indicators and each feature weight information; A system performance feedback value acquisition sub - module, configured to acquire the system performance feedback value during the actual operation of the system based on the current thread block size candidate item; A thread block configuration sub - module, configured to perform adaptive trade - off reinforcement learning between exploration and exploitation according to the system performance feedback value to obtain the target thread block size, and configure the thread block of the system to the target thread block size.

[0111] In a specific implementation manner of the present application, the candidate item acquisition sub - module may include: A hardware resource constraint acquisition unit, configured to acquire the system hardware resource constraint; A candidate item determination unit, configured to determine each thread block size candidate item according to the system hardware resource constraint.

[0112] For the description of the features in the embodiments corresponding to the thread block configuration device, reference may be made to the relevant description of the embodiments corresponding to the thread block configuration method, which will not be elaborated here one by one.

[0113] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described thread block configuration method embodiments.

[0114] An embodiment of the present application further provides a non-volatile computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described thread block configuration method embodiments when running.

[0115] In an exemplary embodiment, the above non-volatile computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.

[0116] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described thread block configuration method embodiments are implemented.

[0117] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described thread block configuration method embodiments are implemented.

[0118] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0119] The above has introduced in detail a thread block configuration method, device, readable storage medium, and program product provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A thread block configuration method, characterized in that, including: Collecting the operation status indicators of each system and obtaining various preset bottleneck features; Calculating the feature scores for each type of bottleneck feature based on the operation status indicators of each system to obtain each feature score; Calculating the feature weights for each feature score respectively to obtain the feature weight information corresponding to each type of bottleneck feature; Performing thread block configuration according to the operation status indicators of each system and the feature weight information.

2. The thread block configuration method according to claim 1, wherein Collecting the operation status indicators of each system, including: Obtaining the operation status indicators of each system by performing asynchronous sampling of indicators.

3. The thread block configuration method according to claim 1, wherein After collecting the operation status indicators of each system, it further includes: Obtaining the preset indicator weights corresponding to each preset indicator dimension; Normalizing the operation status indicators of each system to obtain each normalized indicator value; Calculating the dimension scores corresponding to each preset indicator dimension according to each normalized indicator value; Performing weighted summation calculation on each dimension score according to each preset indicator weight to obtain a comprehensive score; Adjusting the indicator sampling frequency according to the comprehensive score.

4. The thread block configuration method according to claim 3, wherein Adjusting the indicator sampling frequency according to the comprehensive score, including: Obtaining the preset upper limit and lower limit of the comprehensive score; When the comprehensive score exceeds the upper limit of the comprehensive score, adjusting the indicator sampling frequency to the first preset frequency; When the comprehensive score is lower than the lower limit of the comprehensive score, adjusting the indicator sampling frequency to the second preset frequency; wherein, the first preset frequency is less than the second preset frequency.

5. The thread block configuration method according to claim 1, wherein After collecting the operation status indicators of each system, it further includes: Smoothing the operation status indicators of each system.

6. The thread block configuration method according to claim 1, wherein After collecting the operation status indicators of each system, it further includes: Obtaining the indicator thresholds corresponding to the operation status indicators of each system; When it is determined that the system is in a bottleneck scenario according to the operation status indicators of each system and the indicator thresholds, adjusting the thread block size.

7. The thread block configuration method according to claim 6, wherein Obtaining the indicator thresholds corresponding to the operation status indicators of each system, including: Obtaining the historical indicator sets corresponding to the operation status indicators of each system; Calculating the mean values of each historical indicator set respectively to obtain the historical indicator means corresponding to each historical indicator set; Calculating the variances of each historical indicator set respectively to obtain the historical indicator variances corresponding to each historical indicator set; Calculating the indicator thresholds corresponding to the operation status indicators of each system according to the operation status indicators of each system, the historical indicator means of each historical indicator set, and the historical indicator variances of each historical indicator set.

8. The thread block configuration method according to claim 6, wherein, When it is determined that the system is in a bottleneck scenario according to the operation status indicators of each system and the indicator thresholds, adjusting the thread block size, including: When the memory bandwidth utilization rate in the operation status indicators of each system exceeds the current memory bandwidth utilization rate threshold and the first-level cache hit rate is lower than the current first-level cache hit rate threshold, determining the current bottleneck scenario as a memory-limited scenario; Downscaling the thread block size according to the memory-limited scenario.

9. The thread block configuration method according to claim 6, wherein When it is determined that the system is in a bottleneck scenario according to the operation status indicators of each system and the indicator thresholds, adjusting the thread block size, including: When the occupancy rate of the stream processor in the operation status indicators of each system is lower than the current stream processor occupancy rate threshold and the warp efficiency exceeds the current memory bandwidth utilization rate threshold, determining the current bottleneck scenario as a compute-limited scenario; Increase the thread block size according to the computationally limited scenario.

10. The thread block configuration method according to claim 6, wherein When it is determined that the system is in a bottleneck scenario based on each system operating status indicator and each indicator threshold, perform thread block size adjustment, including: When it is determined that the current bottleneck scenario is a computationally limited and high computational density scenario based on each system operating status indicator and each indicator threshold, increase the thread block size.

11. The thread block configuration method according to claim 1, wherein Perform thread block configuration according to each system operating status indicator and each feature weight information, including: Obtain each preset thread block size candidate; Select the current thread block size candidate from each thread block size candidate according to each system operating status indicator and each feature weight information; Obtain the system performance feedback value in the actual operation of the system based on the current thread block size candidate; Perform adaptive trade-off reinforcement learning between exploration and exploitation according to the system performance feedback value to obtain the target thread block size, and configure the thread block of the system as the target thread block size.

12. The thread block configuration method according to claim 11, wherein Obtain each preset thread block size candidate, including: Obtain the system hardware resource constraint; Determine each thread block size candidate according to the system hardware resource constraint.

13. An electronic device, characterized in that, Include: A memory for storing a computer program; A processor for implementing the steps of the thread block configuration method according to any one of claims 1 to 12 when executing the computer program.

14. A non-volatile computer-readable storage medium, characterized in that, A computer program is stored in the non-volatile computer-readable storage medium, wherein the computer program implements the steps of the thread block configuration method according to any one of claims 1 to 12 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the thread block configuration method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Patent Citations

  • Design for computer performance self-adjusting system

    CN103116538A

  • Thread load-based software performance bottleneck indication method and system

    CN111949482A

  • Thread allocation method and device, computer equipment and storage medium

    CN112162861A

  • Data acquisition method, device and system

    CN113129473A

  • Software acquisition data storage performance improvement method and system

    CN117420967A