Method, apparatus, and storage medium for adjusting data chunk size
Patent Information
- Application Number
- CN202610692050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-29
AI Technical Summary
这导致许多面向具体场景快速开发的算子无法充分覆盖各类输入情形,难以获得理想的运行性能,与Triton降低算子开发难度、减轻开发者心智负担的设计初衷产生矛盾
[0015]本申请实施例还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述调整数据分块大小的方法。
Smart Images

Figure CN122837841A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device, and storage medium for adjusting the size of data blocks. Background Technology
[0002] Triton is a high-performance programming language and compiler framework for AI. Its programming model is based on the concept of block partitioning, allowing operator developers to implement high-performance operators simply by describing data blocks and computational logic, without having to manually handle low-level details such as thread scheduling and shared memory allocation. In Triton kernel functions, block size is a key parameter affecting runtime performance; using different block sizes for the same kernel function can produce significant performance differences.
[0003] To address this, Triton offers two parameter specification methods: automatic tuning and heuristics. Automatic tuning involves developers providing multiple candidate block configurations, with the compiler sequentially compiling and running each set of parameters and selecting the shortest-running set as the final configuration. Heuristics, on the other hand, allows developers to provide parameter values or calculation rules, which directly participate in the compilation and execution. However, the quality of parameter configuration in both methods heavily relies on the developer's experience in optimizing performance for specific hardware. Whether defining the candidate set range for automatic tuning or writing the value selection rules for heuristics, developers need a comprehensive understanding of operator behavior under different input shapes and the hardware characteristics of different AI chips to consistently achieve optimal configurations under dynamically changing input conditions. When operators face a wide range of combinations of input dimensions such as batch size, sequence length, and hidden layer dimensions, the workload of manual enumeration and calculation increases exponentially. This results in many operators designed for rapid development in specific scenarios failing to adequately cover various input scenarios and achieving ideal performance, contradicting Triton's design goal of reducing operator development difficulty and alleviating the mental burden on developers.
[0004] Therefore, a solution is urgently needed to address the above problems. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and storage medium for adjusting the size of data blocks to address the deficiencies in the prior art.
[0006] This application provides a method for adjusting the size of data blocks, including: Static analysis was performed on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; Obtain the actual input tensor of the kernel function, and determine the dimension value of the actual input tensor; Based on the mapping relationship, the upper bound of the block size parameter to be used by the kernel function is adjusted using the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; Based on the instruction constraints of the target hardware platform, the lower bound of the block size parameter after the upper bound adjustment is adjusted so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0007] According to an embodiment of this application, a method for adjusting data block size is provided, wherein static analysis of the source code of the kernel function is performed to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function, including: Parse the abstract syntax tree of the kernel function to identify memory access operations and matrix multiplication operations in the kernel function; The addressing expressions used in the memory access operation and the matrix multiplication operation are traced back to their variables, and the correspondence between each block size parameter and the input tensor dimension is determined based on the tracing results to obtain the mapping relationship.
[0008] According to an embodiment of this application, a method for adjusting the size of data blocks is provided, wherein the memory access operation includes a loading operation based on a tensor descriptor; The process of tracing variables in the addressing expressions used in the memory access operation and the matrix multiplication operation, and determining the correspondence between each block size parameter and the input tensor dimension based on the tracing results, to obtain the mapping relationship, includes: Parse the assignment statements about descriptor block shapes in the kernel function and associated pre-hook functions to obtain at least one set of candidate block shapes corresponding to the descriptor; The set of block size parameters appearing in the addressing expression of the loading operation is matched with the candidate block shape, and the block shape actually associated with the loading operation is determined as the block size parameter corresponding to the descriptor. The mapping relationship is constructed based on the block size parameters associated with the canonical access mode of the descriptor.
[0009] According to an embodiment of this application, a method for adjusting the size of data blocks is provided, wherein the upper bound of the block size parameter to be used by the kernel function is adjusted based on the mapping relationship and the dimension value, including: For any group of block size parameters in the mapping relationship and the dimension value of the actual input tensor of the mapping, if the current value of the block size parameter is greater than the dimension value, the value of the block size parameter is adjusted to the minimum value that is not less than the dimension value and is an integer power of 2.
[0010] According to an embodiment of this application, a method for adjusting the size of data blocks, wherein the lower bound adjustment of the block size parameter after the upper bound adjustment is based on the instruction constraints of the target hardware platform, includes: For the inner dimension parameter in the block size parameter associated with the matrix multiplication operation, the value is restricted to be no less than the minimum inner dimension value required by the hardware matrix multiplication accumulation instruction; For the outer dimension parameter in the block size parameter associated with the matrix multiplication operation, obtain the hardware-specified alignment byte count, the adjusted inner dimension parameter value, and the number of bytes occupied by the tensor element type. Based on the alignment byte count, the adjusted inner dimension value, and the number of bytes occupied, calculate the minimum allowable value of the outer dimension that satisfies the hardware alignment constraint, and adjust the value of the outer dimension parameter to be no less than the minimum allowable value of the outer dimension.
[0011] According to an embodiment of this application, a method for adjusting the size of data blocks is provided, wherein the current block size parameter to be used is determined by a set of candidate block configurations provided by an automatic tuning mechanism; The method further includes: The final block configuration obtained after completing the upper and lower bound adjustments is deduplicated. If the final chunk configuration is determined to be the same as a historical configuration that has been evaluated under the same input conditions, the runtime performance evaluation for the final chunk configuration is skipped, and the stored best-performing configuration is reused.
[0012] According to an embodiment of this application, a method for adjusting the size of data blocks, when the kernel function simultaneously applies both an automatic tuning decorator and a heuristic decorator, the method further includes: The upper bound adjustment, the lower bound adjustment, and the deduplication process are performed on the candidate block configuration controlled by the automatic tuning decorator. When the heuristic decorator is run, the upper bound adjustment and the lower bound adjustment are performed again on the block size parameters calculated by the heuristic decorator; The static analysis of the source code of the kernel function is performed only once, and the resulting mapping relationship is shared in both adjustment processes.
[0013] This application also provides an apparatus for adjusting the size of data blocks, comprising: The analysis and identification module is used to perform static analysis on the source code of the kernel function and identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function. The dimension determination module is used to obtain the actual input tensor of the kernel function and determine the dimension value of the actual input tensor. The upper bound adjustment module is used to adjust the upper bound of the block size parameter to be used by the kernel function according to the mapping relationship and the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; The lower bound adjustment module is used to adjust the lower bound of the block size parameters after the upper bound adjustment based on the instruction constraints of the target hardware platform, so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0014] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for adjusting the data block size as described above.
[0015] This application also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method for adjusting the size of data blocks as described above.
[0016] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method for adjusting data block size as described above.
[0017] This application provides a method, apparatus, electronic device, and storage medium for adjusting data block size. The method involves statically analyzing the source code of a kernel function to identify the mapping relationship between block size parameters and input tensor shape parameters; obtaining the actual input tensor of the kernel function and determining its dimension value; adjusting the upper bound of the currently used block size parameters of the kernel function based on the mapping relationship and the dimension value, ensuring that the value of any adjusted block size parameter is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; and adjusting the lower bound of the adjusted block size parameters based on the instruction constraints of the target hardware platform, ensuring that the value of any adjusted block size parameter is not less than the minimum allowable value determined by the hardware platform. Therefore, this application embodiment automatically identifies the mapping relationship between the block size parameter and the shape of the input tensor in the kernel function through static analysis, so that the compiler can know which block parameter corresponds to which tensor dimension without relying on the developer's experience. On this basis, the block size parameter is automatically shrunk based on the upper bound of the input size in combination with the actual input tensor dimension value, and the block size parameter is protected by the lower bound according to the instruction constraints of the target hardware platform. Thus, when the shape of the input tensor changes dynamically over a wide range, the same kernel function can automatically obtain a block configuration that does not exceed the actual tensor size and is not lower than the minimum effective granularity of the hardware instruction. Without the need for the developer to manually customize multiple sets of candidate parameters or calculation rules for different input conditions, it takes into account both the running performance of the operator and the execution efficiency of the hardware instruction. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the method for adjusting data block size provided in an embodiment of this application.
[0020] Figure 2 This is a two-stage adjustment flowchart for the Autotune × Heuristic nested scenario provided in the embodiments of this application.
[0021] Figure 3 This is a complete flowchart of the method for adjusting the size of data blocks provided in the embodiments of this application.
[0022] Figure 4 This is a schematic diagram of the structure of the device for adjusting the size of data blocks provided in the embodiments of this application.
[0023] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the embodiments of this application.
[0025] The following description, in conjunction with the accompanying drawings, describes a method, apparatus, electronic device, and storage medium for adjusting the size of data blocks according to embodiments of this application.
[0026] Figure 1 This is a flowchart illustrating the method for adjusting data block size provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following: Step 100: Perform static analysis on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function.
[0027] It should be noted that the method for adjusting the data block size provided in this application embodiment is mainly applied to the scenario of automatically optimizing the block size parameter of the kernel function in the Triton compiler framework.
[0028] Specifically, a kernel function refers to a computation function declared through decorators in the Triton programming language and executed in parallel on AI chips such as GPUs (Graphics Processing Units). Internally, it performs piecewise computation on the input tensor using a block-based approach. The block size parameter refers to the variable in the kernel function used to define the size of the data block processed in each iteration. It usually appears in the form of "BLOCK_M", "BLOCK_N", "BLOCK_K", etc., and they are key operating parameters that affect register usage, shared memory usage, and instruction execution efficiency. The input tensor shape parameter refers to the parameter describing the size of each dimension of the input tensor. For example, in matrix multiplication, M, N, and K represent the number of rows, columns, and inner dimensions of the two input matrices, respectively.
[0029] This embodiment first performs static analysis on the source code of the kernel function. That is, during the compilation stage, the mapping relationship between the block size parameter and the input tensor shape parameter is automatically established by parsing the abstract syntax tree. For example, if there is a memory access expression of the form "tl.load(A + ... + tl.arange(0, BLOCK_M)" in the kernel function, and the addressing variable is traced up to the row dimension parameter M of matrix A, then it is determined that there is a correspondence between BLOCK_M and M.
[0030] Step 200: Obtain the actual input tensor of the kernel function and determine the dimension value of the actual input tensor.
[0031] Step 300: Based on the mapping relationship, use the dimension value to adjust the upper bound of the block size parameter to be used by the kernel function, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter.
[0032] In the actual operation phase, after obtaining the actual input tensor of the kernel function and determining its dimensions, the upper bound is adjusted according to the aforementioned mapping relationship: for example, when the M dimension of the input tensor is 1, if BLOCK_M in the candidate block configuration is 32, it is adjusted down to the minimum value of 2 that is no greater than 1, that is, adjusted to 1, so as to avoid allocating invalid blocks that exceed the actual required data range.
[0033] Step 400: Based on the instruction constraints of the target hardware platform, adjust the lower bound of the block size parameters after the upper bound adjustment, so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0034] Based on this, the lower bound is further adjusted according to the instruction constraints of the target hardware platform. The instruction constraints include the minimum effective dimension value required by the hardware matrix multiplication and accumulation unit and the alignment requirements of the TMA (Tensor Memory Accelerator) data transport unit. For example, on the NVIDIA Hopper architecture chip, in order to ensure the normal execution of the matrix multiplication and accumulation instruction, the block size corresponding to the inner dimension K must not be less than 16. If the value is 8 after the upper bound adjustment, it needs to be increased to 16. The block size corresponding to the M dimension needs to be protected by calculating a minimum allowable value based on the number of alignment bytes, the current K dimension size and the number of element bytes.
[0035] By automatically identifying the mapping relationship and making bidirectional adjustments based on the actual input size and hardware constraints, this method enables the same kernel function to obtain a block configuration that neither exceeds the actual tensor boundary nor violates the minimum execution granularity requirement when facing dynamically changing input shapes, without requiring developers to manually write complex parameter candidate sets or calculation rules. This balances operator performance and instruction execution efficiency.
[0036] The above describes the steps of the method for adjusting data block size provided in the embodiments of this application. As can be seen from the above description, the method for adjusting data block size provided in the embodiments of this application involves: statically analyzing the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; obtaining the actual input tensor of the kernel function and determining the dimension value of the actual input tensor; adjusting the upper bound of the block size parameter currently to be used in the kernel function based on the mapping relationship and the dimension value, so that the value of any adjusted block size parameter is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; and adjusting the lower bound of the block size parameter after the upper bound adjustment based on the instruction constraints of the target hardware platform, so that the value of any adjusted block size parameter is not less than the minimum allowable value determined by the hardware platform. Therefore, this application embodiment automatically identifies the mapping relationship between the block size parameter and the shape of the input tensor in the kernel function through static analysis, so that the compiler can know which block parameter corresponds to which tensor dimension without relying on the developer's experience. On this basis, the block size parameter is automatically shrunk based on the upper bound of the input size in combination with the actual input tensor dimension value, and the block size parameter is protected by the lower bound according to the instruction constraints of the target hardware platform. Thus, when the shape of the input tensor changes dynamically over a wide range, the same kernel function can automatically obtain a block configuration that does not exceed the actual tensor size and is not lower than the minimum effective granularity of the hardware instruction. Without the need for the developer to manually customize multiple sets of candidate parameters or calculation rules for different input conditions, it takes into account both the running performance of the operator and the execution efficiency of the hardware instruction.
[0037] Based on the above embodiments, in this embodiment, step 100 performs static analysis on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function, including: Step 110: Parse the abstract syntax tree of the kernel function and identify the memory access operations and matrix multiplication operations in the kernel function.
[0038] Step 120: Perform variable tracing on the addressing expressions used in the memory access operation and the matrix multiplication operation, and determine the correspondence between each block size parameter and the input tensor dimension based on the tracing result to obtain the mapping relationship.
[0039] It should be noted that the memory access operation includes loading operations based on tensor descriptors.
[0040] Step 120 specifically includes: Step 121: Parse the assignment statements about the descriptor block shape in the kernel function and the associated pre-hook function to obtain at least one set of candidate block shapes corresponding to the descriptor.
[0041] Step 122: Match the set of block size parameters appearing in the addressing expression of the loading operation with the candidate block shape, and determine the block shape actually associated with the loading operation as the block size parameter corresponding to the descriptor.
[0042] Step 123: Construct the mapping relationship based on the block size parameters associated with the canonical access mode of the descriptor.
[0043] Specifically, memory access operations and matrix multiplication operations are identified by parsing the abstract syntax tree (AST) of the kernel function. The AST is a tree-like data structure generated by the compiler after parsing the source code, which can structurally express variable definitions, function calls, and computational expressions in the code. Memory access operations refer to the behavior of loading input tensors in the kernel function, including the regular tl.load loading operation and loading operations based on tensor descriptors. A tensor descriptor is a structure created through the tl.make_tensor_descriptor interface that encapsulates information such as the tensor's base address, shape, stride, and block shape, and is used to cooperate with the TMA hardware unit to achieve efficient multidimensional data transfer. Matrix multiplication operations refer to tensor-matrix multiplication operations performed in the kernel function through the tl.dot interface.
[0044] After identifying these key operations, further investigation was conducted on the addressing expressions used within them. Addressing expressions are expressions that calculate the index of the data loading or storage location, such as "pid_m". BLOCK_M +tl.arange(0, BLOCK_M)”; Variable tracing refers to starting from each variable in the addressing expression, recursively tracing its definition source upwards to determine the correspondence between the block size parameter and the input tensor dimension.
[0045] Specifically, when memory access operations involve tensor descriptors, the assignment statements regarding the descriptor block shape in the kernel function and its associated pre-hook functions are first parsed. Pre-hook functions are auxiliary functions that run before the kernel function is executed and are often used to dynamically set attributes such as block_shape of the descriptor during runtime. For example, a statement like "nargs["a_desc"].block_shape = [BLOCK_M, BLOCK_K]" may have multiple candidate assignments under different conditional branches due to different layouts. By parsing these assignment statements, at least one set of candidate block shapes corresponding to the descriptor can be obtained. Subsequently, the set of actual block size parameters appearing in the addressing expression of the current loading operation is matched with the above candidate block shapes to accurately select the set that best matches the current addressing mode. This avoids incorrect mappings caused by the existence of two usage modes of descriptors: canonical access and transpose access. Canonical access refers to loading according to the block shape dimension order originally defined by the descriptor, while transpose access refers to swapping dimensions after loading using the tl.trans operation. This step uses the block size parameters associated with the canonical access mode as the standard to construct the final mapping relationship, thereby ensuring that the mapping information on which subsequent upper and lower bound adjustments are based is accurate.
[0046] The method for adjusting data block size provided in this embodiment automatically establishes a mapping relationship between block size parameters and input tensor dimensions by parsing the abstract syntax tree and tracing the addressing expressions in memory access and matrix multiplication operations. This allows the compiler to know the tensor dimensions corresponding to each block parameter without relying on manual annotation by the developer. Based on this, for loading operations based on tensor descriptors, the method parses the candidate block shape assignment statements in the pre-hook function and matches them with the actual addressing expressions. Then, it constructs a mapping based on the canonical access mode, effectively solving the problem of mapping ambiguity when descriptors have different block shape candidates under multiple conditional branches and when canonical access and transpose access coexist. This ensures that the mapping information on which subsequent automatic adjustment depends is accurate and reliable.
[0047] Based on the above embodiments, in this embodiment, step 300, according to the mapping relationship and using the dimension value, adjusts the upper bound of the block size parameter currently to be used by the kernel function, including: Step 310: For any group of block size parameters in the mapping relationship and the dimension value of the actual input tensor of the mapping, if the current value of the block size parameter is greater than the dimension value, adjust the value of the block size parameter to the minimum value that is not less than the dimension value and is an integer power of 2.
[0048] Specifically, the upper bound adjustment is implemented as follows: For each block size parameter in the mapping relationship established by static analysis and its mapped actual input tensor dimension value, when the current value of the block size parameter is greater than the corresponding dimension value, it is adjusted down to the minimum value that is not less than the dimension value and is an integer power of 2. Here, the current value refers to the value of the block size parameter in the block configuration currently to be used by the kernel function. It may come from the candidate block configuration set provided by the automatic tuning mechanism, or it may come from the result calculated by heuristic rules.
[0049] Taking the matrix multiplication kernel function as an example, if static analysis has determined that the block size parameter BLOCK_M maps to the M dimension of the input tensor A, but the actual runtime input tensor A has an M dimension value of 1, then if the current value of BLOCK_M is 32, since 32 is greater than 1, it will be adjusted to 1 according to the minimum value of an integer power of 2 that is not less than 1; if the M dimension is 3 and BLOCK_M is 32, it will be adjusted to the minimum value of an integer power of 2 that is not less than 3, which is 4; if the M dimension is 128 and BLOCK_M is 64, since 64 is not greater than 128, no adjustment will be triggered. The reason for using an integer power of 2 as the adjustment target is that the block size is usually taken as a power of 2, which can better adapt to the memory alignment requirements and thread bundle scheduling granularity of the GPU hardware, avoiding the memory bandwidth reduction caused by unaligned access.
[0050] The method for adjusting the size of data blocks provided in this embodiment can automatically shrink the candidate block configuration, which was originally designed for a larger input size, to a reasonable range when faced with a small input size by adjusting the upper bound. This avoids a large number of invalid padding calculations caused by the blocks exceeding the actual tensor boundaries, thereby reducing the waste of registers and shared memory and improving the operator execution efficiency under small-scale inputs.
[0051] Based on the above embodiments, in this embodiment, step 400, based on the instruction constraints of the target hardware platform, adjusts the lower bound of the block size parameters after the upper bound adjustment, including: Step 410: For the inner dimension parameter in the block size parameter associated with the matrix multiplication operation, restrict the value to be no less than the minimum inner dimension value required by the hardware matrix multiplication accumulation instruction.
[0052] Step 420: For the outer dimension parameter in the block size parameter associated with the matrix multiplication operation, obtain the hardware-specified alignment byte count, the adjusted inner dimension parameter value, and the number of bytes occupied by the tensor element type. Based on the alignment byte count, the adjusted inner dimension value, and the number of bytes occupied, calculate the minimum allowable value of the outer dimension that satisfies the hardware alignment constraint, and adjust the value of the outer dimension parameter to be no less than the minimum allowable value of the outer dimension.
[0053] Specifically, the implementation of the lower bound adjustment further considers the execution constraints of hardware instructions to avoid the hardware failing to function effectively due to excessively small block sizes after the upper bound adjustment. Matrix multiplication refers to the tensor matrix multiplication and accumulation operation executed in the kernel function through the tl.dot interface. The two input operands involved have semantically outer and inner dimensions: taking the multiplication of matrix A (M, K) and matrix B (K, N) as an example, K is the common inner dimension of the two operands, while M and N are their respective outer dimensions. Correspondingly, the block size parameters associated with matrix multiplication operations are also divided into inner dimension parameters and outer dimension parameters. For example, BLOCK_K is the inner dimension parameter, while BLOCK_M and BLOCK_N are the outer dimension parameters. For the inner dimension parameter, the hardware matrix multiplication and accumulation instruction has a minimum valid dimension requirement. For example, on the NVIDIA Hopper architecture chip, in order to ensure the normal execution of the matrix multiplication and accumulation instruction and to give full play to the hardware computing power, the minimum allowed value of the inner dimension is usually 16. If BLOCK_K is shrunk to 8 after the upper bound adjustment, it needs to be further adjusted to 16 to meet the constraint. For the outer dimension parameter, its lower bound constraint comes from the data alignment requirements of the TMA tensor memory accelerator. Specifically, when TMA performs data transfer, it requires that the amount of data loaded each time meets certain byte alignment conditions. This alignment byte number is usually 128 bytes on the Hopper architecture. When determining the minimum allowable value of the outer dimension, the alignment byte number, the actual value of the current inner dimension parameter after adjustment, and the number of bytes occupied by the tensor element type need to be considered. For example, when BLOCK_K is 32 and the element type is bf16 (each element occupies 2 bytes), in order to ensure that the total amount of data loaded by TMA in a single time reaches the alignment requirement of 128 bytes, the outer dimension should be at least 128 divided by 32 and then divided by 2, which is 2, and the minimum allowable value is at least 1. If the outer dimension parameter BLOCK_M after upper bound adjustment is 1, it needs to be adjusted up to the calculated minimum allowable value of 2.
[0054] The method for adjusting the data block size provided in this embodiment ensures, through lower bound adjustment, that the adjusted block size parameters not only adapt to the actual input size, but also meet the minimum requirements of the target hardware platform for effective execution of matrix multiplication and accumulation instructions and alignment of TMA data transfer, thus avoiding the block size falling into the range where hardware refuses to execute or performance is severely degraded.
[0055] Based on the above embodiments, in this embodiment, the currently used block size parameter is determined by the candidate block configuration set provided by the automatic tuning mechanism; The method further includes: The final block configuration obtained after completing the upper and lower bound adjustments is deduplicated. If the final chunk configuration is determined to be the same as a historical configuration that has been evaluated under the same input conditions, the runtime performance evaluation for the final chunk configuration is skipped, and the stored best-performing configuration is reused.
[0056] Specifically, the block size parameter to be used is determined by the set of candidate block configurations provided by the automatic tuning mechanism. The automatic tuning mechanism refers to the runtime optimization process in the Triton compiler, which declares multiple sets of candidate runtime parameters through decorators, compiles and runs each set of parameters in sequence when a new input condition is encountered for the first time, and selects the one with the shortest execution time as the final configuration. The set of candidate block configurations is a combination of multiple blocks size parameters pre-written by the developer through the decorator.
[0057] Under this mechanism, each candidate block configuration, after undergoing the aforementioned upper and lower bound adjustments, yields a final block configuration. This embodiment avoids redundant performance evaluation overhead by deduplicating the final block configurations obtained after adjustment. Specifically, the adjusted complete parameter combination is hashed to obtain the corresponding identifier key, and a storage structure recording the evaluated configurations is maintained. Before each performance evaluation, it is first determined whether the identifier key of the current final block configuration already exists in this storage structure. If it exists, it means that the same configuration has already completed the performance evaluation under the same input conditions. In this case, the time-consuming compilation and execution process is skipped, and the stored performance results are reused. Compilation and timing are only actually executed when the identifier key does not exist. This deduplication mechanism significantly reduces the pre-running time bloat caused by the expansion of the candidate set. In a typical matrix multiplication scenario, hundreds of candidate configurations are merged into single-digit effective combinations after upper and lower bound adjustments under small-scale inputs, thus reducing the total time for automatic optimization by several times.
[0058] The method for adjusting the size of data blocks provided in this embodiment deduplicatively processes the final block configuration obtained after adjusting the upper and lower bounds, and directly reuses the stored performance results when it is determined that the current configuration is the same as the historical configuration evaluated under the same input conditions. This avoids the time bloat caused by repeated compilation and timing when multiple candidate configurations converge to the same effective configuration after adjustment. Thus, even when the candidate set size is expanded to cover a wider range of input dynamics, the total time for automatic tuning can still be kept low.
[0059] Based on the above embodiments, in this embodiment, when the kernel function simultaneously applies both an automatic tuning decorator and a heuristic decorator, the method further includes: The upper bound adjustment, the lower bound adjustment, and the deduplication process are performed on the candidate block configuration controlled by the automatic tuning decorator. When the heuristic decorator is run, the upper bound adjustment and the lower bound adjustment are performed again on the block size parameters calculated by the heuristic decorator; The static analysis of the source code of the kernel function is performed only once, and the resulting mapping relationship is shared in both adjustment processes.
[0060] Specifically, the kernel function simultaneously applies both autotuning decorators and heuristic decorators, a common practice in Triton operator development. The outer layer uses an autotuning decorator to control the selection of candidate block size parameters, while the inner layer uses a heuristic decorator to control the rule-based calculation of another set of parameters. The two layers of decorators are nested and expanded to form a structure where Autotuner wraps Heuristics wraps JITFunction. In this nested scenario, this embodiment employs a two-stage adjustment strategy: In the first stage, when the autotuning decorator iterates through each group of candidate block configurations, it performs upper bound adjustment, lower bound adjustment, and deduplication on the block size parameters it controls, writes the adjusted parameter values into the current configuration, and completes the candidate selection accordingly. In the second stage, when the heuristic decorator runs, the user's heuristic rules first calculate the block size parameter values they are responsible for. Then, in this embodiment, these calculated parameters undergo upper bound adjustment and lower bound adjustment again, the adjusted values are written back to the parameter list, and passed to the final kernel function for execution. The two adjustment phases are functionally independent. The first phase fixes the block parameters for automatic tuning control, while the second phase adjusts the heuristic parameters based on the parameters determined in the first phase. Crucially, the static analysis of the kernel function source code is performed only once. The mapping relationship between the obtained block size parameters and the input tensor shape parameters is cached and shared for reuse in both adjustments. The heuristic path avoids repeatedly incurring the overhead of abstract syntax tree parsing and variable tracing, thus achieving dual-path source coverage while maintaining extremely low additional runtime overhead.
[0061] Figure 2 This is a two-stage adjustment flowchart for a nested Autotune × Heuristic scenario provided in this application embodiment, as follows: Figure 2As shown, during execution, the system first enters the benchmark testing phase of Autotuner, traversing each group of candidate block configurations: Inside the _bench function, the meta-information of the current candidate configuration is merged with the specific block parameters into a complete parameter dictionary current. The automatic adjustment module is then called to perform upper and lower bound adjustments on it. The adjustment results are written back to current and config.kwargs to ensure that subsequent kernel function calls use the adjusted values. Subsequently, a hash identifier key is calculated for the adjusted complete parameter dictionary, and duplicates are removed using the seen_tuned_metas dictionary. Only when the identifier key has not appeared is do_bench actually executed for compilation and timing. Inside do_bench, self.fn.run is called, where self.fn is the inner Heuristics object. When `Heuristics.run` is called, the user-written heuristic rule function is executed as usual. It calculates the block parameter values it is responsible for based on the input parameters and writes them to `kwargs`. Then, a new automatic adjustment step is introduced: the input parameters are merged with all currently determined block parameter values to form the `current` dictionary. The `auto_adjust_block_sizes_for_heuristics` function is called. This function reuses the same static analysis results and performs upper and lower bound adjustments on the block parameters calculated by the heuristic rules. The adjustment results are written back to the parameter key name in `kwargs` that is responsible for by the heuristic decorator. Finally, the innermost `JITFunction` is called with the adjusted complete parameters to perform the actual calculation.
[0062] The method for adjusting the size of data blocks provided in this embodiment only executes the parsing of the abstract syntax tree of the kernel function source code and the identification of the mapping relationship once throughout the process. The result is shared between the Autotuner path and the Heuristics path through a caching mechanism with the kernel function identifier as the key. This achieves the dual-path same-source adjustment capability while avoiding the additional compilation overhead caused by repeated parsing.
[0063] Figure 3 This is a complete flowchart of the method for adjusting data block size provided in the embodiments of this application. The following is a summary of the process. Figure 3 This application provides a complete description of the method for adjusting the size of data blocks.
[0064] Taking the `mm_kernel_general_host_tma` general matrix multiplication kernel function under the FlagGems Hopper architecture as an example, this kernel function adopts the TMA host style and dynamically sets the block shape of the tensor descriptor in different conditional branches according to the matrix memory layout parameters (row-major or column-major) through the pre-hook function `matmul_tma_set_block_size_hook`. Figure 3As shown, the kernel function is decorated with both `@autotune` and `@heuristics`, forming a typical structure of outer automatic tuning and inner heuristic rule nesting. Its complete execution flow is as follows: When the kernel function is called for the first time and the input shape combination triggers a new automatic tuning key value, the compiler starts the pre-run optimization process. At this time, a one-time static analysis is first performed on the kernel function source code and its associated pre-hook functions, parsing the abstract syntax tree and identifying three types of key operations: the regular loading operation `tl.load`, the loading operation based on the TMA descriptor `desc.load`, and the matrix multiplication operation `tl.dot`. By tracing the variable definition chain in the addressing expression of each loading operation, and combining the parsing and candidate matching of the descriptor block shape assignment statements in the pre-hook functions, the mapping relationship between BLOCK_M and the input tensor M dimension, BLOCK_N and N dimension, and BLOCK_K and K dimension is automatically established, and this mapping relationship is cached with the kernel function identifier as the key. The process then proceeds to the Autotuner's _bench stage, which iterates through the candidate block configuration set provided by the developer. For each candidate configuration, it is first merged with the input metadata to form a complete parameter dictionary. Then, based on the above mapping relationship and the values of each dimension of the actual input tensor, an upper bound adjustment is performed. For example, when the actual M value is 1 and the candidate BLOCK_M is 32, it is shrunk to the minimum value of 2 that is not less than 1, i.e., 1. Next, a lower bound adjustment is performed. For the inner dimension BLOCK_K associated with tl.dot, it is ensured that it is not less than 16, which is required by the hardware matrix multiplication and accumulation instruction. For the outer dimensions BLOCK_M and BLOCK_N, the minimum allowed value is calculated based on the TMA alignment requirements (Hopper architecture alignment bytes are 128) combined with the current BLOCK_K value and the number of bytes of the element type for protection. After the adjustment is completed, the final complete parameter combination is hashed to remove duplicates. If the current adjusted configuration is exactly the same as a previously evaluated configuration, the existing timing results are directly reused to avoid repeated compilation and execution. After all candidate configurations have been evaluated, the outer Autotuner selects the one with the shortest execution time as the optimal configuration, and then calls the inner Heuristics.run. Inside Heuristics.run, the user-written heuristic rule function is first executed to calculate the values of auxiliary block parameters (such as SPLIT_K). Then, the input parameters are merged with all the currently determined block parameters. The automatic adjustment module is called again to perform the same upper and lower bound adjustments on the block parameters heuristically calculated in the previous step. This step reuses the mapping relationship cached in the static analysis stage, so there is no need to re-parse the source code. Finally, the adjusted complete parameters are passed to the innermost JITFunction to perform the actual calculation.
[0065] It should be noted that during variable tracing, the analyzer employs hierarchical dependency tracking and immediately terminates its traversal to deeper levels when it encounters the input parameters of the kernel function. For example, when tracing the definition of the addressing variable ram, its calculation expression is rm % M, where rm is further derived from pid_m. The definition of BLOCK_M + tl.arange(0, BLOCK_M) means that if short-circuit control is not applied and tracing upwards continues, the M associated with the thread coordinate variable pid_m will also be included in the mapping relationship. This would incorrectly map BLOCK_M to both the input tensor dimension M and another unrelated dimension. By terminating the tracing at this layer immediately upon encountering the input parameter M, this kind of "pseudo-association" is avoided from polluting the mapping accuracy. This ensures that only the input tensor dimension that has a direct index relationship with the block size parameter is recorded in the mapping relationship.
[0066] Through the collaborative work of the above two stages, the entire technical solution achieves automatic identification, bidirectional adjustment and candidate deduplication of block parameters without changing the external API and without being completely transparent to operator developers. This ensures that the same kernel function can always obtain an effective block configuration that is both adapted to the actual size and meets the hardware instruction constraints in scenarios where the shape of the input tensor changes drastically. At the same time, it avoids the problem of time inflation in automatic tuning caused by the expansion of the candidate set.
[0067] The following describes the improvement effects of the method for adjusting the data block size provided in the embodiments of this application on multiple levels.
[0068] 1. Performance level.
[0069] (1) Significantly expands the applicable input dynamic range of the operator and improves performance: Without modifying the operator source code, it can still provide effective configuration and improve the operator performance when the dimensions such as M, N, and K are smaller than the candidate block size.
[0070] (2) Avoid hardware instruction degradation: By using the combined lower bound of MMA + TMA swizzle, the block size is prevented from entering the failure range of wgmma / mma.sync, thus maintaining the effectiveness of TMA and MMA asynchronous pipeline.
[0071] The effectiveness of this optimization method is illustrated by the test results of the two-dimensional tensor multiplication operator (mm) in the typical large model qwen3-next on NVIDIA H100 and Hygon BW1000. The mm operator is a binary operator. In the case of no transpose, the tensor shape of the first input is (M, K), and the tensor shape of the second input is (K, N). The shapes of the mm operator used in each layer of qwen3-next range from M = 1 to 16384, and the values of K and N are shown in Table 1.
[0072] Table 1
[0073] The final test results show that performance did not decline in all shapes, but performance improved significantly in some shapes (when the size is smaller in one of the dimensions of M / N / K).
[0074] Furthermore, performance monitoring tools can be used to observe that the use of registers and shared memory is significantly reduced after applying this optimization method, which is an important reason for the performance improvement. Taking the mm operator on the NVIDIA H100 chip as an example, the monitoring results of some examples using the Nsight Compute tool are shown in Table 2.
[0075] Table 2
[0076] While optimizing kernel function execution time, the Autotune process time not only doesn't increase (due to no expansion of the candidate set size), but may even decrease (due to candidate set deduplication logic). Taking the mm operator on the NVIDIA H100 chip as an example, see Table 3: Table 3
[0077] 2. Development efficiency.
[0078] (1) The writing of candidate sets can be greatly simplified: Operator developers do not need to customize candidate sets or Heuristic rules for each input range or each hardware backend; a general candidate set can obtain effective combinations under each input.
[0079] (2) No migration burden from existing operators: Triton operators, which are developed quickly for specific scenarios in the open source ecosystem, can directly enjoy automatic adjustment without rewriting each operator.
[0080] 3. Autotune time consumption.
[0081] (1) Seen_tuned_metas candidate deduplication, so that the expansion of the candidate set no longer linearly increases the Autotune time; in scenarios with a large number of candidates (such as hundreds of candidates in typical mm), the time consumption usually decreases by several times.
[0082] (2) AST analysis is cached globally by kernel function, and the analysis overhead is amortized over all input conditions, so its impact on a single run is negligible.
[0083] 4. Engineering level.
[0084] (1) It integrates with Triton’s existing Autotune protocol in a non-intrusive manner, and is completely transparent to the external APIs of the @autotune and @heuristic decorators; it can be toggled with a single click by knobs.autotuning.adjust_block_size, which facilitates grayscale introduction and rollback.
[0085] (2) The analysis and adjustment rules are concentrated in a single module, adjust_kernel_param.py, and can be reused across operators; when adding new hardware backends, only parameterized thresholds such as limit_k and limit_bytes need to be adjusted, and the rules do not need to be rewritten.
[0086] (3) Through mechanisms such as _match_bshape_by_addr and canonical / trans distinction, the analysis is still accurate under complex usages such as pre_hook containing branches and desc.load having both transposed and non-transposed access, covering all use cases of typical TMA host style GEMM in FlagGems.
[0087] 5. Autotune and Heuristic dual-path same-source coverage.
[0088] (1) One analysis, one rule, dual path coverage: For operators on the Heuristic path, there is no need to enumerate different values or calculation formulas for different input conditions and different hardware backends, thus avoiding another type of repetitive development besides the aforementioned Autotune candidate deduplication based on the adjusted parameters.
[0089] (2) Nested use with zero additional protocol cost: In the common use of stacking @autotune × @heuristics decorators, the two-stage adjustment is automatically completed by the natural nesting order of the decorators, without requiring operator developers to write any additional code or explicitly declare the stage order.
[0090] (3) Static analysis overhead is not paid repeatedly: the two paths share the same AST parsing results cached by kernel function ID, and the runtime overhead added by the Heuristic path is approximately zero (only the constant time of lower bound comparison and dictionary write-back remains).
[0091] (4) Observability of dual-path conflict: When the Heuristic calculation result conflicts with the hardware lower bound, this application prioritizes the correctness of the hardware and exposes the adjustment facts through knobs.autotuning.print to avoid silent rewriting causing difficult-to-locate accuracy or performance problems.
[0092] The apparatus for adjusting the size of data blocks provided in the embodiments of this application will be described below. The apparatus for adjusting the size of data blocks described below can be referred to in correspondence with the method for adjusting the size of data blocks described above.
[0093] Figure 4 This is a schematic diagram of the structure of the device for adjusting the size of data blocks provided in the embodiments of this application, as shown below. Figure 4 As shown in the embodiment of this application, the apparatus for adjusting the size of data blocks includes: The analysis and identification module 401 is used to perform static analysis on the source code of the kernel function and identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; The dimension determination module 402 is used to obtain the actual input tensor of the kernel function and determine the dimension value of the actual input tensor; The upper bound adjustment module 403 is used to adjust the upper bound of the block size parameter to be used by the kernel function according to the mapping relationship and the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; The lower bound adjustment module 404 is used to adjust the lower bound of the block size parameters after the upper bound adjustment based on the instruction constraints of the target hardware platform, so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0094] The apparatus for adjusting data block size provided in this application embodiment identifies the mapping relationship between block size parameters and input tensor shape parameters in the kernel function through static analysis of the kernel function's source code; obtains the actual input tensor of the kernel function and determines the dimension value of the actual input tensor; adjusts the upper bound of the block size parameters currently to be used in the kernel function based on the mapping relationship and the dimension value, so that the value of any adjusted block size parameter is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; and adjusts the lower bound of the block size parameters after the upper bound adjustment based on the instruction constraints of the target hardware platform, so that the value of any adjusted block size parameter is not less than the minimum allowable value determined by the hardware platform. Therefore, this application embodiment automatically identifies the mapping relationship between the block size parameter and the shape of the input tensor in the kernel function through static analysis, so that the compiler can know which block parameter corresponds to which tensor dimension without relying on the developer's experience. On this basis, the block size parameter is automatically shrunk based on the upper bound of the input size in combination with the actual input tensor dimension value, and the block size parameter is protected by the lower bound according to the instruction constraints of the target hardware platform. Thus, when the shape of the input tensor changes dynamically over a wide range, the same kernel function can automatically obtain a block configuration that does not exceed the actual tensor size and is not lower than the minimum effective granularity of the hardware instruction. Without the need for the developer to manually customize multiple sets of candidate parameters or calculation rules for different input conditions, it takes into account both the running performance of the operator and the execution efficiency of the hardware instruction.
[0095] Based on the above embodiments, in this embodiment, the analysis and identification module 401 is specifically used for: Parse the abstract syntax tree of the kernel function to identify memory access operations and matrix multiplication operations in the kernel function; The addressing expressions used in the memory access operation and the matrix multiplication operation are traced back to their variables, and the correspondence between each block size parameter and the input tensor dimension is determined based on the tracing results to obtain the mapping relationship.
[0096] Based on the above embodiments, in this embodiment, the memory access operation includes a loading operation based on a tensor descriptor; The device further includes a mapping relationship determination module, specifically used for: Parse the assignment statements about descriptor block shapes in the kernel function and associated pre-hook functions to obtain at least one set of candidate block shapes corresponding to the descriptor; The set of block size parameters appearing in the addressing expression of the loading operation is matched with the candidate block shape, and the block shape actually associated with the loading operation is determined as the block size parameter corresponding to the descriptor. The mapping relationship is constructed based on the block size parameters associated with the canonical access mode of the descriptor.
[0097] Based on the above embodiments, in this embodiment, the upper limit adjustment module 403 is specifically used for: For any group of block size parameters in the mapping relationship and the dimension value of the actual input tensor of the mapping, if the current value of the block size parameter is greater than the dimension value, the value of the block size parameter is adjusted to the minimum value that is not less than the dimension value and is an integer power of 2.
[0098] Based on the above embodiments, in this embodiment, the lower boundary adjustment module 404 is specifically used for: For the inner dimension parameter in the block size parameter associated with the matrix multiplication operation, the value is restricted to be no less than the minimum inner dimension value required by the hardware matrix multiplication accumulation instruction; For the outer dimension parameter in the block size parameter associated with the matrix multiplication operation, obtain the hardware-specified alignment byte count, the adjusted inner dimension parameter value, and the number of bytes occupied by the tensor element type. Based on the alignment byte count, the adjusted inner dimension value, and the number of bytes occupied, calculate the minimum allowable value of the outer dimension that satisfies the hardware alignment constraint, and adjust the value of the outer dimension parameter to be no less than the minimum allowable value of the outer dimension.
[0099] Based on the above embodiments, in this embodiment, the currently used block size parameter is determined by the candidate block configuration set provided by the automatic tuning mechanism; The device also includes a deduplication module, specifically used for: The final block configuration obtained after completing the upper and lower bound adjustments is deduplicated. If the final chunk configuration is determined to be the same as a historical configuration that has been evaluated under the same input conditions, the runtime performance evaluation for the final chunk configuration is skipped, and the stored best-performing configuration is reused.
[0100] Based on the above embodiments, in this embodiment, the device further includes an application module, specifically used for: When both auto-tuning decorators and heuristic decorators are applied to the kernel function, The upper bound adjustment, the lower bound adjustment, and the deduplication process are performed on the candidate block configuration controlled by the automatic tuning decorator. When the heuristic decorator is run, the upper bound adjustment and the lower bound adjustment are performed again on the block size parameters calculated by the heuristic decorator; The static analysis of the source code of the kernel function is performed only once, and the resulting mapping relationship is shared in both adjustment processes.
[0101] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device can be a robot or other electronic device. This electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 550. The processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a method for adjusting the size of data blocks, including: Static analysis was performed on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; Obtain the actual input tensor of the kernel function, and determine the dimension value of the actual input tensor; Based on the mapping relationship, the upper bound of the block size parameter to be used by the kernel function is adjusted using the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; Based on the instruction constraints of the target hardware platform, the lower bound of the block size parameter after the upper bound adjustment is adjusted so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0102] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application embodiment, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in at least one embodiment of this application embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0103] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the methods for adjusting data block size provided by the above methods, including: Static analysis was performed on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; Obtain the actual input tensor of the kernel function, and determine the dimension value of the actual input tensor; Based on the mapping relationship, the upper bound of the block size parameter to be used by the kernel function is adjusted using the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; Based on the instruction constraints of the target hardware platform, the lower bound of the block size parameter after the upper bound adjustment is adjusted so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0104] In another aspect, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method for adjusting the data block size provided by the methods described above, including: Static analysis was performed on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; Obtain the actual input tensor of the kernel function, and determine the dimension value of the actual input tensor; Based on the mapping relationship, the upper bound of the block size parameter to be used by the kernel function is adjusted using the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; Based on the instruction constraints of the target hardware platform, the lower bound of the block size parameter after the upper bound adjustment is adjusted so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0106] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of this application, and are not intended to limit them; although the embodiments of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for adjusting the size of data blocks, characterized in that, include: Static analysis was performed on the source code of the kernel function to identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function; Obtain the actual input tensor of the kernel function, and determine the dimension value of the actual input tensor; Based on the mapping relationship, the upper bound of the block size parameter to be used by the kernel function is adjusted using the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; Based on the instruction constraints of the target hardware platform, the lower bound of the block size parameter after the upper bound adjustment is adjusted so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
2. The method for adjusting data block size according to claim 1, characterized in that, The static analysis of the kernel function's source code to identify the mapping relationship between the block size parameters and the input tensor shape parameters includes: Parse the abstract syntax tree of the kernel function to identify memory access operations and matrix multiplication operations in the kernel function; The addressing expressions used in the memory access operation and the matrix multiplication operation are traced back to their variables, and the correspondence between each block size parameter and the input tensor dimension is determined based on the tracing results to obtain the mapping relationship.
3. The method for adjusting data block size according to claim 2, characterized in that, The memory access operations include loading operations based on tensor descriptors; The process of tracing variables in the addressing expressions used in the memory access operation and the matrix multiplication operation, and determining the correspondence between each block size parameter and the input tensor dimension based on the tracing results, to obtain the mapping relationship, includes: Parse the assignment statements about descriptor block shapes in the kernel function and associated pre-hook functions to obtain at least one set of candidate block shapes corresponding to the descriptor; The set of block size parameters appearing in the addressing expression of the loading operation is matched with the candidate block shape, and the block shape actually associated with the loading operation is determined as the block size parameter corresponding to the descriptor. The mapping relationship is constructed based on the block size parameters associated with the canonical access mode of the descriptor.
4. The method for adjusting data block size according to claim 1, characterized in that, The step of adjusting the upper bound of the block size parameter to be used by the kernel function based on the mapping relationship and the dimension value includes: For any group of block size parameters in the mapping relationship and the dimension value of the actual input tensor of the mapping, if the current value of the block size parameter is greater than the dimension value, the value of the block size parameter is adjusted to the minimum value that is not less than the dimension value and is an integer power of 2.
5. The method for adjusting data block size according to claim 2, characterized in that, The instruction constraints based on the target hardware platform, which adjust the lower bound of the block size parameters after the upper bound adjustment, include: For the inner dimension parameter in the block size parameter associated with the matrix multiplication operation, the value is restricted to be no less than the minimum inner dimension value required by the hardware matrix multiplication accumulation instruction; For the outer dimension parameter in the block size parameter associated with the matrix multiplication operation, obtain the hardware-specified alignment byte count, the adjusted inner dimension parameter value, and the number of bytes occupied by the tensor element type. Based on the alignment byte count, the adjusted inner dimension value, and the number of bytes occupied, calculate the minimum allowable value of the outer dimension that satisfies the hardware alignment constraint, and adjust the value of the outer dimension parameter to be no less than the minimum allowable value of the outer dimension.
6. The method for adjusting data block size according to any one of claims 1-5, characterized in that, The currently used block size parameter is determined by the set of candidate block configurations provided by the automatic tuning mechanism; The method further includes: The final block configuration obtained after completing the upper and lower bound adjustments is deduplicated. If the final chunk configuration is determined to be the same as a historical configuration that has been evaluated under the same input conditions, the runtime performance evaluation for the final chunk configuration is skipped, and the stored best-performing configuration is reused.
7. The method for adjusting data block size according to claim 6, characterized in that, When both automatic tuning decorators and heuristic decorators are applied to the kernel function, the method further includes: The upper bound adjustment, the lower bound adjustment, and the deduplication process are performed on the candidate block configuration controlled by the automatic tuning decorator. When the heuristic decorator is run, the upper bound adjustment and the lower bound adjustment are performed again on the block size parameters calculated by the heuristic decorator; The static analysis of the source code of the kernel function is performed only once, and the resulting mapping relationship is shared in both adjustment processes.
8. A device for adjusting the size of data blocks, characterized in that, include: The analysis and identification module is used to perform static analysis on the source code of the kernel function and identify the mapping relationship between the block size parameter and the input tensor shape parameter in the kernel function. The dimension determination module is used to obtain the actual input tensor of the kernel function and determine the dimension value of the actual input tensor. The upper bound adjustment module is used to adjust the upper bound of the block size parameter to be used by the kernel function according to the mapping relationship and the dimension value, so that the value of any block size parameter after adjustment is not greater than the corresponding dimension value of the actual input tensor mapped by the block size parameter; The lower bound adjustment module is used to adjust the lower bound of the block size parameters after the upper bound adjustment based on the instruction constraints of the target hardware platform, so that the value of any block size parameter after adjustment is not less than the minimum allowable value determined by the hardware platform.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for adjusting the size of data blocks as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for adjusting the size of data blocks as described in any one of claims 1 to 7.