Data processing method and device, equipment, storage medium and program product

By determining the index of kernel functions and thread blocks in a large language model and processing discontinuous and continuous tensor data, the problems of computing complexity and delay in the prior art are solved, and more efficient computing speed and performance improvements are achieved.

CN120523745APending Publication Date: 2025-08-22MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510457754.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The prior art cannot effectively process tensor data with discontinuous addresses, resulting in increased computational complexity and delays, and the existing optimization technology is not optimized for specific application scenarios of large language models.

Method used

By determining the kernel function and thread block of the target tensor, the thread index in the thread block, the number of elements in the element area and the step size of adjacent element, the index of each element is calculated, and the kernel function is called through the thread block to process the tensor data, supporting the input of discontinuous and continuous tensors, reducing additional continuous operations.

Benefits of technology

Improves computing efficiency, reduces computing delay and memory transfer times, and improves computing speed, especially in large-scale language model pre-training tasks, which significantly improves performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523745A_ABST
    Figure CN120523745A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, equipment, a storage medium and a program product, and the data processing method comprises the steps: determining a kernel function and a thread block which are needed when a target tensor is processed; the target tensor comprises a plurality of element areas, and threads in the thread block and the element areas have a one-to-one correspondence relationship; determining the index of each element in the target tensor based on the index of the thread in the thread block, the element number of the element region in the target dimension and the step length between adjacent elements in the element region; and calling the kernel function through the thread block, and processing the target tensor according to the index of each element. Therefore, not only can the input of discontinuous tensor data be supported, but also the input of continuous tensor data can be supported, and the application range is wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of computer technology, and in particular to a data processing method, apparatus, device, storage medium, and program product. Background Art

[0002] With the development of artificial intelligence and deep learning technologies, large language models have demonstrated powerful data processing capabilities in many fields such as natural language processing, question-answering systems, and machine translation.

[0003] Typically, data for model training is input as tensors, and kernel functions are used to sequentially operate on each element in the tensor to implement convolution, pooling, and other processing on the model data. However, existing solutions can only process tensors with consecutive addresses and cannot process tensors with discontinuous addresses. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide at least one data processing method, apparatus, device, storage medium, and program product.

[0005] The technical solution of the embodiment of the present disclosure is implemented as follows:

[0006] On the one hand, an embodiment of the present disclosure provides a data processing method, which includes: determining a kernel function and a thread block required for processing a target tensor; the target tensor includes multiple element regions, and the threads in the thread block have a one-to-one correspondence with the element regions; based on the index of the thread in the thread block, the number of elements in the element region in the target dimension, and the step size between adjacent elements in the element region, determining the index of each element in the target tensor; calling the kernel function through the thread block, and processing the target tensor according to the index of each element.

[0007] On the other hand, an embodiment of the present disclosure provides a data processing device, which includes: a determination module for determining the kernel function and thread block required for processing a target tensor; the target tensor includes multiple element regions, and the threads in the thread block have a one-to-one correspondence with the element regions; a processing module for determining the index of each element in the target tensor based on the index of the thread in the thread block, the number of elements in the element region in the target dimension, and the step size between adjacent elements in the element region; the processing module is also used to call the kernel function through the thread block to process the target tensor according to the index of each element.

[0008] On the other hand, an embodiment of the present disclosure provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0009] On the other hand, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which implements part or all of the steps in the above method when executed by a processor.

[0010] On the other hand, an embodiment of the present disclosure provides a computer program, including computer-readable codes. When the computer-readable codes are executed in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0011] On the other hand, an embodiment of the present disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method.

[0012] In the disclosed embodiment, the position of each element in a tensor is determined based on the step size between adjacent elements. This supports the input of both discontinuous and continuous tensor data, making it widely applicable. Furthermore, no additional serialization operations are required, reducing computational latency and improving computational efficiency. Cache optimization of the input tensor data before executing the operator can reduce the number of data transfers between memory and the GPU, thereby increasing computational speed.

[0013] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0015] Figure 1 A schematic diagram of the implementation process of a data processing method provided in the embodiment of the present disclosure Figure 1 ;

[0016] Figure 2 A schematic diagram of the implementation process of a data processing method provided in the embodiment of the present disclosure Figure 2 ;

[0017] Figure 3 A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure;

[0018] Figure 4 A schematic diagram of a hardware entity of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions of the present disclosure are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting the present disclosure. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0020] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0021] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are for the purpose of describing the present disclosure only and are not intended to limit the present disclosure.

[0023] In order to better understand the data processing method provided by the embodiments of the present disclosure, the solutions in the related art are first described below.

[0024] In related technologies, non-continuous input tensors need to be converted to continuous input tensors before calculation. This means that additional continuous operations are required to handle the discontinuity, increasing the complexity and latency of the calculation.

[0025] In addition, existing memory bandwidth optimization technologies focus on improving memory access efficiency and reducing data transmission time, and there is no solution specifically for optimizing memory-bottleneck operators; existing GPU hardware optimization technologies do not have solutions specifically for optimizing specific application scenarios such as large language model pre-training; existing compiler optimization technologies do not have solutions for optimizing memory-bottleneck operators; and existing software stack adaptation technologies do not have solutions for optimizing large language model pre-training on non-Nvidia GPUs.

[0026] To this end, an embodiment of the present disclosure provides a data processing method that can be executed by a processor of a computer device. The computer device may refer to a server, a laptop, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable gaming device), or other devices with data processing capabilities. The processor may refer to a graphics processing unit (GPU) or a central processing unit (CPU). Figure 1 As shown, the method includes the following steps 101 to 103:

[0027] Step 101: Determine a kernel function and a thread block required for processing a target tensor; the target tensor includes multiple element regions, and the threads in the thread block have a one-to-one correspondence with the element regions.

[0028] In neural networks, data typically exists in the form of tensors; for example, input features and hidden layer activation values. Operators are functions that perform tensor operations, defining how to perform mathematical operations on tensors. Kernel functions are functions written in low-level code that implement the specific operation logic of operators and are used to execute the operations defined by operators on hardware (e.g., GPUs, CPUs).

[0029] The target tensor refers to the tensor currently being processed. A tensor is a multidimensional array used to represent data and parameters in a neural network. For example, a tensor can include, but is not limited to, input data (text, images, speech, etc.), weights, biases, and intermediate activation values. The dimensions of a tensor can be a scalar (0-dimensional), a vector (1-dimensional), a matrix (2-dimensional), or even a higher-dimensional array.

[0030] The kernel function required for processing the target tensor refers to the underlying function that performs specific operations on the target tensor and is used to run on multiple processing cores of the GPU.

[0031] A thread block is a collection of threads. Threads within a thread block can execute on the same multiprocessor, share data, and execute synchronously. The size and number of thread blocks can be adjusted based on the computing task and hardware characteristics. Thread blocks can be one-dimensional, two-dimensional, three-dimensional, or even higher-dimensional. For example, block(16,16) represents a two-dimensional thread block, each containing 16×16 threads.

[0032] When executing operations on the GPU, a tensor can be split into multiple element regions, and each thread in a thread block is responsible for processing an element region of the tensor, where each element region includes one or more elements.

[0033] In some embodiments, the specific implementation method of step 101 can be: determining a neural network operator based on the processing requirements of the target tensor; configuring the kernel function required for processing the target tensor based on the neural network operator; and determining the thread block required for processing the target tensor based on the size of the target tensor.

[0034] Neural network operators are operators that perform specific operations on target tensors. They can be unary (e.g., activation functions) or binary (e.g., matrix multiplication). They can be memory-bound or compute-intensive. The performance of memory-bound operators is limited by memory access speed, while the performance of compute-intensive operators is limited by processor processing speed.

[0035] Memory-bound operators are operators whose performance is limited by memory access speed. When an operator is memory-bound, its execution speed depends primarily on the speed of reading and writing memory, rather than the processing speed of the processor. Examples of memory-bound operators include, but are not limited to, unary operators and binary operators.

[0036] Computationally Intensive Operators refer to operators that require a large amount of computing resources to execute. Computationally intensive operators usually involve a large number of mathematical operations; for example, matrix multiplication, convolution operations, etc.

[0037] Since neural network operators may be memory-limited operators or computationally intensive operators, the kernel function may be determined based on the memory-limited operators or the computationally intensive operators.

[0038] Step 102: Determine the index of each element in the target tensor based on the index of the thread in the thread block, the number of elements in the element region in the target dimension, and the stride between adjacent elements in the element region.

[0039] Each thread block has a unique index, which is used to identify and locate the thread block in the parallel computing environment. Each thread within a thread block also has a unique index, which is used to identify and locate the thread in the parallel computing environment. The stride between adjacent elements refers to the distance between two adjacent elements in memory. The index of each element represents its position within the element region.

[0040] In some implementations, the target dimension may be an x ​​dimension. In this case, the number of elements in the element region in the x dimension may be expressed as x_size.

[0041] In some embodiments, step 102 may be specifically implemented by determining the index of each element in the target tensor based on the thread index, the number of elements in the element region in the target dimension, and the stride between adjacent elements in different dimensions.

[0042] Step 103: Call the kernel function through the thread block and process the target tensor according to the index of each element.

[0043] In some embodiments, step 103 may be specifically implemented by: obtaining the elements in the element region corresponding to each thread according to the position of each element; and processing the elements in the element region corresponding to each thread in turn by calling a kernel function through each thread.

[0044] It should be noted that the operations performed on the target tensor in this application can be executed on the CPU or on the GPU.

[0045] It should be noted that memory-constrained operators such as unary operators and binary operators only support continuous tensor inputs and do not support discontinuous tensor inputs; and the embodiment of the present disclosure passes the step length between adjacent elements into the kernel, and determines the position of each element in the tensor based on the step length between adjacent elements. In this way, even if the address is discontinuous, the position of each element can be accurately obtained based on the step length, thereby realizing the processing of tensors with discontinuous addresses; when the addresses are continuous, the step length between adjacent elements can be regarded as one to realize the processing of tensors with continuous addresses.

[0046] In the disclosed embodiment, the position of each element in a tensor is determined based on the step size between adjacent elements. This supports the input of both discontinuous and continuous tensor data, making it widely applicable. Furthermore, no additional serialization operations are required, reducing computational latency and improving computational efficiency. Cache optimization of the input tensor data before executing the operator can reduce the number of data transfers between memory and the GPU, thereby increasing computational speed.

[0047] The present disclosure provides a data processing method, which can be executed by a processor of a computer device. Figure 2 As shown, the method includes the following steps 201 to 206:

[0048] Step 201: Determine a kernel function and a thread block required for processing a target tensor; the target tensor includes multiple element regions, and the threads in the thread block have a one-to-one correspondence with the element regions.

[0049] Here, the above step 201 corresponds to the above step 101, and the specific implementation of the above step 101 can be referred to during implementation.

[0050] In some embodiments, the specific implementation method of "determining the kernel function required for processing the target tensor" can be: determining multiple neural network operators required for processing the target tensor; based on the association relationship between the multiple neural network operators, fusing the multiple neural network operators to obtain a target neural network operator; based on the target neural network operator, configuring the kernel function.

[0051] The target neural network operator refers to the operator formed by fusing multiple neural network operators.

[0052] Fusing multiple related neural network operators into one operator can reduce the number of executed operators, thereby reducing memory access and computational overhead.

[0053] Step 202: Determine the index of each element in the target dimension based on the thread index, the number of elements in the element region in the target dimension, and the first step length.

[0054] The step length between adjacent elements includes a first step length of the target dimension and a second step length of the remaining dimensions except the target dimension.

[0055] In some embodiments, if the element region is two-dimensional, the first step size refers to the step size between adjacent elements in the x-dimension. The second step size refers to the step size between adjacent elements in the y-dimension. If the element region is three-dimensional, the first step size refers to the step size between adjacent elements in the x-dimension. The second step size refers to the step size between adjacent elements in the y-dimension and the step size between adjacent elements in the z-dimension.

[0056] In some embodiments, step 202 may be specifically implemented by performing a modulo operation on the thread index and the number of elements in the element region on the target dimension; and performing a multiplication operation on the result of the modulo operation and the first step length to obtain the index of each element on the target dimension.

[0057] For example, the formula for calculating the index of each element in the target dimension can be: size_t tx = tid% x_size*stride_x. Here, tid represents the global index of the thread; x_size represents the first step length; tid% x_size represents the number of elements to be retrieved; stride_x represents the first step length; and size_t tx represents the index of each element in the target dimension.

[0058] In some embodiments, a specific implementation method for determining the index of a thread may be: based on the index of the thread block on the target dimension in the computing space and the number of threads of the thread block on the target dimension, determining the starting index of the thread block on the target dimension; based on the starting index of the target dimension and the index of each thread on the target dimension in the thread block, determining the index of each thread.

[0059] Compute space refers to the logical or physical space used to perform computational tasks. For example, in GPU programming, compute space refers to the GPU's memory space, which includes global memory, shared memory, registers, etc. The compute space can be a grid.

[0060] If the computation space is two-dimensional, the computation space includes two dimensions: x-dimension and y-dimension. If the computation space is three-dimensional, the computation space includes three dimensions: x-dimension, y-dimension, and z-dimension.

[0061] The index of a thread block in the target dimension refers to the x-coordinate of the thread block in computation space and can be expressed as blockIdx.x. The number of threads in a thread block in the target dimension refers to the number of threads in the thread block in the x-direction and can be expressed as blockDim.x. The index of each thread in the thread block in the target dimension refers to the x-coordinate within the thread block and can be expressed as threadIdx.x.

[0062] The thread block index along the target dimension and the index of each thread within the thread block along the target dimension are built-in variables automatically provided by the parallel computing platform and programming model (CUDA) when executing the kernel function. The thread block size can be specified by the programmer and includes the number of threads in each dimension of the thread block.

[0063] In some embodiments, a specific implementation method of "determining the starting index of the thread block in the target dimension based on the index of the thread block in the target dimension in the computing space and the number of threads of the thread block in the target dimension" may be: multiplying the index of the thread block in the target dimension and the number of threads of the thread block in the target dimension to obtain the starting index of the thread block in the target dimension.

[0064] In some embodiments, the specific implementation method of "determining the index of each thread based on the starting index of the target dimension and the index of each thread on the target dimension in the thread block" can be: adding the starting index on the target dimension and the index of each thread on the target dimension in the thread block to obtain the index of each thread.

[0065] For example, the formula for calculating the thread index can be: int64_t tid = (int64_t)blockIdx.x * blockDim.x + threadIdx.x. In this example, int64_t is a data type representing a 64-bit integer; blockIdx.x represents the index of the thread block on the target dimension in the computation space; blockDim.x represents the number of threads in each thread block on the target dimension; threadIdx.x represents the index of any thread on the target dimension of the thread block; and tid represents the global index of the thread, which is used to distinguish all threads in the computation space.

[0066] Step 203: Determine the index of each element in the remaining dimensions based on the thread index, the number of elements in the element region in the target dimension, and the second step size.

[0067] In some embodiments, step 203 may be specifically implemented by performing a division operation on the thread index and the number of elements in the element region in the target dimension; and performing a multiplication operation on the result of the division and the second step length to obtain the index of each element in the remaining dimensions.

[0068] For example, let's assume the element region is two-dimensional. In this case, the second stride is the stride between adjacent elements in the y-dimension. The formula for calculating each element's y-dimension index is: size_tty = tid / x_size * stride_y. Here, size_tty represents the y-dimension index of each element, tid / x_size represents the row in which the element is located, and stride_y represents the second stride.

[0069] For non-contiguous tensor inputs, calculating the index of each element involves division, which is complex and slow. Therefore, to further improve computational efficiency and reduce computational complexity, the disclosed embodiment replaces division with shift operations.

[0070] Specifically, for any element on the last dimension of the target tensor, when the number of elements on the last dimension is 2 to the power of n, the index of the thread is shifted right by n bits to obtain the starting address of the last dimension; n is a positive integer; based on the starting address of the last dimension and the second step size, the index of the any element on the last dimension is determined.

[0071] For example, if the size of the last dimension is 4 (that is, 2 to the power of 2), the index can be shifted right by 2 bits instead of dividing by 4. This not only improves the computational efficiency, but also reduces the computational complexity, thereby improving the overall performance.

[0072] Step 204: Determine the index of each element based on the index of each element in the target dimension and the index of each element in the remaining dimensions.

[0073] Here, the above steps 202 to 204 correspond to the above step 102, and the specific implementation of the above step 102 may be referred to during implementation.

[0074] In some embodiments, the index of each element in the target dimension is combined with the index of each element in the remaining dimensions in a target format to obtain the index of each element. The target format can be an array, a coordinate, etc.

[0075] Step 205: Obtain a continuous output tensor corresponding to the target tensor according to the index of each element.

[0076] The target tensor includes a first tensor and a second tensor with discontinuous addresses. The continuous output tensor refers to a tensor with continuous addresses obtained by processing the discontinuous first tensor and the second tensor.

[0077] In some embodiments, the specific implementation method of step 205 may be: based on the size of the first tensor and the size of the second tensor, determine the memory space that meets the tensor storage requirements; read the elements in the first tensor according to the index of each element in the first tensor to obtain a first output tensor; read the elements in the second tensor according to the index of each element in the second tensor to obtain a second output tensor; store the first output tensor at a first address in the memory space, and store the second output tensor at a second address adjacent to the first address to obtain a continuous output tensor.

[0078] Memory space is used to store tensor data. The first output tensor refers to the first tensor read. The second output tensor refers to the second tensor read. The first address refers to any address in memory space, and the second address refers to an address adjacent to the first address. If the first address is determined, the second address is also determined.

[0079] In some embodiments, the memory space may refer to the memory space allocated on the host, i.e., a section of memory space allocated on the host for tensor processing. Alternatively, the memory space may refer to the memory space allocated on the GPU, i.e., a section of memory space allocated on the host for tensor processing.

[0080] In some embodiments, in order to further improve the processing performance of tensors, after obtaining the first output tensor and the second output tensor, the first output tensor needs to be stored at a first address in the memory space, and the second output tensor needs to be stored at a second address adjacent to the first address to obtain continuous output tensors.

[0081] When tensor data is discontinuous, the get_offset function is called during calculations. This function involves division, which is time-consuming, especially when the index is of type int64. Therefore, this can be optimized. Specifically, when requesting memory on the host side to initialize the output tensor, a contiguous memory space can be allocated. This way, the two inputs (the first and second tensors) are discontinuous, while the output (the continuous output tensor) is continuous. This can reduce the number of get_offset function calls by one-third, resulting in significant performance improvements.

[0082] Because GPUs have superior computing power compared to host CPUs, the processing of two inputs can be performed on the GPU. Specifically, a contiguous memory space is allocated on the GPU, the two input tensors are copied to the GPU memory space, the two input tensors are processed on the GPU, and the two output tensors are stored in contiguous addresses to obtain continuous output.

[0083] In some embodiments, the processor performs the operation of obtaining the continuous output tensors corresponding to the target tensor according to the index of each element; or, the target tensor is copied to a graphics processor, and the graphics processor performs the operation of obtaining the continuous output tensors corresponding to the target tensor according to the index of each element.

[0084] Step 206: Call the kernel function through the thread block to process the continuous output tensor.

[0085] Here, the above steps 205 to 206 correspond to the above step 103, and the specific implementation of the above step 103 may be referred to during implementation.

[0086] In some embodiments, data can also be cached and optimized, specifically: continuous output tensors are stored in a cache area; and continuous output tensors in the cache area are processed by calling kernel functions.

[0087] Before executing an operator, you can preprocess the input tensor data; for example, by optimizing the data cache. This can reduce the number of data transfers between memory and the GPU, improving computation speed. Experiment with different configurations to select the cache optimization solution that yields the best performance.

[0088] In the disclosed embodiment, the position of each element in the tensor is determined according to the step size between adjacent elements, which can support the input of non-continuous tensor data and continuous tensor data, and has a wide range of applications; and no additional serialization operations are required, which reduces computational delays and improves computational efficiency. Before executing the operator, the input tensor data is cached and optimized, which can reduce the number of data transfers between the memory and the GPU and increase the computing speed. Multiple related operators are merged into one operator to reduce the number of executed operators, thereby reducing memory access and computational overhead. When the size of the last dimension of the tensor is a power of 2, the shift operation is used instead of the division operation, which can reduce 1 / 3 of the get_offset function calls and improve the processing performance of the tensor.

[0089] The following describes the application of the data processing method provided by the embodiments of the present disclosure in actual scenarios, taking its application in the muDNN operator library as an example.

[0090] The disclosed embodiments provide a method for optimizing the performance of the muDNN operator library on GPUs. By reducing the number of get_offset calls, replacing division operations with shift operations, and directly supporting non-contiguous inputs, the muDNN library improves its computational efficiency on GPUs, particularly when handling high-performance computing tasks such as large-scale language model pre-training. These optimizations help enhance the competitiveness of the muDNN operator library and promote the development of the high-performance computing field.

[0091] Unary and binary operators when handling non-contiguous inputs. These optimizations include reducing the number of kernel get_offset calls, replacing division operations with shift operations, and optimizing performance bottlenecks in specific situations. Furthermore, the disclosed embodiments propose an innovative method that enables unary and binary operators to directly support non-contiguous inputs, thus avoiding additional serialization operations and kernel calls, further improving computational efficiency.

[0092] The specific solutions of the embodiments of the present disclosure are as follows:

[0093] 1. Directly support non-continuous input;

[0094] Currently, unary and binary operators do not support non-continuous inputs. Non-continuous tensors need to be serialized before being fed into the kernel for calculation. The method proposed in the disclosed embodiment is to pass the stride of each input dimension into the kernel and then calculate the address (index) of each element based on the stride, so that unary and binary operators can directly support non-continuous inputs. This method of calculating the address of each element based on the stride avoids additional serialization operations and kernel calls, reduces computational latency, and improves computational efficiency.

[0095] Before optimization, it was necessary to call a kernel function for continuous operations to convert the two input tensors into continuous tensors, and then call the kernel function for processing the input tensors. This process required calling three kernel functions. After optimization using the method of the embodiment of the present disclosure, only the kernel function for processing the input tensors needs to be called, saving two kernel function calls.

[0096] The following code explains:

[0097] int64_t tid=(int64_t)blockIdx.x*blockDim.x+threadIdx.x;

[0098] size_t tx=tid%x_size*stride_x;

[0099] size_t ty=tid / x_size*stride_y;

[0100] The line "int64_t tid = (int64_t)blockIdx.x * blockDim.x + threadIdx.x" calculates the global index (thread ID) of the current thread. blockIdx.x is the x-coordinate of the thread block in computation space, and blockDim.x is the number of threads in the thread block in the x-direction. Multiplying these two values ​​yields the starting global index of the current thread block in the x-direction. Then, threadIdx.x is the index of the current thread within the thread block. Adding these two values ​​yields the global index of the current thread.

[0101] The line of code “size_t tx=tid%x_size*stride_x” is used to calculate the local index of the current thread in the x dimension.

[0102] The line of code "size_t ty = tid / x_size*stride_y" is used to calculate the local index of the current thread in the y dimension.

[0103] 2. Optimize kernel index calculation method;

[0104] In the PyTorch framework, the logic for calling the muDNN operator library involves execution on the host (the CPU), where output tensors are typically created within the framework. When tensor data is discontinuous, the get_offset function is called during calculation. This involves division, which is time-consuming, especially since indices are typically int64 types. Therefore, this can be optimized. Specifically, when allocating memory on the host to initialize the output tensor, a contiguous memory space can be allocated. During calculation, when the two inputs are discontinuous and the output is contiguous, this can reduce the number of get_offset function calls by one-third, resulting in a significant performance improvement. Because index calculations are relatively time-consuming, this can significantly improve performance for discontinuous inputs. For example, with unary optimization, bandwidth increases from 85GB / s to 110GB / s when the discontinuous vlen is 1; and from 258GB / s to 308GB / s when the discontinuous vlen is aligned to 4.

[0105] 3. Use shift operations instead of division calculations;

[0106] In the muDNN operator library, processing non-contiguous inputs typically involves calculating the index of each element. However, when the last dimension of the input (dim[-1]) is a power of 2, a shift operation can be used instead of a division operation. For example, if dim[-1] is 4 (2 to the power of 2), the division by 4 can be replaced by shifting the index right by 2 bits. This not only improves computational efficiency but also reduces computational complexity, thereby improving overall performance. After unary optimization, performance in the non-contiguous model scenario increased from 308 GB / s to 346 GB / s, approaching the bandwidth of 366 GB / s in the continuous scenario under the same conditions. After binary optimization, performance in the non-broadcast non-contiguous model scenario increased from 340 GB / s to 376 GB / s; in the broadcast non-contiguous model scenario, performance increased from 425 GB / s to 464 GB / s.

[0107] It should be noted that the embodiments of the present disclosure include at least the following innovations:

[0108] 1. Pass the stride of each dimension of the input into the kernel, calculate the address of each element according to the stride, and reduce the number of kernel get_offset calls.

[0109] 2. When the last dimension of the input (dim[-1]) is a power of 2, use shift operations instead of division calculations.

[0110] 3. Open up a continuous memory space. During calculation, the two inputs are discontinuous and the output is continuous, so as to optimize the performance bottleneck in the discontinuous model scenario.

[0111] It should be noted that the embodiments of the present disclosure can at least achieve the following technical effects:

[0112] 1. Significantly improve computing efficiency: By optimizing memory bandwidth and reducing unnecessary calculations, the disclosed embodiments significantly improve the computing efficiency of the muDNN operator library on the GPU. This improvement is particularly evident when processing large-scale computing tasks such as pre-training of large language models.

[0113] 2. Reduce computational latency: By reducing the number of kernel get_offset calls and replacing division operations with shift operations, we effectively reduce computational latency and make the operator faster when processing discontinuous inputs.

[0114] 3. The disclosed embodiments have better competitive advantages in the GPU market, especially in supporting deep learning and large-scale computing.

[0115] 4. Promote the development of large language models: By optimizing operator performance on the GPU, the disclosed embodiments provide better hardware support for the training and application of large-scale language models, promoting the development of natural language processing and artificial intelligence.

[0116] 5. Improve computing power utilization efficiency: The disclosed embodiments improve the utilization efficiency of GPU hardware resources by optimizing memory usage and operator execution efficiency, which is particularly important for application scenarios that require high-performance computing with limited resources.

[0117] 6. Simplify the adaptation and maintenance of the software stack: By directly supporting non-continuous input, the disclosed embodiments simplify the adaptation process of the software stack, reduce maintenance costs, and make the update and optimization of the operator library more convenient.

[0118] Based on the foregoing embodiments, the embodiments of the present disclosure provide a data processing device, which includes the various units included and the various modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0119] Figure 3 A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure is shown in FIG. Figure 3 As shown, the data processing device 300 includes: a determination module 310 and a processing module 320, wherein:

[0120] A determination module 310 is configured to determine a kernel function and a thread block required for processing a target tensor; the target tensor includes a plurality of element regions, and threads in the thread block have a one-to-one correspondence with the element regions;

[0121] A processing module 320 is configured to determine an index of each element in the target tensor based on an index of a thread in the thread block, a number of elements in the element region in a target dimension, and a stride between adjacent elements in the element region.

[0122] The processing module 320 is further configured to call the kernel function through the thread block and process the target tensor according to the index of each element.

[0123] In some embodiments, the processing module 320 is further used to: determine the starting index of the thread block in the target dimension based on the index of the thread block in the target dimension in the computing space and the number of threads of the thread block in the target dimension; determine the index of each thread based on the starting index of the target dimension and the index of each thread in the thread block in the target dimension.

[0124] In some embodiments, the step size includes a first step size of the target dimension and a second step size of the remaining dimensions except the target dimension; the processing module 320 is specifically used to: determine the index of each element in the target dimension based on the index of the thread, the number of elements in the element area in the target dimension, and the first step size; determine the index of each element in the remaining dimensions based on the index of the thread, the number of elements in the element area in the target dimension, and the second step size; determine the index of each element based on the index of each element in the target dimension and the index of each element in the remaining dimensions.

[0125] In some embodiments, the processing module 320 is specifically used to: for any element on the last dimension of the target tensor, when the number of elements on the last dimension is 2 to the power of n, shift the index of the thread right by n bits to obtain the starting address of the last dimension; based on the starting address of the last dimension and the second step size, determine the index of the any element on the last dimension.

[0126] In some embodiments, the processing module 320 is specifically used to: obtain the continuous output tensor corresponding to the target tensor according to the index of each element; and process the continuous output tensor by calling the kernel function through the thread block.

[0127] In some embodiments, the target tensor includes a first tensor and a second tensor with discontinuous addresses; the processing module 320 is specifically used to: determine a memory space that meets the tensor storage requirements based on the size of the first tensor and the size of the second tensor; read the elements in the first tensor according to the index of each element in the first tensor to obtain a first output tensor; read the elements in the second tensor according to the index of each element in the second tensor to obtain a second output tensor; store the first output tensor at a first address in the memory space, and store the second output tensor at a second address adjacent to the first address to obtain a continuous output tensor.

[0128] In some embodiments, the processor performs the operation of obtaining the continuous output tensors corresponding to the target tensor according to the index of each element; or, the target tensor is copied to a graphics processor, and the graphics processor performs the operation of obtaining the continuous output tensors corresponding to the target tensor according to the index of each element.

[0129] In some embodiments, the processing module 320 is specifically used to: determine the multiple neural network operators required for processing the target tensor; based on the association relationship between the multiple neural network operators, fuse the multiple neural network operators to obtain the target neural network operator; and configure the kernel function based on the target neural network operator.

[0130] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided by the embodiment of the present disclosure can be used to perform the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present disclosure, please refer to the description of the method embodiment of the present disclosure for understanding.

[0131] It should be noted that, in the embodiments of the present disclosure, if the above-mentioned data processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present disclosure is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0132] An embodiment of the present disclosure provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0133] The present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0134] An embodiment of the present disclosure provides a computer program, including computer-readable codes. When the computer-readable codes are executed in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0135] The present disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0136] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the embodiments, and reference can be made to the similarities or similarities between them. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above-mentioned method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product disclosed herein, please refer to the description of the method embodiments disclosed herein for understanding.

[0137] It should be noted that Figure 4 A schematic diagram of a hardware entity of a computer device in an embodiment of the present disclosure is shown in FIG. Figure 4 As shown, the hardware entity of the computer device 400 includes: a processor 401, a communication interface 402 and a memory 403, wherein:

[0138] Processor 401 generally controls the overall operation of computer device 400 .

[0139] The communication interface 402 enables the computer device to communicate with other terminals or servers through a network.

[0140] Memory 403 is configured to store instructions and applications executable by processor 401 and to cache data to be processed or processed by processor 401 and various modules in computer device 400 (e.g., image data, audio data, voice communication data, and video communication data). This can be implemented using flash memory (FLASH) or random access memory (RAM). Data can be transmitted between processor 401, communication interface 402, and memory 403 via bus 404.

[0141] It should be understood that “one embodiment” or “an embodiment” mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, “in one embodiment” or “in an embodiment” appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present disclosure, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure are for description only and do not represent the advantages and disadvantages of the embodiments.

[0142] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0143] In the several embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0144] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0145] In addition, all functional units in the embodiments of the present disclosure may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0146] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0147] Alternatively, if the above-mentioned integrated unit of the present disclosure is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0148] The above is only an embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the present disclosure, and they should all be covered by the protection scope of the present disclosure.

Claims

1. A data processing method, characterized in that: The data processing method includes: Determining a kernel function and a thread block required for processing a target tensor; the target tensor includes a plurality of element regions, and threads in the thread block have a one-to-one correspondence with the element regions; determining an index of each element in the target tensor based on an index of a thread in the thread block, a number of elements in the element region in a target dimension, and a stride between adjacent elements in the element region; The kernel function is called by the thread block to process the target tensor according to the index of each element.

2. The data processing method according to claim 1, wherein: The data processing method further includes: Determining a start index of the thread block in the target dimension based on an index of the thread block in the target dimension in the computation space and the number of threads of the thread block in the target dimension; An index of each thread is determined based on the starting index of the target dimension and the index of each thread on the target dimension in the thread block.

3. The data processing method according to claim 1 or 2, characterized in that: The step size includes a first step size of the target dimension and a second step size of the remaining dimensions except the target dimension; Determining the index of each element in the target tensor based on the index of the thread in the thread block, the number of elements in the element region in the target dimension, and the stride between adjacent elements in the element region includes: Determine an index of each element in the target dimension based on the thread index, the number of elements in the element region in the target dimension, and the first step length; Determine the index of each element in the remaining dimensions based on the thread index, the number of elements in the element region in the target dimension, and the second step size; The index of each element is determined based on the index of each element in the target dimension and the index of each element in the remaining dimensions.

4. The data processing method according to claim 3, wherein: The determining, based on the index of the thread, the number of elements in the element region in the target dimension, and the second step size, the index of each element in the remaining dimensions includes: For any element on the last dimension of the target tensor, when the number of elements on the last dimension is 2 to the power of n, shift the thread index right by n bits to obtain the starting address of the last dimension; n is a positive integer; Based on the starting address of the last dimension and the second step length, the index of the any element in the last dimension is determined.

5. The data processing method according to claim 1, wherein: The calling the kernel function through the thread block and processing the target tensor according to the index of each element includes: Obtaining a continuous output tensor corresponding to the target tensor according to the index of each element; The kernel function is called by the thread block to process the continuous output tensor.

6. The data processing method according to claim 5, characterized in that: The target tensor includes a first tensor and a second tensor whose addresses are discontinuous; Obtaining the continuous output tensor corresponding to the target tensor according to the index of each element includes: Determining a memory space that meets tensor storage requirements based on a size of the first tensor and a size of the second tensor; Reading elements in the first tensor according to the index of each element in the first tensor to obtain a first output tensor; Reading the elements in the second tensor according to the index of each element in the second tensor to obtain a second output tensor; The first output tensor is stored at a first address in the memory space, and the second output tensor is stored at a second address adjacent to the first address, to obtain continuous output tensors.

7. The data processing method according to claim 6, characterized in that: The processor executes the operation of obtaining the continuous output tensor corresponding to the target tensor according to the index of each element; or The target tensor is copied to a graphics processor, and the graphics processor performs the operation of obtaining the continuous output tensors corresponding to the target tensor according to the index of each element.

8. The data processing method according to any one of claims 1 or 2, or 4 to 7, characterized in that: The data processing method further includes: Determining a plurality of neural network operators required for processing the target tensor; Based on the correlation relationship between the multiple neural network operators, the multiple neural network operators are fused to obtain a target neural network operator; Based on the target neural network operator, the kernel function is configured.

9. A data processing device, characterized in that: The data processing device includes: a determination module, configured to determine a kernel function and a thread block required for processing a target tensor; the target tensor comprising a plurality of element regions, wherein threads in the thread block have a one-to-one correspondence with the element regions; a processing module configured to determine an index of each element in the target tensor based on an index of a thread in the thread block, a number of elements in the element region in a target dimension, and a stride between adjacent elements in the element region; The processing module is further configured to call the kernel function through the thread block and process the target tensor according to the index of each element.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is read and executed by a computer, the steps of the method according to any one of claims 1 to 8 are implemented.