Tensor calculation method, electronic device, storage medium and product

By determining the way multiple streams start the kernel in parallel, the problems of low resource utilization and time-consuming calculation in tensor calculation are solved, and more efficient tensor calculation is achieved.

CN120066751BActive Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510562376.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In the prior art, multiple small-scale tensors cannot fill the accelerator calculation unit when performing bit-by-bit operations, resulting in low resource utilization and each set of tensors has a corresponding startup core, resulting in time-consuming calculation, which in turn leads to low tensor calculation efficiency.

Method used

By obtaining the target tensor calculation task, multiple streams are determined based on the number of tensors, the dependencies between streams are managed using the event synchronization mechanism, and the kernel corresponding to each stream is started in parallel to complete the calculation of the sub-task.

Benefits of technology

This improves resource utilization, reduces the time-consuming process of starting the kernel of tensors, and thus improves the efficiency of tensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066751B_ABST
    Figure CN120066751B_ABST
Patent Text Reader

Abstract

The present application discloses a tensor calculation method, an electronic device, a storage medium, and a product, relating to the technical field of data processing, including obtaining a target tensor calculation task; determining a plurality of streams based on the number of tensors in the target tensor calculation task; determining subtasks assigned to each stream based on the target tensor calculation task and the plurality of streams; using an event synchronization mechanism to manage the dependency relationships between the plurality of streams, and starting the kernels corresponding to each stream in parallel to complete the calculation of the subtasks, which solves the technical problem in the related solutions that when multiple tensors with smaller scales perform bit-by-bit operations, they cannot fully occupy the accelerator calculation units, resulting in low resource utilization, and each group of tensors has a corresponding kernel to be started, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, and achieves the technical effect of improving resource utilization, reducing the time consumption of starting kernels for tensors, and further improving the tensor calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a tensor calculation method, an electronic device, a storage medium, and a product. Background Art

[0002] In related tensor calculation solutions, when multiple tensors with small sizes perform bit-by-bit operations, they cannot fully occupy the accelerator calculation units, resulting in low resource utilization. Moreover, when each group of tensors performs bit-by-bit operations, the corresponding kernels need to be started, resulting in the accumulation of tensor calculation time consumption, and thus low tensor calculation efficiency. Summary of the Invention

[0003] This application provides a tensor calculation method, an electronic device, a storage medium, and a product, so as to at least solve the problem in related technologies that when multiple tensors with small scales perform bit-by-bit operations, they cannot fully occupy the accelerator calculation units, resulting in low resource utilization, and each group of tensors has a corresponding starting kernel, resulting in the accumulation of calculation time consumption, and thus low tensor calculation efficiency.

[0004] This application provides a tensor calculation method, including:

[0005] Obtain a target tensor calculation task;

[0006] Based on the number of tensors in the target tensor calculation task, determine multiple streams;

[0007] Based on the target tensor calculation task and the multiple streams, determine the subtasks assigned to each stream;

[0008] Use an event synchronization mechanism to manage the dependency relationships between the multiple streams, and parallelly start the kernels corresponding to each stream to complete the calculation of the subtasks.

[0009] This application also provides a tensor calculation device, including:

[0010] An obtaining unit, configured to obtain a target tensor calculation task;

[0011] A first determination unit, configured to determine multiple streams based on the number of tensors in the target tensor calculation task;

[0012] A second determination unit, configured to determine the subtasks assigned to each stream based on the target tensor calculation task and the multiple streams;

[0013] A calculation unit, configured to use an event synchronization mechanism to manage the dependency relationships between the multiple streams, and parallelly start the kernels corresponding to each stream to complete the calculation of the subtasks.

[0014] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above tensor calculation methods when executing the computer program.

[0015] The present application also provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps of any of the above tensor calculation methods.

[0016] The present application also provides a computer program product including a computer program, where the computer program, when executed by a processor, implements the steps of any of the above tensor calculation methods.

[0017] Through the present application, the present application discloses a tensor calculation method, an electronic device, a storage medium, and a product, relating to the technical field of data processing, including obtaining a target tensor calculation task; determining a plurality of streams based on the number of tensors in the target tensor calculation task; determining subtasks assigned to each stream based on the target tensor calculation task and the plurality of streams; using an event synchronization mechanism to manage the dependency relationships between the plurality of streams, and starting the kernels corresponding to each stream in parallel to complete the calculation of the subtasks, solving the technical problems in the related solutions that when performing bit-by-bit operations on a plurality of tensors with small scales, the accelerator computing units cannot be fully occupied, resulting in low resource utilization, and each group of tensors has a corresponding kernel startup, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, achieving the technical effects of improving resource utilization, reducing the kernel startup time of tensors, and further improving the tensor calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a schematic flowchart of a tensor calculation method provided by an embodiment of the present application;

[0020] Figure 2 It is a schematic diagram of a subtask allocation method provided by an embodiment of the present application;

[0021] Figure 3 It is a schematic diagram of a thread waiting method provided by an embodiment of the present application;

[0022] Figure 4 It is a schematic flowchart of a tensor calculation method provided by an embodiment of the present application;

[0023] Figure 5 It is a schematic structural diagram of a tensor calculation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.

[0025] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0026] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0027] The optimal substructure means that the optimal substructure requires the solution of the problem to be recursive, that is, the optimal solution of the original problem depends on the optimal solutions of the subproblems.

[0028] A stream refers to that on the accelerator chip, the stream can ensure that the tasks placed on this stream are executed in the order in which they are placed, and the order between different streams is not affected.

[0029] The greedy allocation algorithm refers to an algorithm strategy that makes the optimal (i.e., most favorable) choice in each step of selection, hoping to lead to a globally optimal solution. The greedy algorithm is applicable to problems that can be decomposed into multiple subproblems, and the optimal solutions of each subproblem can be combined into a globally optimal solution.

[0030] A kernel refers to a function that is called on the host but executed on a device or other accelerator. Here, the host is usually a Central Processing Unit (CPU), and the accelerator is usually a Graphics Processing Unit (GPU). These functions are designed to be able to process a large amount of data in parallel, taking advantage of the parallel computing power of the accelerator.

[0031] With the rise of artificial intelligence and the popularization of large models, various operators will emerge. In artificial intelligence deep learning frameworks such as PyTorch and TensorFlow, there is a class of batch processing operation operators. For example, an operator like foreach_add is used for batch processing of tensors. Its input is two lists of tensors (input 1 and input 2), and the output is a list of tensors with element-wise operation results. That is, the corresponding tensors in input 1 and input 2 are operated on in sequence and placed in each tensor in the result tensor list. The tensors in these lists are usually stored in different memory locations and cannot be uniformly processed using a single kernel. Moreover, it is often the case that the individual sizes of these tensors are not very large. If calculated sequentially, the kernel needs to be started sequentially, resulting in a relatively large amount of time wasted on kernel startup. And due to the small number of individual tensors, the utilization rate of the acceleration chip is not high, causing waste and relatively more time consumption.

[0032] The following briefly introduces several solutions for tensor calculation methods in related technologies:

[0033] In Scheme A, the conventional approach is to process them sequentially on a single stream.

[0034] In Scheme B, an attempt is made to process them in a single kernel, but it has low generality, too many constraints, and is only applicable in extremely rare cases.

[0035] In Scheme C, multiple small tensors are concatenated into a large tensor and then a single bitwise operation is performed. However, it is limited by the requirement of tensor shape consistency. Even if the conditions are met, there will be additional overheads brought by concatenation and splitting, which may be more time-consuming than the conventional approach.

[0036] In the above solutions, the following defects exist:

[0037] Low utilization rate of accelerator computing power: When the scale of a single tensor is small (for example, with a size of 32x32), the addition operations of each group of tensors cannot fully occupy the accelerator computing units, resulting in idle resources and low utilization rate.

[0038] Accumulation of execution latency: In the serial execution mode, the total latency is the sum of the latencies of all operations, and the multi-task parallel ability of the GPU cannot be utilized.

[0039] High kernel startup time consumption: Starting the corresponding kernel for each group of tensors, the serial startup time will accumulate.

[0040] To solve the problems existing in related solutions, an embodiment of the present application provides a tensor calculation method, including: obtaining a target tensor calculation task; determining a plurality of streams based on the number of tensors in the target tensor calculation task; determining subtasks assigned to each stream based on the target tensor calculation task and the plurality of streams; using an event synchronization mechanism to manage the dependencies between the plurality of streams, and starting the kernels corresponding to each stream in parallel to complete the calculation of the subtasks, which solves the technical problems in related solutions that when multiple tensors of relatively small scales perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding kernel startup, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, and achieves the technical effects of improving resource utilization, reducing the kernel startup time of tensors, and further improving the tensor calculation efficiency.

[0041] A tensor calculation method provided by an embodiment of the present disclosure can be executed by an accelerator, such as a GPU, a Tensor Processing Unit (TPU), etc. The tensor calculation method provided by the embodiment of the present disclosure can be applied to fields such as deep learning, scientific computing, and financial evaluation. Taking the application in the field of deep learning as an example, during the training process of a deep learning model (such as a convolutional neural network, a recurrent neural network, etc.), a large number of tensor operations are involved, such as matrix multiplication, convolutional operations, etc. This method can determine a plurality of streams according to the number of tensors, and allocate the calculation tasks of different layers or the calculation tasks of different batches of data to each stream for parallel execution, and use an event synchronization mechanism to ensure the correctness of the calculation order, thereby significantly accelerating the training process of the model. For example, when training an image classification model, the convolutional calculation tasks of different convolutional layers can be allocated to different streams for parallel processing.

[0042] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] Figure 1 It is a flowchart of a tensor calculation method provided by an embodiment of the present disclosure.

[0044] As Figure 1 shown, the method includes the following steps:

[0045] Step 101, obtaining a target tensor calculation task;

[0046] In some embodiments, the target tensor calculation task includes two or more input tensor lists and an output tensor list. The two or more input tensor lists may be stored in different memory addresses, and may have different shapes and sizes. Each tensor in the output tensor list is the element-wise operation result (such as addition, multiplication, etc.) of the tensors at the corresponding positions in the input tensor lists.

[0047] In some embodiments, the target tensor calculation task is used to perform bitwise operations on the input tensor, and the bitwise operations may include element-wise addition, element-wise subtraction, etc.

[0048] In some embodiments, the foreach_add operator in the deep learning framework is taken as an example for illustration below. Other foreach_sub operators and other operators use a similar scheme, which will not be elaborated in this application. For the two input tensor lists, tensor list A (InputA_list) and tensor list B (InputB_list), both tensor list A and tensor list B have N tensors, and the output tensor list (Output_list). The operation performed on each tensor is:

[0049] Output_list[i]=InputA_list[i]+InputB_list[i];

[0050] where i represents the i-th tensor in the tensor list, InputA_list[i] represents the i-th tensor in tensor list A, InputB_list[i] represents the i-th tensor in tensor list B, and Output_list[i] represents the i-th tensor in the output tensor list.

[0051] Step 102, determine multiple streams based on the number of tensors in the target tensor calculation task;

[0052] In some embodiments, the number of tensors in the target tensor calculation task refers to the total number of tensors participating in the calculation task in a single tensor list.

[0053] In some embodiments, when the task involves a large number of small-scale tensors, a single stream cannot fully utilize the hardware resources, while multiple streams can start multiple kernels simultaneously to improve the utilization rate of computing resources.

[0054] In some embodiments, the number of required streams can be dynamically calculated according to the characteristics of the chip itself and the situation of the tensors in the current target tensor calculation task.

[0055] In some embodiments, the number of streams is also affected by hardware limitations and resource contention. Among them, hardware limitations include the maximum number of streams, computing units, and memory bandwidth, and resource contention includes memory resource contention and computing resource contention. Specifically, different hardware platforms (such as GPUs, TPUs) support a limited maximum number of concurrent streams. If the number of tensors is large, the number of streams can be limited according to the maximum number of streams supported by the hardware to avoid excessive management costs. Each stream occupies a certain amount of computing resources and video memory bandwidth. If too many streams are created, it may cause the computing unit or memory bandwidth to become a bottleneck, thereby reducing performance. When multiple streams are running simultaneously, if they all require a large amount of video memory or other types of memory, it may lead to memory resource contention, which will not only slow down the data transfer speed but also may cause out-of-memory errors. When multiple streams need to use the same computing resources simultaneously, resource contention will occur, affecting the overall efficiency.

[0056] In some embodiments, by determining multiple streams based on the number of tensors in the target tensor computing task, the parallel computing capabilities of the hardware resources can be fully utilized.

[0057] Step 103: Based on the target tensor computing task and multiple streams, determine the subtasks assigned to each stream;

[0058] In some embodiments, based on the target tensor computing task and multiple streams, the target tensor computing task can be evenly distributed to each stream so that the computing tasks on multiple streams are almost completed simultaneously, thereby improving the efficiency of tensor computing.

[0059] In some embodiments, by determining the subtasks assigned to each stream based on the target tensor computing task and multiple streams, the problems of low resource utilization rate and cumulative kernel startup time caused by processing the target tensor computing task on a single stream can be avoided, effectively improving the utilization rate of hardware resources, reducing resource contention, and ensuring the efficiency and stability of task execution.

[0060] Step 104: Use the event synchronization mechanism to manage the dependencies between multiple streams, and parallelly start the kernels corresponding to each stream to complete the calculation of the subtasks.

[0061] In some embodiments, the event synchronization mechanism is a technical means for managing the execution order of concurrent tasks. Usually, an event is used to mark a specific operation point or state. For example, in GPU programming, an event is a lightweight object that can be used to record the completion status of operations in a certain stream and allow other streams to wait for the occurrence of this event, ensuring that the tasks between different streams are executed in the correct order and avoiding data contention or resource conflicts.

[0062] In some embodiments, the dependency relationship between multiple streams refers to a certain order or constraint condition existing between tasks in different streams. For example, the task in stream A must start after the task in stream B is completed, or the task in stream C needs to wait until both stream A and stream B are completed before it can be executed.

[0063] In some embodiments, starting the kernels corresponding to each stream in parallel means starting different kernel tasks simultaneously on multiple streams to make full use of hardware resources.

[0064] In some embodiments, starting the kernel corresponding to each stream is usually to call an API on the host to start the kernel, which is executed by the device.

[0065] In some embodiments, by using an event synchronization mechanism to manage the dependency relationship between multiple streams and starting the kernels corresponding to each stream in parallel to complete the calculation of subtasks, the utilization rate of computing resources can be improved, the computing latency can be reduced, the resource utilization can be optimized, and the correctness and execution order of tasks can be ensured.

[0066] Through this application, a target tensor calculation task is obtained; based on the number of tensors in the target tensor calculation task, multiple streams are determined; based on the target tensor calculation task and multiple streams, the subtasks assigned to each stream are determined; an event synchronization mechanism is used to manage the dependency relationship between multiple streams, and the kernels corresponding to each stream are started in parallel to complete the calculation of subtasks, solving the technical problem in the related solutions that when multiple tensors of relatively small scales perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding started kernel, resulting in the accumulation of computing time consumption and thus low tensor calculation efficiency, achieving the technical effect of improving resource utilization, reducing the time consumption of starting the kernel for tensors, and thus improving the tensor calculation efficiency.

[0067] In some embodiments, determining multiple streams based on the number of tensors in the target tensor calculation task includes:

[0068] Based on the tensors in the target tensor calculation task, determine the number of tensors smaller than a preset size, where the size is the number of elements in the tensor;

[0069] In some embodiments, the preset size can be determined by the smallest size that can theoretically fill the GPU. Specifically, the mathematical expression for determining the preset size is:

[0070] M = Size * 25%;

[0071] where M represents the preset size, Size represents the smallest size that can theoretically fill the GPU, and 25% is a preset value. Generally, it is considered that when the utilization rate reaches more than 50%, the computing resources of the GPU can be considered to be utilized relatively fully, and the grouping is at least two groups. Therefore, the preset value is determined to be 25%, and it can also be determined according to needs. This application does not limit this.

[0072] Further, the minimum size that can theoretically fill the GPU can be determined by the number of streaming multiprocessors and the recommended typical value of the thread block. Specifically, the mathematical expression for determining the minimum size that can theoretically fill the GPU is:

[0073] Size = S * B;

[0074] Wherein, S represents the number of multiprocessors, B represents the recommended typical value of the thread block, and Size represents the minimum size that can theoretically fill the GPU.

[0075] In some embodiments, taking S as 108 and B as 512 as an example, M = 108 * 512 * 25% = 13824, that is, the preset size is 13824.

[0076] In some embodiments, a two-dimensional tensor with a shape of (3, 4) contains 12 elements, and its size is 12; a three-dimensional tensor with a shape of (2, 2, 2) contains 8 elements, and its size is 8.

[0077] In some embodiments, by calculating the tensors in the target tensor calculation task to determine the number of tensors smaller than the preset size, small-scale tasks that may not be able to fully utilize the computing power of the hardware accelerator can be identified.

[0078] Based on the number of tensors smaller than the preset size and the number of preset tensor groups, determine the target grouping number.

[0079] Based on the target grouping number, determine multiple streams.

[0080] In some embodiments, the number of preset tensor groups is usually 2, that is, when the data volume of more than 2 groups of tensors is less than the preset size theoretically, grouping can be considered. However, considering the unevenness of grouping and the additional overhead brought by grouping, the comprehensive benefit may not be achieved when the number is small. In this application, the specific value of the number of preset tensor groups is not limited. Taking the number of preset tensor groups in this application as 8 as an example, that is, when the data of more than 8 groups of tensors is less than the preset size, grouping starts.

[0081] In some embodiments, controlling the maximum number of groups in this application to be 8 can meet most scenarios. Of course, the maximum number of groups can also be specifically determined according to specific hardware and tests.

[0082] In some embodiments, determine the target grouping number as the number of streams, that is, finally use the target grouping number of streams to process the target tensor calculation task.

[0083] In some embodiments, determining the number of tensors smaller than the preset size based on the tensors in the target tensor calculation task includes:

[0084] Determine the dimensions of each group of tensors based on the tensors in the target tensor calculation task;

[0085] In some embodiments, the target tensor calculation task includes at least two input tensor lists.

[0086] In some embodiments, each group of tensors is the tensors for bitwise operations in the input tensor list.

[0087] Determine the number of tensors smaller than the preset dimension based on the dimensions of each group of tensors and the preset dimension.

[0088] In some embodiments, compare the dimensions of each group of tensors with the preset dimension, filter out the tensors with dimensions smaller than the preset dimension, and determine the number of the filtered tensors.

[0089] In some embodiments, determining the target grouping number based on the number of tensors smaller than the preset dimension and the number of preset tensor groups includes:

[0090] In response to the number of tensors smaller than the preset dimension being lower than the number of preset tensor groups, determine the first preset value as the target grouping number.

[0091] In some embodiments, the number of tensors smaller than the preset dimension being lower than the number of preset tensor groups means that the number of small tensors in the target tensor calculation task is not very large (does not meet the grouping requirements). Therefore, directly using one stream to calculate the target tensor calculation task can make better use of the computing resources. Therefore, the first preset value is 1.

[0092] In some embodiments, determining the target grouping number based on the number of tensors smaller than the preset dimension and the number of preset tensor groups includes:

[0093] In response to the number of tensors smaller than the preset dimension not being lower than the number of preset tensor groups, compare the number of tensors smaller than the preset dimension with the number of preset tensor groups to obtain a first ratio;

[0094] In some embodiments, the number of preset tensor groups represents the expected number of small tensor groups, which is used to measure whether the number of small tensors in the current task is large enough to determine whether multiple streams need to be used for parallel calculation.

[0095] In some embodiments, the number of tensors smaller than the preset dimension not being lower than the number of preset tensor groups indicates that there are more small-scale tensors in the target tensor calculation task. More streams can be created to increase the parallelism and make full use of the hardware resources.

[0096] In some embodiments, the first ratio is obtained by dividing the number of tensors smaller than the preset dimension by the number of preset tensor groups.

[0097] Determine the target number of groups based on the first ratio.

[0098] In some embodiments, use the first ratio as the target number of groups, that is, the target number of groups can be dynamically determined based on the first ratio. Specifically, during the task execution, the target number of groups can be dynamically adjusted according to the real-time monitored first ratio. For example, if the first ratio increases, the number of groups can be dynamically increased to handle more small tensor tasks; if the first ratio decreases, the number of groups can be reduced to save resources.

[0099] In some embodiments, based on the target tensor calculation task and multiple streams, determining the subtasks assigned to each stream includes:

[0100] Based on the target tensor calculation task, obtain the number of elements in each tensor and the continuity of each tensor;

[0101] In some embodiments, the continuity is used to indicate the input / output continuity of the elements in the tensor.

[0102] In some embodiments, the number of elements in each tensor refers to the total number of all elements in the tensor. In deep learning frameworks (such as PyTorch or TensorFlow), the built-in functions can be directly used to obtain the number of elements in the tensor.

[0103] In some embodiments, the input / output continuity of the elements in the tensor means that all elements of the tensor are continuously arranged in memory, then the tensor is called continuous. In this case, the efficiency of accessing the data of the tensor is higher. If the tensor undergoes some operations (such as slicing, transposing, etc.), its elements may become discontinuous in memory. For discontinuous tensors, when performing bit-by-bit calculations, since the position information needs to be calculated additionally, the time consumption is longer than that of continuous tensors, which will increase the time overhead of data access.

[0104] Based on the number of elements in each tensor and the continuity of each tensor, determine the heaviness value of each group of tensors in the target tensor calculation task. The size of the heaviness value is used to indicate the computational heaviness of the task;

[0105] In some embodiments, the more elements in the tensor, the higher the computational complexity, and the greater the corresponding heaviness value of the tensor.

[0106] In some embodiments, different weights can be determined for the heaviness value of the tensor according to the specific type of computational operation (such as addition, multiplication, convolution, etc.).

[0107] Based on the heaviness value of each group of tensors in the target tensor calculation task and multiple streams, determine the subtasks assigned to each stream.

[0108] In some embodiments, based on the heaviness value of each group of tensors and multiple streams in the target tensor calculation task, the target tensor calculation task can be allocated to each stream as evenly as possible according to the magnitude of the heaviness value. Here, each stream corresponds to one or more groups of tensors, and each group of tensors corresponds to a stream.

[0109] In some embodiments, by determining the subtasks allocated to each stream based on the heaviness value of each group of tensors and multiple streams in the target tensor calculation task, load balancing can be ensured, resource contention can be reduced, and the utilization rate of hardware resources can be maximized.

[0110] In some embodiments, determining the heaviness value of each group of tensors in the target tensor calculation task based on the number of elements in each tensor and the continuity of each tensor includes:

[0111] Based on the continuity of each tensor, determining the number of discontinuous tensors in the target tensor calculation task;

[0112] In some embodiments, in PyTorch, the is_contiguous function can be used to determine whether a tensor is continuous. For each tensor, if the is_contiguous function returns false, it is counted as a discontinuous tensor, and the number of all discontinuous tensors in the target tensor calculation task is statistically calculated.

[0113] Based on the number of elements in each tensor and the number of discontinuous tensors, determining the heaviness value of each group of tensors in the target tensor calculation task.

[0114] In some embodiments, discontinuous tensors will increase the memory access overhead, so an additional penalty factor is required. The mathematical expression for determining the heaviness value of a tensor is as follows:

[0115] Weight[i]=InputA_list[i].numel()*(1+Num_discontinuous);

[0116] Where, Weight[i] represents the heaviness value, InputA_list[i].numel() represents the number of elements of the i-th group of tensors participating in the calculation. Since the number of elements of the three tensors in this group of bitwise operations is the same size, any one of the tensors in this group can be taken. Num_discontinuous represents the number of discontinuous tensors. For example, in the foregoing embodiments, there are 2 input tensors, 1 output tensor, input tensor list A (InputA_list) and tensor list B (InputB_list), output tensor list (Output_list). If 2 of these 3 tensors are discontinuous and 1 tensor is continuous, then the value of Num_discontinuous is 2.

[0117] In some embodiments, determining the subtasks assigned to each stream based on the heaviness values of each group of tensors and multiple streams in a target tensor calculation task includes:

[0118] Sorting the heaviness values of each group of tensors in the target tensor calculation task in descending order of the heaviness values to obtain the sorted heaviness values;

[0119] In some embodiments, the heaviness value reflects the computational complexity and memory access overhead of each group of tensors. The larger the heaviness value, the more burdensome the calculation task of that group of tensors.

[0120] In some embodiments, by sorting the heaviness values in descending order, tasks with higher computational complexity can be processed preferentially, thereby optimizing task scheduling and resource allocation.

[0121] In some embodiments, by sorting in descending order of the heaviness values to obtain the sorted heaviness values, tasks with higher computational complexity can be effectively identified, and based on this, task scheduling and resource allocation can be optimized.

[0122] Determining the subtasks assigned to each stream based on the sorted heaviness values and multiple streams.

[0123] In some embodiments, after the assignment based on the sorted heaviness values and multiple streams is completed, the total amount of tasks assigned to each stream is roughly equal.

[0124] In some embodiments, by based on the sorted heaviness values and multiple streams, tasks with higher heaviness values are preferentially assigned to the stream with the lightest current load, avoiding overloading of some streams while other streams are idle.

[0125] In some embodiments, by determining the subtasks assigned to each stream based on the sorted heaviness values and multiple streams, the efficiency of parallel computing and resource utilization can be improved.

[0126] In some embodiments, during the task execution process, if it is found that some streams are completed faster, the remaining tasks can be dynamically reassigned. For example, an event synchronization mechanism can be used to monitor the completion status of each stream and adjust the task assignment in real time.

[0127] In some embodiments, if some tasks have higher priorities, higher weights can be given to these tasks during assignment, and they are preferentially assigned to the streams with more idle resources.

[0128] In some embodiments, determining the subtasks assigned to each stream based on the sorted heaviness values and multiple streams includes:

[0129] Based on the sorted heavy values and multiple streams, the target tensor calculation task is allocated to multiple streams using a preset algorithm to obtain subtasks allocated to each stream. The preset algorithm is used to allocate the heavy value to the stream with the lowest load each time.

[0130] In some embodiments, the preset algorithm in this application takes the greedy algorithm as an example. For example, Figure 2 as shown, Figure 2 is a schematic diagram of a subtask allocation method provided by an embodiment of this application. Specifically, Figure 2 taking 6 groups of tensor data in it and using 3 streams (stream1, stream2, and stream3 from left to right) for calculation as an example, Figure 2 the numbers in the gray grids in it are the calculated computational weights (heaviness). Steps 1 to 2 are the sorting process, and steps 3 to 8 are the task allocation process. After completing the task division through the greedy algorithm, it can be seen from Figure 2 the last step that the task amounts on each stream are as follows: stream1: 400, stream2: 400, stream3: 380. It can be seen that the task amounts allocated to each stream after division are roughly equivalent, and the task allocation has an optimal substructure. The solution obtained using the greedy algorithm is the global optimal solution.

[0131] In some embodiments, by allocating the target tensor calculation task to multiple streams based on the sorted heavy values and multiple streams using a preset algorithm to obtain subtasks allocated to each stream, it can ensure that the task loads of each stream are as balanced as possible, thereby improving the efficiency of parallel computing and resource utilization.

[0132] In some embodiments, after determining the subtasks allocated to each stream based on the target tensor calculation task and multiple streams, the tensor calculation method further includes:

[0133] Create a main thread and multiple sub-threads;

[0134] In some embodiments, the main thread is used for task initialization, scheduling, and coordination. Each sub-thread is bound to a stream and is responsible for executing the tasks in that stream.

[0135] In some embodiments, the main thread is responsible for creating events and allocating the events to each stream. When the sub-threads execute tasks, they coordinate the task execution order of different streams through the events.

[0136] Based on multiple streams, determine the streams corresponding to the main thread and multiple sub-threads;

[0137] In some embodiments, each stream maintains a task queue.

[0138] Put the subtasks into the task pools of the corresponding streams and notify the threads of the corresponding streams.

[0139] In some embodiments, the task pool of a stream is used to store the subtasks allocated to that stream.

[0140] In some embodiments, after the main thread puts a task into the task pool of a stream, it notifies the corresponding child thread to start execution.

[0141] In some embodiments, when a thread receives information that a task has arrived, it retrieves the task from the task pool of the corresponding stream.

[0142] In some embodiments, using an event synchronization mechanism to manage the dependencies between multiple streams and starting the kernels corresponding to each stream in parallel to complete the calculation of subtasks includes:

[0143] Create a first event for the stream corresponding to the child thread, where the first event is used to synchronize the subtasks between multiple streams;

[0144] In some embodiments, the first event is a synchronization point used to mark the completion status of tasks in a certain stream. When a stream completes some tasks, this event is recorded, and other streams can wait for this event to complete before continuing to execute subsequent tasks. The first event can ensure that the calculation on this stream starts only after the calculation on the main stream is completed.

[0145] In some embodiments, the first event (event) can be reused later to avoid the additional overhead caused by creating it each time.

[0146] In some embodiments, by creating the first event for the stream corresponding to the child thread and using this event to synchronize the subtasks between multiple streams, the dependency relationship between tasks and the resource contention problem can be effectively solved. The reuse mechanism of the event further improves the efficiency and reduces the resource overhead.

[0147] Based on the first event, determine a second event on the stream corresponding to the main thread. The second event includes a first recording event and a first waiting event. The first recording event is used to record the key points of the stream corresponding to the main thread, and the first waiting event is used to wait for the subtasks on the stream corresponding to the main thread to complete;

[0148] In some embodiments, the first recording event can be implemented through the recordevent function of the accelerator.

[0149] In some embodiments, the first waiting event can be implemented through the wait event function of the accelerator.

[0150] In response to the completion of the subtasks on the stream corresponding to the main thread, start the kernels on the streams corresponding to multiple child threads in parallel to complete the calculation of the subtasks.

[0151] In some embodiments, in parallel computing, after the subtask on the stream corresponding to the main thread is completed, the kernel startup on the streams corresponding to multiple sub-threads can be triggered to complete the remaining subtask calculations. This mechanism can make full use of hardware resources, thereby improving computing efficiency.

[0152] In some embodiments, for each stream thread, to ensure that the calculation on this stream starts after the main stream calculation is completed, an event is created, and the recordevent function of the accelerator is called, and then the wait event function of the accelerator is called, and then the main thread is notified. This operation does not block the thread, but keeps a record to ensure that the kernel started on this stream later will not be started before the previous task of the main thread is completed, thereby using dirty data or causing a program crash, and then the kernel assigned to this stream is started in sequence to perform calculations.

[0153] In some embodiments, managing the dependencies between multiple streams using an event synchronization mechanism, and starting the kernel corresponding to each stream in parallel to complete the calculation of the subtask includes:

[0154] In response to receiving notification messages of the streams corresponding to the plurality of child threads, starting a kernel on the stream corresponding to the main thread;

[0155] In some embodiments, the notification message of the flow corresponding to the multiple sub-threads means that each sub-thread generates a notification after its corresponding task is completed, indicating that the task has been successfully completed.

[0156] In some embodiments, in addition to a simple task completion notification, the notification message may also include information about the task execution status, such as whether the task is successful, any errors or exceptions encountered, etc.

[0157] In some embodiments, in parallel computing, when the streams corresponding to multiple sub-threads complete their tasks, notification messages can be sent to the main thread. After receiving these notification messages, the main thread starts the kernel on its corresponding stream to complete subsequent computing tasks, which can achieve dependency management between tasks and fully utilize the parallel capabilities of multi-threading and multi-stream.

[0158] Obtaining a third event on the stream corresponding to the plurality of sub-threads, the third event comprising a second recording event and a second waiting event, the second recording event being used to record a key point of the stream corresponding to the sub-thread, and the second waiting event being used to wait for completion of a sub-task on the stream corresponding to the sub-thread;

[0159] In some embodiments, the second recording event and the second waiting event may be implemented in the same manner as the aforementioned first recording event and the first waiting event, which will not be described in detail herein.

[0160] In some embodiments, a third event is generated by the streams corresponding to multiple child threads and is used to coordinate the task execution order among the child threads.

[0161] In response to the completion of subtasks on the streams corresponding to multiple child threads, a first task on the stream corresponding to the main thread is executed to complete the calculation of the subtasks.

[0162] In some embodiments, the first task refers to other tasks in the main thread.

[0163] In some embodiments, for the main thread, after waiting for all other stream notification messages, the tasks of the kernels to be executed on the main thread are started in sequence. After the task startup is completed, the recordevent functions of other streams and the functions waiting for these events are called in sequence to ensure that before the kernels on the main stream are executed later, the kernels on stream1 and stream2 have been completed. In this way, other tasks on the main thread stream can be continued later, and the calculated results can be obtained. This operation does not block the main thread either, but only ensures the execution order on the stream.

[0164] In some embodiments, as Figure 3 shown, Figure 3 FIG. is a schematic diagram of a thread waiting method provided by an embodiment of the present application. Specifically, after calling to record the event on the main stream, the kernel 1 that has been initiated on the main stream at this moment will be recorded. When calling to wait for this event, it is ensured that the kernel 3 on stream 1 will continue to execute only after the kernel 1 on the main stream has been executed. Similarly, the kernel on the main stream will continue to execute only after the kernel 3 on stream 1 has been executed. In this way, the execution order is ensured, and the calling thread is not blocked.

[0165] In some embodiments, as Figure 4 shown, Figure 4 FIG. is a schematic flow diagram of a tensor calculation method provided by an embodiment of the present application. The tasks on each stream are calculated and allocated, and the allocated tasks are respectively placed in the main stream (main stream) task pool, stream 1 task pool, and stream 2 task pool. For thread 1 and thread 2, the stream events of the main thread are respectively recorded, and then the functions waiting for the main thread events are respectively called to notify the main thread and start the kernels of the child threads; for the main thread, after receiving the notification messages from thread 1 and thread 2, the kernel of the main thread is started, other stream events are recorded, and other tasks on the main thread are executed after waiting for the stream events of other threads.

[0166] Through this application, a target tensor calculation task is obtained; based on the number of tensors in the target tensor calculation task, multiple streams are determined; based on the target tensor calculation task and the multiple streams, subtasks assigned to each stream are determined; an event synchronization mechanism is used to manage the dependencies between the multiple streams, and the kernels corresponding to each stream are started in parallel to complete the calculation of the subtasks, solving the technical problems in related solutions that when multiple tensors of a small scale perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding started kernel, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, achieving the technical effects of improving resource utilization, reducing the time consumption of starting kernels for tensors, and further improving tensor calculation efficiency.

[0167] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0168] An embodiment of this application also provides a tensor calculation device 500. Figure 5 It is a schematic structural diagram of a tensor calculation device provided by an embodiment of the present disclosure, as Figure 5 shown, including:

[0169] An acquisition unit 501, configured to acquire a target tensor calculation task;

[0170] A first determination unit 502, configured to determine multiple streams based on the number of tensors in the target tensor calculation task;

[0171] A second determination unit 503, configured to determine subtasks assigned to each stream based on the target tensor calculation task and the multiple streams;

[0172] A calculation unit 504, configured to use an event synchronization mechanism to manage the dependencies between the multiple streams, and start the kernels corresponding to each stream in parallel to complete the calculation of the subtasks.

[0173] Through this application, a target tensor calculation task is obtained; based on the number of tensors in the target tensor calculation task, multiple streams are determined; based on the target tensor calculation task and the multiple streams, subtasks assigned to each stream are determined; an event synchronization mechanism is used to manage the dependencies between the multiple streams, and the kernels corresponding to each stream are started in parallel to complete the calculation of the subtasks, solving the technical problems in related solutions that when multiple tensors of a small scale perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding started kernel, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, achieving the technical effects of improving resource utilization, reducing the time consumption of starting kernels for tensors, and further improving tensor calculation efficiency.

[0174] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determination unit 502 is configured to:

[0175] Based on the tensors in the target tensor calculation task, determine the number of tensors smaller than a preset size, where the size is the number of elements in the tensor;

[0176] Based on the number of tensors smaller than the preset size and the number of preset tensor groups, determine the target grouping number;

[0177] Based on the target grouping number, determine multiple streams.

[0178] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determination unit 502 is configured to:

[0179] Based on the tensors in the target tensor calculation task, determine the size of each group of tensors. The target tensor calculation task includes at least two input tensor lists, and each group of tensors is the tensors that perform bitwise operations in the input tensor list;

[0180] Based on the size of each group of tensors and the preset size, determine the number of tensors smaller than the preset size.

[0181] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determination unit 502 is configured to:

[0182] In response to the number of tensors smaller than the preset size being lower than the number of preset tensor groups, determine the first preset value as the target grouping number.

[0183] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determination unit 502 is configured to:

[0184] In response to the number of tensors smaller than the preset size not being lower than the number of preset tensor groups, compare the number of tensors smaller than the preset size with the number of preset tensor groups to obtain a first ratio;

[0185] Based on the first ratio, determine the target grouping number.

[0186] Further, in a possible implementation manner of the embodiments of the present disclosure, the second determination unit 503 is configured to:

[0187] Based on the target tensor calculation task, obtain the number of elements in each tensor and the continuity situation of each tensor. The continuity situation is used to indicate the input and output continuity situation of the elements in the tensor;

[0188] Based on the number of elements in each tensor and the continuity situation of each tensor, determine the heaviness value of each group of tensors in the target tensor calculation task. The size of the heaviness value is used to indicate the computational heaviness degree of the task;

[0189] Determine the subtasks assigned to each stream based on the heavy values and multiple streams of each group of tensors in the target tensor calculation task.

[0190] Further, in a possible implementation manner of the embodiment of the present disclosure, the second determination unit 503 is configured to:

[0191] Determine the number of discontinuous tensors in the target tensor calculation task based on the continuity of each tensor;

[0192] Determine the heavy value of each group of tensors in the target tensor calculation task based on the number of elements in each tensor and the number of discontinuous tensors.

[0193] Further, in a possible implementation manner of the embodiment of the present disclosure, the second determination unit 503 is configured to:

[0194] Sort the heavy values of each group of tensors in the target tensor calculation task in descending order of heavy value to obtain the sorted heavy values;

[0195] Determine the subtasks assigned to each stream based on the sorted heavy values and multiple streams.

[0196] Further, in a possible implementation manner of the embodiment of the present disclosure, the second determination unit 503 is configured to:

[0197] Allocate the target tensor calculation task to multiple streams based on the sorted heavy values and multiple streams by using a preset algorithm, and obtain the subtasks assigned to each stream. The preset algorithm is used to allocate the heavy value to the stream with the lowest load each time.

[0198] Further, in a possible implementation manner of the embodiment of the present disclosure, the tensor calculation device 500 further includes a notification unit, and the notification unit is configured to:

[0199] Create a main thread and multiple sub-threads;

[0200] Determine the streams corresponding to the main thread and multiple sub-threads based on multiple streams;

[0201] Put the subtasks into the task pools of the corresponding streams and notify the threads of the corresponding streams.

[0202] Further, in a possible implementation manner of the embodiment of the present disclosure, the calculation unit 504 is configured to:

[0203] Create a first event for the stream corresponding to the sub-thread, and the first event is used to synchronize the subtasks between multiple streams;

[0204] Based on the first event, determine a second event on the stream corresponding to the main thread. The second event includes a first recording event and a first waiting event. The first recording event is used to record key points of the stream corresponding to the main thread, and the first waiting event is used to wait for a subtask on the stream corresponding to the main thread to complete;

[0205] In response to the completion of a subtask on the stream corresponding to the main thread, parallelly start kernels on the streams corresponding to multiple child threads to complete the calculation of the subtask.

[0206] Further, in a possible implementation manner of the embodiments of the present disclosure, the computing unit 504 is configured to:

[0207] In response to receiving a notification message of the streams corresponding to multiple child threads, start a kernel on the stream corresponding to the main thread;

[0208] Obtain a third event on the streams corresponding to multiple child threads. The third event includes a second recording event and a second waiting event. The second recording event is used to record key points of the stream corresponding to the child thread, and the second waiting event is used to wait for a subtask on the stream corresponding to the child thread to complete;

[0209] In response to the completion of a subtask on the streams corresponding to multiple child threads, execute a first task on the stream corresponding to the main thread to complete the calculation of the subtask.

[0210] For the description of the features in the corresponding embodiments of the tensor calculation device, reference can be made to the relevant description in the corresponding embodiments of the tensor calculation method, which will not be elaborated here one by one.

[0211] The embodiments of the present application further provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned embodiments of the tensor calculation method.

[0212] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-mentioned embodiments of the tensor calculation method when running.

[0213] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks and other various media that can store computer programs.

[0214] The embodiments of the present application further provide a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above-mentioned embodiments of the tensor calculation method are implemented.

[0215] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described tensor calculation method embodiments are implemented.

[0216] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0217] The above has introduced in detail a tensor calculation method, an electronic device, a storage medium, and a product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A tensor calculation method, characterized in that, Including: Obtaining a target tensor calculation task; Based on the tensors in the target tensor calculation task, determining the number of tensors smaller than a preset size, where the size is the number of elements in the tensor; Based on the number of tensors smaller than the preset size and the number of preset tensor groups, determining a target grouping number; Based on the target grouping number, determining multiple streams; Based on the target tensor calculation task, obtaining the number of elements in each tensor and the continuity situation of each tensor, where the continuity situation is used to indicate the input / output continuity situation of the elements in the tensor; Based on the continuity situation of each tensor, determining the number of discontinuous tensors in the target tensor calculation task; Based on the number of elements in each tensor and the number of discontinuous tensors, determining the workload value of each group of tensors in the target tensor calculation task, where the magnitude of the workload value is used to indicate the computational workload of the task; Based on the workload value of each group of tensors in the target tensor calculation task and the multiple streams, determining the subtasks assigned to each stream; Using an event synchronization mechanism to manage the dependency relationships between the multiple streams, and starting the kernels corresponding to each stream in parallel to complete the calculation of the subtasks.

2. The tensor calculation method according to claim 1, wherein The determining the number of tensors smaller than a preset size based on the tensors in the target tensor calculation task includes: Based on the tensors in the target tensor calculation task, determining the size of each group of tensors, where the target tensor calculation task includes at least two input tensor lists, and each group of tensors is the tensors for bitwise operations in the input tensor lists; Based on the size of each group of tensors and the preset size, determining the number of tensors smaller than the preset size.

3. The tensor calculation method according to claim 1, characterized in that The determining the target grouping number based on the number of tensors smaller than the preset size and the number of preset tensor groups includes: In response to the number of tensors smaller than the preset size being lower than the number of preset tensor groups, determining the first preset value as the target grouping number.

4. The tensor calculation method according to claim 1, wherein The determining the target grouping number based on the number of tensors smaller than the preset size and the number of preset tensor groups includes: In response to the number of tensors smaller than the preset size not being lower than the number of preset tensor groups, comparing the number of tensors smaller than the preset size with the number of preset tensor groups to obtain a first ratio; Based on the first ratio, determining the target grouping number.

5. The tensor calculation method according to claim 1, characterized in that, The determining the subtasks assigned to each stream based on the workload value of each group of tensors in the target tensor calculation task and the multiple streams includes: Based on the workload value of each group of tensors in the target tensor calculation task, sorting in descending order according to the workload value to obtain the sorted workload value; Based on the sorted workload value and the multiple streams, determining the subtasks assigned to each stream.

6. The tensor calculation method according to claim 5, characterized in that, The determining the subtasks assigned to each stream based on the sorted workload value and the multiple streams includes: Based on the sorted workload value and the multiple streams, using a preset algorithm to allocate the target tensor calculation task to the multiple streams to obtain the subtasks assigned to each stream, where the preset algorithm is used to allocate the workload value to the stream with the lowest load each time.

7. The tensor calculation method according to claim 1, wherein After determining the heavy value of each group of tensors and the multiple streams in the target tensor calculation task, and determining the subtasks assigned to each stream, the method further includes: Create a main thread and multiple sub-threads; Based on the multiple streams, determine the streams corresponding to the main thread and the multiple sub-threads; Put the subtasks into the task pools of the corresponding streams, and notify the threads of the corresponding streams.

8. The tensor calculation method according to claim 7, characterized in that, The managing the dependency relationships between the multiple streams by using an event synchronization mechanism and parallelly starting the kernels corresponding to each stream to complete the calculation of the subtasks includes: Create a first event for the stream corresponding to the sub-thread, where the first event is used to synchronize the subtasks between the multiple streams; Based on the first event, determine a second event on the stream corresponding to the main thread, where the second event includes a first recording event and a first waiting event, the first recording event is used to record the key points of the stream corresponding to the main thread, and the first waiting event is used to wait for the completion of the subtasks on the stream corresponding to the main thread; In response to the completion of the subtasks on the stream corresponding to the main thread, parallelly start the kernels on the streams corresponding to the multiple sub-threads to complete the calculation of the subtasks.

9. The tensor calculation method according to claim 8, wherein The managing the dependency relationships between the multiple streams by using an event synchronization mechanism and parallelly starting the kernels corresponding to each stream to complete the calculation of the subtasks includes: In response to receiving the notification messages of the streams corresponding to the multiple sub-threads, start the kernel on the stream corresponding to the main thread; Obtain a third event on the streams corresponding to the multiple sub-threads, where the third event includes a second recording event and a second waiting event, the second recording event is used to record the key points of the stream corresponding to the sub-thread, and the second waiting event is used to wait for the completion of the subtasks on the stream corresponding to the sub-thread; In response to the completion of the subtasks on the streams corresponding to the multiple sub-threads, execute the first task on the stream corresponding to the main thread to complete the calculation of the subtasks.

10. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the tensor calculation method according to any one of claims 1 to 9 when executing the computer program.

11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the tensor calculation method according to any one of claims 1 to 9 when executed by a processor.

12. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the tensor calculation method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • End-side cloud collaborative distributed computing method, device, equipment and medium

    CN117931447A

  • Computing power engine construction method and device, equipment and storage medium

    CN119358617A