Tensor calculation method, electronic equipment, storage medium and product

By determining the parallel processing of multiple streams and event synchronization mechanism management in tensor calculation, the problem of low computational efficiency of small tensors is solved, and more efficient resource utilization and calculation speed is achieved.

CN120066751AActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510562376.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In the prior art, multiple small-scale tensors cannot fill the accelerator calculation unit when performing bit-by-bit operations, resulting in low resource utilization and each set of tensors has a corresponding startup core, resulting in time-consuming calculation, which in turn leads to low tensor calculation efficiency.

Method used

By obtaining the target tensor calculation task, multiple streams are determined based on the number of tensors, subtasks are allocated to each stream, and the dependencies between multiple streams are managed using the event synchronization mechanism, and the kernel corresponding to each stream is started in parallel to complete the calculation of the subtask.

Benefits of technology

This improves resource utilization, reduces the time-consuming process of starting the kernel of tensors, and thus improves the efficiency of tensor calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066751A_ABST
    Figure CN120066751A_ABST
Patent Text Reader

Abstract

The invention discloses a tensor calculation method, electronic equipment, a storage medium and a product, and relates to the technical field of data processing, and the method comprises the steps: obtaining a target tensor calculation task; determining a plurality of streams based on the number of tensors in the target tensor calculation task; on the basis of the target tensor calculation task and the multiple streams, sub-tasks distributed to each stream are determined; an event synchronization mechanism is utilized to manage the dependency relationship among a plurality of streams, a kernel corresponding to each stream is started in parallel to complete calculation of subtasks, and the problems that in a related scheme, when a plurality of tensors with small scales are subjected to bit-by-bit operation, an accelerator calculation unit cannot be fully occupied, the resource utilization rate is low, each group of tensors have corresponding starting kernels, and the calculation efficiency is low are solved. The technical problem that the tensor calculation efficiency is low due to the fact that calculation time consumption is accumulated is solved, and the technical effects of improving the resource utilization rate, reducing the time consumption for starting the kernel by the tensor and further improving the tensor calculation efficiency are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly relates to a tensor calculation method, an electronic device, a storage medium, and a product. Background Art

[0002] In related tensor calculation solutions, when multiple tensors with small sizes perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization. Moreover, when each group of tensors performs bit-by-bit operations, the corresponding kernels need to be started, resulting in the accumulation of tensor calculation time consumption, and thus low tensor calculation efficiency. Summary of the Invention

[0003] This application provides a tensor calculation method, an electronic device, a storage medium, and a product, so as to at least solve the problem in related technologies that when multiple tensors with small scales perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding starting kernel, resulting in the accumulation of calculation time consumption, and thus low tensor calculation efficiency.

[0004] This application provides a tensor calculation method, including: Obtaining a target tensor calculation task; Determining multiple streams based on the number of tensors in the target tensor calculation task; Determining subtasks assigned to each stream based on the target tensor calculation task and the multiple streams; Using an event synchronization mechanism to manage the dependency relationships between the multiple streams, and parallelly starting the kernels corresponding to each stream to complete the calculation of the subtasks.

[0005] This application further provides a tensor calculation device, including: An obtaining unit, configured to obtain a target tensor calculation task; A first determining unit, configured to determine multiple streams based on the number of tensors in the target tensor calculation task; A second determining unit, configured to determine subtasks assigned to each stream based on the target tensor calculation task and the multiple streams; A calculation unit, configured to use an event synchronization mechanism to manage the dependency relationships between the multiple streams, and parallelly start the kernels corresponding to each stream to complete the calculation of the subtasks.

[0006] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above tensor calculation methods when executing the computer program.

[0007] This application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above tensor calculation methods.

[0008] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any one of the above tensor calculation methods.

[0009] Through the present application, the present application discloses a tensor calculation method, an electronic device, a storage medium, and a product, relating to the technical field of data processing. It includes obtaining a target tensor calculation task; determining a plurality of streams based on the number of tensors in the target tensor calculation task; determining subtasks assigned to each stream based on the target tensor calculation task and the plurality of streams; and using an event synchronization mechanism to manage the dependency relationships between the plurality of streams, and starting the kernels corresponding to each stream in parallel to complete the calculation of the subtasks, which solves the technical problems in the related solutions that when a plurality of tensors with small scales perform bit-by-bit operations, they cannot fully occupy the accelerator calculation units, resulting in low resource utilization rate, and each group of tensors has a corresponding started kernel, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, and achieves the technical effects of improving resource utilization rate, reducing the time consumption of starting kernels for tensors, and further improving tensor calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 It is a schematic flowchart of a tensor calculation method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a subtask allocation method provided by an embodiment of the present application; Figure 3 It is a schematic diagram of a thread waiting method provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of a tensor calculation method provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of a tensor calculation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0013] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0014] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0015] The optimal substructure means that the optimal substructure requires the solution of the problem to be recursive, that is, the optimal solution of the original problem depends on the optimal solutions of the subproblems.

[0016] A stream refers to a stream on the accelerator chip where it can ensure that the tasks placed on this stream are executed in the order in which they are placed, and the order between different streams does not affect each other.

[0017] The greedy allocation algorithm refers to an algorithm strategy that makes the optimal (i.e., most favorable) choice at each step, hoping to lead to a globally optimal solution. The greedy algorithm is applicable to problems that can be decomposed into multiple subproblems, and the optimal solutions of each subproblem can be combined into a globally optimal solution.

[0018] A kernel refers to a function that is called on the host but executed on a device or other accelerator. Here, the host is usually a Central Processing Unit (CPU), and the accelerator is usually a Graphics Processing Unit (GPU). These functions are designed to process a large amount of data in parallel, taking advantage of the parallel computing power of the accelerator.

[0019] With the rise of artificial intelligence and the popularization of large models, various operators will emerge. In artificial intelligence deep learning frameworks such as PyTorch and TensorFlow, there is a class of batch processing operation operators. For example, an operator like foreach_add is used for batch processing of tensors. Its input is two lists of tensors (input 1 and input 2), and the output is a list of tensors with element-wise operation results, that is, the corresponding tensors in input 1 and input 2 are operated on in sequence and placed in the respective tensors in the result tensor list. The tensors in these lists are usually stored in different memory locations and cannot be uniformly processed using a single kernel. Moreover, it is often the case that the individual sizes of these tensors are not very large. If calculated sequentially, the kernel needs to be started sequentially, and a relatively long time is consumed in starting the kernel. Also, due to the small number of individual tensors, the utilization rate of the acceleration chip is not high, resulting in waste and relatively more time consumption.

[0020] The following briefly introduces several solutions for tensor calculation methods in related technologies: In Solution A, the conventional approach is to process them sequentially on a single stream.

[0021] Solution B attempts to process them in a single kernel, but it has low generality and too many constraints, and is only available in very few cases.

[0022] Solution C splices multiple small tensors into a large tensor and then performs a single bitwise operation. However, it is limited by the requirement of tensor shape consistency. Even if the conditions are met, there will be additional overheads brought by splicing and splitting, which may be more time-consuming than the conventional approach.

[0023] In the above solutions, there are the following defects: Low utilization rate of accelerator computing power: When the scale of an individual tensor is small (for example, the size is 32x32), the addition operation of each group of tensors cannot fully occupy the accelerator computing units, resulting in idle resources and low utilization rate.

[0024] Accumulation of execution latency: In the serial execution mode, the total latency is the sum of the latencies of all operations, and the multi-task parallel capabilities of the GPU cannot be utilized.

[0025] High kernel startup time consumption: Starting the corresponding kernel for each group of tensors, the serial startup time will accumulate.

[0026] To solve the problems existing in related solutions, an embodiment of the present application provides a tensor calculation method, including: obtaining a target tensor calculation task; determining a plurality of streams based on the number of tensors in the target tensor calculation task; determining subtasks assigned to each stream based on the target tensor calculation task and the plurality of streams; using an event synchronization mechanism to manage the dependency relationships between the plurality of streams, and starting the kernels corresponding to each stream in parallel to complete the calculation of the subtasks, solving the technical problems in related solutions that when multiple tensors of relatively small scales perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding started kernel, resulting in the accumulation of calculation time consumption, and further resulting in low tensor calculation efficiency, achieving the technical effects of improving resource utilization, reducing the time consumption of starting kernels for tensors, and further improving tensor calculation efficiency.

[0027] A tensor calculation method provided by an embodiment of the present disclosure can be executed by an accelerator, such as a GPU, a Tensor Processing Unit (TPU), etc. The tensor calculation method provided by the embodiment of the present disclosure can be applied to fields such as deep learning, scientific computing, and financial evaluation. Taking the application in the field of deep learning as an example, during the training process of a deep learning model (such as a convolutional neural network, a recurrent neural network, etc.), a large number of tensor operations are involved, such as matrix multiplication, convolution operations, etc. This method can determine a plurality of streams according to the number of tensors, allocate the calculation tasks of different layers or the calculation tasks of different batches of data to each stream for parallel execution, and use an event synchronization mechanism to ensure the correctness of the calculation order, thereby significantly accelerating the training process of the model. For example, when training an image classification model, the convolution calculation tasks of different convolutional layers can be allocated to different streams for parallel processing.

[0028] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] Figure 1 It is a schematic flowchart of a tensor calculation method provided by an embodiment of the present disclosure.

[0030] As Figure 1 shown, the method includes the following steps: Step 101, obtaining a target tensor calculation task; In some embodiments, the target tensor calculation task includes two or more input tensor lists and an output tensor list. The two or more input tensor lists may be stored in different memory addresses, and their shapes and sizes may vary. Each tensor in the output tensor list is the element-wise operation result (such as addition, multiplication, etc.) of the tensors at the corresponding positions in the input tensor lists.

[0031] In some embodiments, the target tensor calculation task is used to perform bitwise operations on the input tensor, and the bitwise operations may include element-wise addition, element-wise subtraction, etc.

[0032] In some embodiments, taking the foreach_add operator in the deep learning framework as an example for illustration below, the other foreach_sub operators and other operators use a similar scheme, which will not be elaborated in this application. For two input tensor lists, tensor list A (InputA_list) and tensor list B (InputB_list), both tensor list A and tensor list B have N tensors, and the output tensor list (Output_list). The operation performed on each tensor is: Output_list[i]=InputA_list[i]+InputB_list[i]; where i represents the i-th tensor in the tensor list, InputA_list[i] represents the i-th tensor in tensor list A, InputB_list[i] represents the i-th tensor in tensor list B, and Output_list[i] represents the i-th tensor in the output tensor list.

[0033] Step 102, determine multiple streams based on the number of tensors in the target tensor calculation task; In some embodiments, the number of tensors in the target tensor calculation task refers to the total number of tensors participating in the calculation task in a single tensor list.

[0034] In some embodiments, when the task involves a large number of small-scale tensors, a single stream cannot fully utilize the hardware resources, while multiple streams can start multiple kernels simultaneously to improve the utilization rate of computing resources.

[0035] In some embodiments, the number of required streams can be dynamically calculated according to the characteristics of the chip itself and the tensor situation in the current target tensor calculation task.

[0036] In some embodiments, the number of streams is also affected by hardware limitations and resource contention. Among them, hardware limitations include the maximum number of streams, computing units, and memory bandwidth, and resource contention includes memory resource contention and computing resource contention. Specifically, different hardware platforms (such as GPUs and TPUs) support a limited maximum number of concurrent streams. If the number of tensors is large, the number of streams can be limited according to the maximum number of streams supported by the hardware to avoid excessive management costs. Each stream occupies a certain amount of computing resources and video memory bandwidth. If too many streams are created, it may cause the computing unit or memory bandwidth to become a bottleneck, thereby reducing performance. When multiple streams are running simultaneously, if they all require a large amount of video memory or other types of memory, it may lead to memory resource contention, which will not only slow down the data transfer speed but also may cause out-of-memory errors. When multiple streams need to use the same computing resources simultaneously, resource contention will occur, affecting the overall efficiency.

[0037] In some embodiments, by determining multiple streams based on the number of tensors in the target tensor computing task, the parallel computing capabilities of hardware resources can be fully utilized.

[0038] Step 103: Based on the target tensor computing task and multiple streams, determine the subtasks assigned to each stream. In some embodiments, based on the target tensor computing task and multiple streams, the target tensor computing task can be evenly distributed to each stream to achieve that the computing tasks on multiple streams are completed almost simultaneously, thereby improving the efficiency of tensor computing.

[0039] In some embodiments, by determining the subtasks assigned to each stream based on the target tensor computing task and multiple streams, the problems of low resource utilization rate and cumulative kernel startup time caused by processing the target tensor computing task on one stream can be avoided, effectively improving the utilization rate of hardware resources, reducing resource contention, and ensuring the efficiency and stability of task execution.

[0040] Step 104: Use the event synchronization mechanism to manage the dependencies between multiple streams, and parallelly start the kernels corresponding to each stream to complete the calculation of subtasks.

[0041] In some embodiments, the event synchronization mechanism is a technical means for managing the execution order of concurrent tasks. Usually, an event is used to mark a specific operation point or state. For example, in GPU programming, an event is a lightweight object that can be used to record the completion status of operations in a certain stream and allow other streams to wait for the occurrence of this event, ensuring that the tasks between different streams are executed in the correct order and avoiding data contention or resource conflicts.

[0042] In some embodiments, the dependency relationship between multiple streams refers to a certain order or constraint condition between tasks in different streams. For example, the task in stream A must start after the task in stream B is completed, or the task in stream C needs to wait until both stream A and stream B are completed before it can be executed.

[0043] In some embodiments, starting the kernels corresponding to each stream in parallel means starting different kernel tasks simultaneously on multiple streams to fully utilize the hardware resources.

[0044] In some embodiments, starting the kernel corresponding to each stream is usually achieved by calling an API on the host to start the kernel, which is then executed by the device.

[0045] In some embodiments, by using an event synchronization mechanism to manage the dependency relationship between multiple streams and starting the kernels corresponding to each stream in parallel to complete the calculation of subtasks, the utilization rate of computing resources can be improved, the computing latency can be reduced, the resource utilization can be optimized, and the correctness and execution order of tasks can be ensured.

[0046] Through this application, a target tensor calculation task is obtained; based on the number of tensors in the target tensor calculation task, multiple streams are determined; based on the target tensor calculation task and multiple streams, the subtasks assigned to each stream are determined; an event synchronization mechanism is used to manage the dependency relationship between multiple streams, and the kernels corresponding to each stream are started in parallel to complete the calculation of subtasks, solving the technical problems in related solutions that when multiple tensors of a small scale perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding kernel startup, resulting in the accumulation of computing time consumption and thus low tensor calculation efficiency, achieving the technical effects of improving resource utilization, reducing the time consumption of tensor kernel startup, and thus improving tensor calculation efficiency.

[0047] In some embodiments, determining multiple streams based on the number of tensors in the target tensor calculation task includes: Based on the tensors in the target tensor calculation task, determining the number of tensors smaller than a preset size, where the size is the number of elements in the tensor; In some embodiments, the preset size can be determined by the smallest size that can theoretically fill the GPU. Specifically, the mathematical expression for determining the preset size is: M = Size * 25%; where M represents the preset size, Size represents the smallest size that can theoretically fill the GPU, and 25% is a preset value. Generally, it is considered that when the utilization rate reaches more than 50%, the computing resources of the GPU can be considered to be utilized relatively fully, and the grouping is at least two groups. Therefore, the preset value is determined to be 25%, and it can also be determined according to needs. This application does not limit this.

[0048] Furthermore, the minimum size that can theoretically fill up the GPU can be determined by the number of streaming multiprocessors and the recommended typical value of thread blocks. Specifically, the mathematical expression for determining the minimum size that can theoretically fill up the GPU is: Size = S * B; where S represents the number of multiprocessors, B represents the recommended typical value of thread blocks, and Size represents the minimum size that can theoretically fill up the GPU.

[0049] In some embodiments, taking S as 108 and B as 512 as an example, M = 108 * 512 * 25% = 13824, that is, the preset size is 13824.

[0050] In some embodiments, a two-dimensional tensor with a shape of (3, 4) contains 12 elements and its size is 12; a three-dimensional tensor with a shape of (2, 2, 2) contains 8 elements and its size is 8.

[0051] In some embodiments, by calculating the tensors in the target tensor calculation task and determining the number of tensors smaller than the preset size, small-scale tasks that may not fully utilize the computing power of the hardware accelerator can be identified.

[0052] Based on the number of tensors smaller than the preset size and the number of preset tensor groups, determine the target grouping number.

[0053] Based on the target grouping number, determine multiple streams.

[0054] In some embodiments, the number of preset tensor groups is usually 2, that is, when the data volume of more than 2 groups of tensors is less than the preset size theoretically, grouping can be considered. However, considering the unevenness of grouping and the additional overhead brought by grouping, the overall benefit may not be achieved when the number is small. In this application, the specific value of the number of preset tensor groups is not limited. Taking the number of preset tensor groups in this application as 8 as an example, grouping starts when the data of more than 8 groups of tensors is less than the preset size.

[0055] In some embodiments, controlling the maximum number of groups in this application to be 8 can meet most scenarios. Of course, the maximum number of groups can also be specifically determined according to specific hardware and tests.

[0056] In some embodiments, determine the target grouping number as the number of streams, that is, finally use the target grouping number of streams to process the target tensor calculation task.

[0057] In some embodiments, determining the number of tensors smaller than the preset size based on the tensors in the target tensor calculation task includes: Based on the tensors in the target tensor calculation task, determine the size of each group of tensors; In some embodiments, the target tensor calculation task includes at least two input tensor lists.

[0058] In some embodiments, each group of tensors is a tensor for performing bitwise operations in the input tensor list.

[0059] Based on the size of each group of tensors and the preset size, determine the number of tensors smaller than the preset size.

[0060] In some embodiments, compare the size of each group of tensors with the preset size, filter out the tensors with a size smaller than the preset size, and determine the number of the filtered tensors.

[0061] In some embodiments, based on the number of tensors smaller than the preset size and the number of preset tensor groups, determining the target grouping number includes: In response to the number of tensors smaller than the preset size being lower than the number of preset tensor groups, determine the first preset value as the target grouping number.

[0062] In some embodiments, the number of tensors smaller than the preset size being lower than the number of preset tensor groups indicates that the number of small tensors in the target tensor calculation task is not very large (not meeting the grouping requirements). Therefore, directly calculating the target tensor calculation task using one stream can make better use of computing resources. Thus, the first preset value is 1.

[0063] In some embodiments, based on the number of tensors smaller than the preset size and the number of preset tensor groups, determining the target grouping number includes: In response to the number of tensors smaller than the preset size not being lower than the number of preset tensor groups, compare the number of tensors smaller than the preset size with the number of preset tensor groups to obtain a first ratio; In some embodiments, the number of preset tensor groups represents the expected number of small tensor groups, which is used to measure whether the number of small tensors in the current task is large enough to determine whether multiple streams need to be used for parallel computing.

[0064] In some embodiments, the number of tensors smaller than the preset size not being lower than the number of preset tensor groups indicates that there are many small-scale tensors in the target tensor calculation task. More streams can be created to increase the parallelism and make full use of the hardware resources.

[0065] In some embodiments, the first ratio is obtained by dividing the number of tensors smaller than the preset size by the number of preset tensor groups.

[0066] Based on the first ratio, determine the target grouping number.

[0067] In some embodiments, the first ratio is used as the target number of groups, that is, the target number of groups can be dynamically determined based on the first ratio. Specifically, during the task execution, the target number of groups can be dynamically adjusted according to the first ratio monitored in real time. For example, if the first ratio increases, the number of groups can be dynamically increased to handle more tasks of small tensors; if the first ratio decreases, the number of groups can be reduced to save resources.

[0068] In some embodiments, based on the target tensor calculation task and multiple streams, determining the subtasks assigned to each stream includes: Based on the target tensor calculation task, obtain the number of elements in each tensor and the continuity of each tensor; In some embodiments, the continuity is used to indicate the input / output continuity of the elements in the tensor.

[0069] In some embodiments, the number of elements in each tensor refers to the total number of all elements in the tensor. In deep learning frameworks (such as PyTorch or TensorFlow), the built-in functions can be directly used to obtain the number of elements in the tensor.

[0070] In some embodiments, the input / output continuity of the elements in the tensor means that all elements of the tensor are continuously arranged in memory, then the tensor is called continuous. In this case, the efficiency of accessing the data of the tensor is higher. If the tensor undergoes some operations (such as slicing, transposing, etc.), its elements may become discontinuous in memory. For discontinuous tensors, when performing bit-by-bit calculations, since additional position information needs to be calculated, the time consumption is longer than that of continuous tensors, which will increase the time overhead of data access.

[0071] Based on the number of elements in each tensor and the continuity of each tensor, determine the heaviness value of each group of tensors in the target tensor calculation task. The size of the heaviness value is used to indicate the computational heaviness of the task; In some embodiments, the more elements in the tensor, the higher the computational complexity, and the greater the heaviness value of the corresponding tensor.

[0072] In some embodiments, different weights can be determined for the heaviness value of the tensor according to the specific type of computational operation (such as addition, multiplication, convolution, etc.).

[0073] Based on the heaviness value of each group of tensors in the target tensor calculation task and multiple streams, determine the subtasks assigned to each stream.

[0074] In some embodiments, based on the heaviness value of each group of tensors in the target tensor calculation task and multiple streams, the target tensor calculation task can be assigned to each stream as evenly as possible according to the size of the heaviness value, where each stream corresponds to one or more groups of tensors, and each group of tensors corresponds to a stream.

[0075] In some embodiments, by determining the heavy value and multiple streams of each group of tensors in the target tensor calculation task, and determining the subtasks assigned to each stream, load balancing can be ensured, resource contention can be reduced, and the utilization rate of hardware resources can be maximized.

[0076] In some embodiments, determining the heavy value of each group of tensors in the target tensor calculation task based on the number of elements in each tensor and the continuity of each tensor includes: Based on the continuity of each tensor, determining the number of discontinuous tensors in the target tensor calculation task; In some embodiments, in PyTorch, the is_contiguous function can be used to determine whether a tensor is continuous. For each tensor, if the is_contiguous function returns false, it is counted as a discontinuous tensor, and the number of all discontinuous tensors in the target tensor calculation task is counted.

[0077] Based on the number of elements in each tensor and the number of discontinuous tensors, determining the heavy value of each group of tensors in the target tensor calculation task.

[0078] In some embodiments, discontinuous tensors increase the memory access overhead, so an additional penalty factor is required. The mathematical expression for determining the heavy value of a tensor is as follows: Weight[i]=InputA_list[i].numel()*(1+Num_discontinuous); Where, Weight[i] represents the heavy value, InputA_list[i].numel() represents the number of elements in the i-th group of tensors participating in the calculation. Since the number of elements in these 3 tensors in the bitwise operation group is the same size, any one of these tensors can be taken. Num_discontinuous represents the number of discontinuous tensors. For example, in the previous embodiments, there are 2 input tensors, 1 output tensor, the input tensor list A (InputA_list) and the tensor list B (InputB_list), and the output tensor list (Output_list). If 2 of these 3 tensors are discontinuous and 1 tensor is continuous, then the value of Num_discontinuous is 2.

[0079] In some embodiments, determining the subtasks assigned to each stream based on the heavy value of each group of tensors in the target tensor calculation task and multiple streams includes: Based on the heavy value of each group of tensors in the target tensor calculation task, sorting them in descending order of the heavy value to obtain the sorted heavy value; In some embodiments, the heaviness value reflects the computational complexity and memory access overhead of each group of tensors. The larger the heaviness value, the more computationally intensive the group of tensors is.

[0080] In some embodiments, by sorting the heaviness values from largest to smallest, tasks with higher computational complexity can be processed first, thereby optimizing task scheduling and resource allocation.

[0081] In some embodiments, by sorting in the order of heaviness values from largest to smallest to obtain the sorted heaviness values, tasks with higher computational complexity can be effectively identified, and task scheduling and resource allocation can be optimized accordingly.

[0082] Based on the sorted heaviness values and multiple streams, determine the subtasks assigned to each stream.

[0083] In some embodiments, after the assignment based on the sorted heaviness values and multiple streams is completed, the total amount of tasks assigned to each stream is approximately equal.

[0084] In some embodiments, by based on the sorted heaviness values and multiple streams, tasks with higher heaviness values are preferentially assigned to the stream with the lightest current load, avoiding overloading of some streams while other streams are idle.

[0085] In some embodiments, by based on the sorted heaviness values and multiple streams to determine the subtasks assigned to each stream, the efficiency of parallel computing and resource utilization can be improved.

[0086] In some embodiments, during the task execution process, if it is found that some streams are completed faster, the remaining tasks can be dynamically reassigned. For example, an event synchronization mechanism can be used to monitor the completion status of each stream and adjust the task assignment in real time.

[0087] In some embodiments, if some tasks have higher priorities, higher weights can be given to these tasks during assignment and they are preferentially assigned to the streams with more idle resources.

[0088] In some embodiments, determining the subtasks assigned to each stream based on the sorted heaviness values and multiple streams includes: Based on the sorted heaviness values and multiple streams, use a preset algorithm to assign the target tensor computation tasks to multiple streams to obtain the subtasks assigned to each stream. The preset algorithm is used to assign the heaviness value to the stream with the lowest load each time.

[0089] In some embodiments, the preset algorithm in this application takes the greedy algorithm as an example, such as Figure 2 shown Figure 2 is a schematic diagram of a subtask assignment method provided by an embodiment of this application. Specifically, Figure 2Taking the calculation of 6 sets of tensor data using 3 streams (stream1, stream2, and stream3 from left to right) as an example, Figure 2 The numbers in the gray cells are the calculated computational weights (heaviness). The process from step 1 to step 2 is the sorting process, and the process from step 3 to step 8 is the task allocation process. After completing the task division through the greedy algorithm, it can be seen from Figure 2 the last step that the task amounts on each stream are as follows: stream1: 400, stream2: 400, stream3: 380. It can be seen that the task amounts allocated to each stream after division are roughly equivalent, and the task allocation has an optimal substructure, and the solution obtained using the greedy algorithm is the global optimal solution.

[0090] In some embodiments, by using a preset algorithm to allocate the target tensor calculation task to multiple streams based on the sorted heaviness values and multiple streams, the subtasks allocated to each stream can be obtained, which can ensure that the task load of each stream is as balanced as possible, thereby improving the efficiency of parallel computing and resource utilization.

[0091] In some embodiments, after determining the subtasks allocated to each stream based on the target tensor calculation task and multiple streams, the tensor calculation method further includes: Creating a main thread and multiple sub-threads; In some embodiments, the main thread is used for task initialization, scheduling, and coordination, and each sub-thread is bound to a stream and is responsible for executing the tasks in that stream.

[0092] In some embodiments, the main thread is responsible for creating events and allocating the events to each stream. When the sub-threads execute tasks, they coordinate the task execution order of different streams through the events.

[0093] Determining the streams corresponding to the main thread and multiple sub-threads based on multiple streams; In some embodiments, each stream maintains a task queue.

[0094] Putting the subtasks into the task pools of the corresponding streams and notifying the threads of the corresponding streams.

[0095] In some embodiments, the task pool of a stream is used to store the subtasks allocated to that stream.

[0096] In some embodiments, after the main thread puts the tasks into the task pool of a stream, it notifies the corresponding sub-thread to start execution.

[0097] In some embodiments, when a thread receives information that a task has arrived, it goes to the task pool of the corresponding stream to retrieve the task.

[0098] In some embodiments, an event synchronization mechanism is used to manage dependencies between multiple streams, and parallelly starting kernels corresponding to each stream to complete the calculation of subtasks includes: Create a first event for the stream corresponding to the child thread, where the first event is used to synchronize subtasks between multiple streams; In some embodiments, the first event is a synchronization point used to mark the completion status of tasks in a certain stream. When a stream completes some tasks, this event is recorded, and other streams can wait for this event to complete before continuing to execute subsequent tasks. The first event can ensure that the calculation on this stream starts only after the calculation on the main stream is completed.

[0099] In some embodiments, the first event can be reused later to avoid the additional overhead caused by creating it each time.

[0100] In some embodiments, by creating the first event for the stream corresponding to the child thread and using this event to synchronize subtasks between multiple streams, the dependency relationship between tasks and resource contention problems can be effectively solved. The reuse mechanism of the event further improves efficiency and reduces resource overhead.

[0101] Based on the first event, determine a second event on the stream corresponding to the main thread. The second event includes a first recording event and a first waiting event. The first recording event is used to record key points of the stream corresponding to the main thread, and the first waiting event is used to wait for subtasks on the stream corresponding to the main thread to complete; In some embodiments, the first recording event can be implemented through the recordevent function of the accelerator.

[0102] In some embodiments, the first waiting event can be implemented through the wait event function of the accelerator.

[0103] In response to the completion of subtasks on the stream corresponding to the main thread, parallelly start kernels on the streams corresponding to multiple child threads to complete the calculation of subtasks.

[0104] In some embodiments, in parallel computing, after the subtasks on the stream corresponding to the main thread are completed, it can trigger the start of kernels on the streams corresponding to multiple child threads to complete the calculation of the remaining subtasks. This mechanism can make full use of hardware resources, thereby improving the calculation efficiency.

[0105] In some embodiments, for the threads of each stream, to ensure that the computations on this stream start only after the computations on the main stream are completed, an event is created, and the function of recording the event of the accelerator is called. Then, the function of waiting for the event of the accelerator is called, and the main thread is notified. This operation does not block the thread but makes records to ensure that the kernels started subsequently on this stream will not start before the previous task of the main thread is completed, thus avoiding the use of dirty data or program crashes. Then, the kernels allocated to this stream are started in sequence to perform computations.

[0106] In some embodiments, an event synchronization mechanism is used to manage the dependencies between multiple streams. Parallelly starting the kernels corresponding to each stream to complete the computation of subtasks includes: In response to receiving the notification messages of the streams corresponding to multiple child threads, starting the kernel on the stream corresponding to the main thread; In some embodiments, the notification messages of the streams corresponding to multiple child threads refer to that after each child thread completes its corresponding task, a notification will be generated indicating that the task has been successfully completed.

[0107] In some embodiments, in addition to the simple task completion notification, the notification message may also contain information about the task execution status, such as whether the task is successful, any errors or exceptions encountered, etc.

[0108] In some embodiments, in parallel computing, when the streams corresponding to multiple child threads complete their tasks, they can send notification messages to the main thread. After receiving these notification messages, the main thread starts the kernel on its corresponding stream to complete the subsequent computation tasks, which can achieve the dependency management between tasks and make full use of the parallel capabilities of multi-threading and multi-streams.

[0109] Obtaining third events on the streams corresponding to multiple child threads, where the third events include second recording events and second waiting events. The second recording events are used to record the key points of the streams corresponding to the child threads, and the second waiting events are used to wait for the subtasks on the streams corresponding to the child threads to complete; In some embodiments, the second recording events and the second waiting events can be implemented in the same way as the aforementioned first recording events and first waiting events, which will not be elaborated here.

[0110] In some embodiments, the third events are generated by the streams corresponding to multiple child threads and are used to coordinate the task execution order between the child threads.

[0111] In response to the completion of the subtasks on the streams corresponding to multiple child threads, performing the first task on the stream corresponding to the main thread to complete the computation of the subtasks.

[0112] In some embodiments, the first task refers to other tasks in the main thread.

[0113] In some embodiments, for the main thread, after waiting for all other stream notification messages, the tasks of the kernels to be executed on the main thread are started in sequence. After the task startup is completed, the recordevent functions of other streams and the functions waiting for these events are called in sequence to ensure that the kernels on stream1 and stream2 have been completed before executing the kernels behind the main stream, so that other tasks on the main thread stream can continue to be executed later and the calculated results that have been completed can be obtained. This operation does not block the main thread either, but only ensures the execution order on the stream.

[0114] In some embodiments, as Figure 3 shown, Figure 3 is a schematic diagram of a thread waiting method provided by an embodiment of the present application. Specifically, after calling to record the event on the main stream, the kernel 1 that has been initiated on the main stream at this moment will be recorded. When calling to wait for this event, it is ensured that the kernel 3 on stream 1 will continue to execute only after the kernel 1 on the main stream has been completed. Similarly, the kernel on the main stream will continue to execute only after the kernel 3 on stream 1 has been completed. In this way, the execution order is ensured and the calling thread is not blocked.

[0115] In some embodiments, as Figure 4 shown, Figure 4 is a schematic flowchart of a tensor calculation method provided by an embodiment of the present application. The tasks on each stream are calculated and allocated, and the allocated tasks are respectively placed in the main stream (main stream) task pool, stream 1 task pool, and stream 2 task pool. For thread 1 and thread 2, the stream events of the main thread are respectively recorded, and then the functions waiting for the main thread events are called respectively to notify the main thread and start the kernels of the child threads; for the main thread, after receiving the notification messages from thread 1 and thread 2, the kernel of the main thread is started, other stream events are recorded, and other tasks on the main thread are executed after waiting for the stream events of other threads.

[0116] Through this application, a target tensor calculation task is obtained; based on the number of tensors in the target tensor calculation task, multiple streams are determined; based on the target tensor calculation task and the multiple streams, subtasks assigned to each stream are determined; an event synchronization mechanism is used to manage the dependencies between the multiple streams, and kernels corresponding to each stream are started in parallel to complete the calculation of the subtasks, solving the technical problem in related solutions that when multiple tensors of relatively small scales perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding kernel startup, resulting in the accumulation of calculation time consumption and further resulting in low tensor calculation efficiency, achieving the technical effect of improving resource utilization, reducing the time consumption of tensor kernel startup, and further improving tensor calculation efficiency.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0118] An embodiment of this application also provides a tensor calculation device 500, Figure 5 which is a schematic structural diagram of a tensor calculation device provided by an embodiment of the present disclosure, as Figure 5 shown, and includes: An acquisition unit 501, configured to acquire a target tensor calculation task; A first determination unit 502, configured to determine multiple streams based on the number of tensors in the target tensor calculation task; A second determination unit 503, configured to determine subtasks assigned to each stream based on the target tensor calculation task and the multiple streams; A calculation unit 504, configured to use an event synchronization mechanism to manage the dependencies between the multiple streams, and start kernels corresponding to each stream in parallel to complete the calculation of the subtasks.

[0119] Through this application, a target tensor calculation task is obtained; based on the number of tensors in the target tensor calculation task, multiple streams are determined; based on the target tensor calculation task and the multiple streams, subtasks assigned to each stream are determined; an event synchronization mechanism is used to manage the dependencies between the multiple streams, and kernels corresponding to each stream are started in parallel to complete the calculation of the subtasks, solving the technical problem in related solutions that when multiple tensors of relatively small scales perform bit-by-bit operations, they cannot fully occupy the accelerator computing units, resulting in low resource utilization, and each group of tensors has a corresponding kernel startup, resulting in the accumulation of calculation time consumption and further resulting in low tensor calculation efficiency, achieving the technical effect of improving resource utilization, reducing the time consumption of tensor kernel startup, and further improving tensor calculation efficiency.

[0120] Further, in a possible implementation manner of an embodiment of the present disclosure, the first determination unit 502 is configured to: Based on the tensors in the target tensor calculation task, determine the number of tensors smaller than a preset size, where the size is the number of elements in the tensor; Based on the number of tensors smaller than the preset size and the number of preset tensor groups, determine the target grouping number; Based on the target grouping number, determine multiple streams.

[0121] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determining unit 502 is configured to: Based on the tensors in the target tensor calculation task, determine the size of each group of tensors. The target tensor calculation task includes at least two input tensor lists, and each group of tensors is the tensors for bitwise operations in the input tensor list; Based on the size of each group of tensors and the preset size, determine the number of tensors smaller than the preset size.

[0122] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determining unit 502 is configured to: In response to the number of tensors smaller than the preset size being lower than the number of preset tensor groups, determine the first preset value as the target grouping number.

[0123] Further, in a possible implementation manner of the embodiments of the present disclosure, the first determining unit 502 is configured to: In response to the number of tensors smaller than the preset size not being lower than the number of preset tensor groups, compare the number of tensors smaller than the preset size with the number of preset tensor groups to obtain a first ratio; Based on the first ratio, determine the target grouping number.

[0124] Further, in a possible implementation manner of the embodiments of the present disclosure, the second determining unit 503 is configured to: Based on the target tensor calculation task, obtain the number of elements in each tensor and the continuity situation of each tensor, where the continuity situation is used to indicate the input-output continuity situation of the elements in the tensor; Based on the number of elements in each tensor and the continuity situation of each tensor, determine the workload value of each group of tensors in the target tensor calculation task, and the magnitude of the workload value is used to indicate the computational workload of the task; Based on the workload value of each group of tensors in the target tensor calculation task and multiple streams, determine the subtasks assigned to each stream.

[0125] Further, in a possible implementation manner of the embodiments of the present disclosure, the second determining unit 503 is configured to: Based on the continuity situation of each tensor, determine the number of discontinuous tensors in the target tensor calculation task; Based on the number of elements in each tensor and the number of discontinuous tensors, determine the workload value of each group of tensors in the target tensor calculation task.

[0126] Further, in a possible implementation manner of the embodiments of the present disclosure, the second determination unit 503 is configured to: Calculate the workload value of each group of tensors in the target tensor calculation task, sort them in descending order of the workload value to obtain the sorted workload values; Based on the sorted workload values and multiple streams, determine the subtasks assigned to each stream.

[0127] Further, in a possible implementation manner of the embodiments of the present disclosure, the second determination unit 503 is configured to: Based on the sorted workload values and multiple streams, use a preset algorithm to allocate the target tensor calculation task to multiple streams to obtain the subtasks assigned to each stream, and the preset algorithm is used to allocate the workload value to the stream with the lowest load each time.

[0128] Further, in a possible implementation manner of the embodiments of the present disclosure, the tensor calculation device 500 further includes a notification unit, and the notification unit is configured to: Create a main thread and multiple sub-threads; Based on multiple streams, determine the streams corresponding to the main thread and multiple sub-threads; Put the subtasks into the task pools of the corresponding streams and notify the threads of the corresponding streams.

[0129] Further, in a possible implementation manner of the embodiments of the present disclosure, the calculation unit 504 is configured to: Create a first event for the stream corresponding to the sub-thread, and the first event is used to synchronize the subtasks between multiple streams; Based on the first event, determine a second event on the stream corresponding to the main thread. The second event includes a first recording event and a first waiting event. The first recording event is used to record the key points of the stream corresponding to the main thread, and the first waiting event is used to wait for the subtasks on the stream corresponding to the main thread to complete; In response to the completion of the subtasks on the stream corresponding to the main thread, parallelly start the kernels on the streams corresponding to multiple sub-threads to complete the calculation of the subtasks.

[0130] Further, in a possible implementation manner of the embodiments of the present disclosure, the calculation unit 504 is configured to: In response to receiving the notification message of the streams corresponding to multiple sub-threads, start the kernel on the stream corresponding to the main thread; Obtain a third event on the streams corresponding to multiple sub-threads. The third event includes a second recording event and a second waiting event. The second recording event is used to record the key points of the stream corresponding to the sub-thread, and the second waiting event is used to wait for the subtasks on the stream corresponding to the sub-thread to complete; In response to the completion of subtasks on the streams corresponding to multiple child threads, execute the first task on the stream corresponding to the main thread to complete the calculation of the subtasks.

[0131] For the description of the features in the corresponding embodiments of the tensor calculation device, reference can be made to the relevant descriptions in the corresponding embodiments of the tensor calculation method, which will not be elaborated here one by one.

[0132] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the tensor calculation method.

[0133] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the tensor calculation method when running.

[0134] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., various media that can store computer programs.

[0135] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the tensor calculation method.

[0136] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the tensor calculation method.

[0137] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0138] The above has introduced in detail a tensor calculation method, an electronic device, a storage medium, and a product provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A tensor calculation method, characterized in that: include: Get the target tensor computing task; Determining a plurality of flows based on the number of tensors in the target tensor computing task; Based on the target tensor computing task and the multiple streams, determining a subtask assigned to each stream; An event synchronization mechanism is used to manage the dependencies between the multiple streams, and the kernel corresponding to each stream is started in parallel to complete the calculation of the subtask.

2. The tensor calculation method according to claim 1, characterized in that: The determining of the plurality of flows based on the number of tensors in the target tensor computing task comprises: Determine, based on the tensors in the target tensor computing task, the number of tensors that are smaller than a preset size, where the size is the number of elements in the tensor; Determining a target number of groups based on the number of tensors smaller than a preset size and the number of preset tensor groups; Based on the target number of packets, a plurality of flows are determined.

3. The tensor calculation method according to claim 2, characterized in that: The determining the number of tensors smaller than a preset size based on the tensors in the target tensor computing task comprises: Determine the size of each group of tensors based on the tensors in the target tensor calculation task, wherein the target tensor calculation task includes at least two input tensor lists, and each group of tensors is a tensor in the input tensor list that performs a bitwise operation; Based on the size of each group of tensors and a preset size, the number of tensors smaller than the preset size is determined.

4. The tensor calculation method according to claim 2, characterized in that: The determining the target number of groups based on the number of tensors smaller than the preset size and the number of preset tensor groups includes: In response to the number of tensors smaller than a preset size being lower than the number of preset tensor groups, a first preset value is determined as the target grouping number.

5. The tensor calculation method according to claim 2, characterized in that: The determining the target number of groups based on the number of tensors smaller than the preset size and the number of preset tensor groups includes: In response to the number of tensors smaller than the preset size being not less than the number of preset tensor groups, comparing the number of tensors smaller than the preset size with the number of preset tensor groups to obtain a first ratio; Based on the first ratio, the target number of groups is determined.

6. The tensor calculation method according to claim 1, characterized in that: The determining, based on the target tensor computing task and the multiple streams, the subtasks allocated to each stream comprises: Based on the target tensor calculation task, obtain the number of elements in each tensor and the continuity of each tensor, where the continuity is used to indicate the input and output continuity of the elements in the tensor; Based on the number of elements in each tensor and the continuity of each tensor, determine the heaviness value of each group of tensors in the target tensor calculation task, the size of the heaviness value is used to indicate the calculation heaviness of the task; The subtask allocated to each stream is determined based on the heavy value of each group of tensors in the target tensor calculation task and the multiple streams.

7. The tensor calculation method according to claim 6, characterized in that: The determining, based on the number of elements in each tensor and the continuity of each tensor, the heavy value of each group of tensors in the target tensor calculation task comprises: Based on the continuity of each tensor, determine the number of discontinuous tensors in the target tensor calculation task; Based on the number of elements in each tensor and the number of discontinuous tensors, a heavy value of each group of tensors in the target tensor calculation task is determined.

8. The tensor calculation method according to claim 6, characterized in that: The determining the subtask assigned to each stream based on the heavy value of each group of tensors in the target tensor computing task and the multiple streams comprises: Based on the heavy value of each group of tensors in the target tensor calculation task, sort the heavy values ​​from large to small to obtain sorted heavy values; The subtask allocated to each flow is determined based on the sorted heavy values ​​and the multiple flows.

9. The tensor calculation method according to claim 8, characterized in that: The determining the subtask assigned to each flow based on the sorted heavy values ​​and the multiple flows comprises: Based on the sorted heavy values ​​and the multiple streams, the target tensor calculation task is allocated to the multiple streams using a preset algorithm to obtain the subtasks allocated to each stream, and the preset algorithm is used to allocate the heavy value to the stream with the lowest load each time.

10. The tensor calculation method according to claim 1, characterized in that: After determining the subtasks assigned to each stream based on the target tensor computing task and the multiple streams, the method further includes: Create a main thread and multiple child threads; Based on the multiple flows, determining flows corresponding to the main thread and the multiple sub-threads; The subtask is placed into the task pool of the corresponding flow, and the thread of the corresponding flow is notified.

11. The tensor calculation method according to claim 10, characterized in that: The method of managing the dependencies between the multiple streams by using an event synchronization mechanism and starting the kernel corresponding to each stream in parallel to complete the calculation of the subtask includes: Creating a first event of the stream corresponding to the sub-thread, wherein the first event is used to synchronize subtasks between the multiple streams; Based on the first event, determining a second event on the stream corresponding to the main thread, the second event comprising a first recording event and a first waiting event, the first recording event being used to record a key point of the stream corresponding to the main thread, and the first waiting event being used to wait for a subtask on the stream corresponding to the main thread to be completed; In response to the completion of the subtask on the stream corresponding to the main thread, kernels on the streams corresponding to the multiple subthreads are started in parallel to complete the calculation of the subtask.

12. The tensor calculation method according to claim 11, characterized in that: The method of managing the dependencies between the multiple streams by using an event synchronization mechanism and starting the kernel corresponding to each stream in parallel to complete the calculation of the subtask includes: In response to receiving notification messages of the streams corresponding to the multiple sub-threads, starting a kernel on the stream corresponding to the main thread; Obtaining a third event on the streams corresponding to the multiple sub-threads, the third event comprising a second recording event and a second waiting event, the second recording event being used to record a key point of the stream corresponding to the sub-thread, and the second waiting event being used to wait for a subtask on the stream corresponding to the sub-thread to be completed; In response to the subtasks on the streams corresponding to the multiple subthreads being completed, executing the first task on the stream corresponding to the main thread to complete the calculation of the subtasks.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the tensor calculation method as claimed in any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the tensor calculation method according to any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the tensor calculation method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • End-side cloud collaborative distributed computing method, device, equipment and medium

    CN117931447A

  • Computing power engine construction method and device, equipment and storage medium

    CN119358617A

  • Cloud data processing system based on artificial intelligence algorithm

    CN119473645A

  • Calculating a Relation Indicator for a Relation Between Entities

    US20160300149A1

Cited By

  • Vector coprocessor and vector calculation method

    CN120255957A