Task scheduling, preprocessing, processing methods and devices, processing units, media

By grouping processing units and allowing them to transmit data to each other, the problem of insufficient cache data multiplexing in sparse tensor multiplication is solved, and efficient cache resource utilization and computing performance improvement is achieved under the large sparse change range.

CN112596872BActive Publication Date: 2025-07-04LYNXI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011483562.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-15
Publication Date
2025-07-04
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

In the prior art, when the graphics processor processes multiplication of sparse tensors and low-rank tensors, it cannot ensure the cache data multiplexing rate, resulting in the impact of operating performance, especially in the case of large sparse change range, cache resources are wasted or the multiplexing rate is low.

Method used

The processing units are divided into packets, each packet contains multiple processing units, and the blocks of tensors are mapped into these packets, so that the processing units in the packets can transmit data to each other, dynamically adjust the cache capacity according to sparseness, and ensure the cache data multiplexing rate.

Benefits of technology

In the case of large sparse change range, satisfactory cache data reuse rate is achieved, avoiding waste of cache resources and improving computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112596872B_ABST
    Figure CN112596872B_ABST
Patent Text Reader

Abstract

The present disclosure provides a scheduling method for computing tasks, including: dividing at least one processing unit group, each of the processing unit groups including a plurality of processing units; mapping at least one block of a first tensor to at least one of the processing unit groups, so that the plurality of processing units in the processing unit group process the computing tasks corresponding to the blocks; different ones of the processing unit groups correspond to different ones of the blocks; wherein, the processing unit has a cache, and the plurality of processing units in the same processing unit group can transmit data to each other. The present disclosure also provides a preprocessing method for computing tasks, a processing method for computing tasks, a controller, a processing unit, an electronic device, a preprocessing device, and a computer-readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a scheduling method for computing tasks, a preprocessing method for computing tasks, a processing method for computing tasks, a controller, a processing unit, an electronic device, a preprocessing device, and a computer-readable medium. Background Art

[0002] Sparse high-dimensional data (such as sparse high-dimensional tensors) are widely used in real life, and the computation between sparse high-dimensional data thus becomes an important workload. The multiplication between tensors is a common computation between sparse high-dimensional data in the process of tensor decomposition.

[0003] In some related technologies, when a graphics processing unit (GPU) processes the multiplication of a sparse tensor and a low-rank tensor, each core of the GPU processes the computation involved in a block of the sparse tensor. The general process is to first load the data of the low-rank tensor corresponding to the block of the sparse tensor into the local cache of the GPU core, and then load the non-zero elements of the block one by one for computation.

[0004] However, the sparsity of sparse tensors varies widely, and the GPU cannot ensure the cache data reuse rate when processing tensor multiplications with different sparsities, seriously affecting the running performance. Summary of the Invention

[0005] Embodiments of the present disclosure provide a scheduling method for computing tasks, a preprocessing method for computing tasks, a processing method for computing tasks, a controller, a processing unit, an electronic device, a preprocessing device, and a computer-readable medium.

[0006] In a first aspect, embodiments of the present disclosure provide a scheduling method for computing tasks, including:

[0007] Dividing at least one processing unit group, each of the processing unit groups including a plurality of processing units;

[0008] Mapping at least one block of a first tensor to at least one of the processing unit groups, so that the plurality of processing units in the processing unit group process the computing tasks corresponding to the block; different ones of the processing unit groups correspond to different ones of the blocks;

[0009] Wherein, the processing unit has a cache, and the plurality of processing units in the same processing unit group can transfer data to each other.

[0010] In some embodiments, the step of mapping at least one block of a first tensor to at least one of the processing unit groups includes:

[0011] Determine the target processing unit group, where the target processing unit group is the processing unit group corresponding to the block;

[0012] Load the data of the second tensor corresponding to the block into the caches of multiple processing units in the target processing unit group;

[0013] Inject the non-zero elements of the block into the processing units in the target processing unit group.

[0014] In some embodiments, the second tensor is obtained by contracting multiple factor tensors, and the data of the second tensor corresponding to the block is the data segments corresponding to the block in each of the factor tensors that are contracted to obtain the second tensor. Each data segment includes multiple sub-segments, and each non-zero element corresponds to one sub-segment of each data segment. The step of loading the data of the second tensor corresponding to the block into the caches of the processing units in the target processing unit group includes:

[0015] Load the multiple data segments corresponding to the block into the caches of multiple processing units in the target processing unit group.

[0016] In some embodiments, the step of loading the multiple data segments corresponding to the block into the caches of multiple processing units in the target processing unit group includes:

[0017] Determine the processing unit corresponding to each sub-segment according to a predetermined rule;

[0018] Load the sub-segment into the cache of the processing unit corresponding to the sub-segment.

[0019] In some embodiments, the predetermined rule is the algebraic relationship between the index of the sub-segment, the index of the non-zero element corresponding to the sub-segment, and the numbers of the processing units in the processing unit group. The step of determining multiple target processing units according to the predetermined rule includes:

[0020] Determine a target number according to the index of the sub-segment, the index of the non-zero element corresponding to the sub-segment, and the algebraic relationship;

[0021] Determine the processing unit with the number of the target number as the processing unit corresponding to the sub-segment.

[0022] In some embodiments, the step of injecting the non-zero elements of the block into the processing units in the target processing unit group includes:

[0023] Inject at least one non-zero element to be injected into a processing unit storing the sub-fragment corresponding to the non-zero element to be injected; wherein the non-zero element to be injected is one of the non-zero elements in the block.

[0024] In some embodiments, multiple sub-fragments corresponding to the non-zero element to be injected correspond one-to-one to consecutive computational stages of a computational task corresponding to the non-zero element to be injected; sub-fragments corresponding to two consecutive computational stages are stored in caches of different processing units; multiple processing units storing multiple sub-fragments corresponding to the non-zero element to be injected in the processing unit group can sequentially process consecutive computational stages of the computational task corresponding to the non-zero element to be injected.

[0025] In some embodiments, the step of injecting at least one non-zero element to be injected into a processing unit storing the sub-fragment corresponding to the non-zero element to be injected includes:

[0026] Inject multiple non-zero elements to be injected into multiple processing units simultaneously; wherein sub-fragments of the same factor tensor corresponding to multiple non-zero elements to be injected are stored in caches of different processing units.

[0027] In some embodiments, the step of dividing a processing unit group includes:

[0028] Divide a target number of processing units into one processing unit group to form the processing unit group, where the target number is determined according to the block size and the cache size of the processing unit.

[0029] In some embodiments, the step of dividing a target number of processing units into one processing unit group to form the processing unit group includes:

[0030] Divide consecutive target numbers of processing units into one processing unit group to form the processing unit group.

[0031] In some embodiments, the step of dividing consecutive target numbers of processing units into one processing unit group to form the processing unit group includes:

[0032] Divide a target number of processing units forming a rectangular topology into one processing unit group to form the processing unit group.

[0033] In a second aspect, an embodiment of the present disclosure provides a preprocessing method for a computational task, including:

[0034] Divide the first tensor into at least one block according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor;

[0035] Determine a target number according to the size of the divided block and the cache size of the processing unit, where the target number is the number of processing units that form a processing unit group;

[0036] Wherein, multiple processing units in the same processing unit group can transmit data to each other, and the processing unit group is used to process the computing task corresponding to the block.

[0037] In some embodiments, the step of dividing the first tensor into multiple blocks according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor includes:

[0038] Determine the size of the block according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor;

[0039] Divide the first tensor into at least one block according to the determined size of the block.

[0040] In some embodiments, the step of determining the target number according to the size of the block and the cache size of the processing unit includes:

[0041] Determine the storage space size required to store the data of the second tensor corresponding to the block according to the size of the block;

[0042] Determine the target number according to the determined storage space size required to store the data of the second tensor corresponding to the block and the cache size of the processing unit.

[0043] In a third aspect, an embodiment of the present disclosure provides a method for processing a computing task, which is applied to a processing unit and includes:

[0044] Receive task data, where the task data corresponds to non-zero elements of a block of the first tensor; the block corresponds to a processing unit group to which the current processing unit belongs, and the processing unit group includes multiple processing units, and the processing unit group is used to process the computing task corresponding to the block;

[0045] Read the data of the second tensor corresponding to the non-zero elements from the cache according to the task data;

[0046] Perform a calculation according to the task data and the data of the second tensor;

[0047] Transmit the calculation result to a first target processing unit, where the first target processing unit is one of the multiple processing units in the processing unit group.

[0048] In some embodiments, the second tensor is obtained by contracting a plurality of factor tensors, the blocks respectively correspond to data segments in each of the factor tensors that contract to obtain the second tensor, and each of the data segments includes a plurality of sub-segments; the non-zero elements respectively correspond to one sub-segment of each of the data segments; at least one sub-segment corresponding to the non-zero element is stored in the cache of the current processing unit; the step of reading, from the cache, data of the second tensor corresponding to the non-zero element according to the task data includes:

[0049] Determine a target sub-segment according to the task data, where the target sub-segment is the sub-segment corresponding to the task data among at least one sub-segment corresponding to the non-zero element stored in the cache;

[0050] Read the target sub-segment from the cache, where the target sub-segment is the data of the second tensor corresponding to the non-zero element.

[0051] In some embodiments, before the step of transmitting the calculation result to the first target processing unit, the processing method further includes:

[0052] Determine the first target processing unit according to a predetermined rule; the first target processing unit stores a second target sub-segment, where the second target sub-segment is one of the plurality of sub-segments corresponding to the non-zero element.

[0053] In some embodiments, the predetermined rule is an algebraic relationship between the index of the sub-segment, the index of the non-zero element corresponding to the sub-segment, and the number of the processing unit in the processing unit grouping; the step of determining the first target processing unit according to the predetermined rule includes:

[0054] Determine a target number according to the index of the second target sub-segment, the index of the non-zero element, and the algebraic relationship;

[0055] Determine the processing unit with the number being the target number as the first target processing unit.

[0056] In some embodiments, before the step of transmitting the calculation result to the first target processing unit, the processing method further includes:

[0057] Judge whether the calculation result is the final result;

[0058] In the case where the calculation result is the final result, transmit the calculation result out of the chip;

[0059] In the case where the calculation result is not the final result, execute the step of transmitting the calculation result to the first target processing unit.

[0060] In some embodiments, the step of receiving task data includes:

[0061] Receive the task data injected into the current processing unit, where the task data is the non-zero element.

[0062] In some embodiments, the step of receiving task data includes:

[0063] Receive the task data from a second target processing unit, where the task data is the calculation result sent by the second target processing unit; wherein, the second target processing unit is one of multiple processing units in the processing unit group.

[0064] Fourthly, an embodiment of the present disclosure provides a controller, including:

[0065] One or more processing modules;

[0066] A storage module, on which one or more programs are stored. When the one or more programs are executed by the one or more processing modules, the one or more processing modules implement the scheduling method of the computing task described in the first aspect of the embodiment of the present disclosure; and / or

[0067] The preprocessing method of the computing task described in the second aspect of the embodiment of the present disclosure.

[0068] Fifthly, an embodiment of the present disclosure provides a processing unit, including a computing unit and a cache;

[0069] The computing unit can read data from the cache to implement the processing method of the computing task described in the third aspect of the embodiment of the present disclosure.

[0070] Sixthly, an embodiment of the present disclosure provides an electronic device, including a controller and multiple processing units;

[0071] The controller is the controller described in the fourth aspect of the embodiment of the present disclosure;

[0072] The processing unit is the processing unit described in the fifth aspect of the embodiment of the present disclosure;

[0073] Wherein, multiple said processing units can transmit data to each other.

[0074] Seventhly, an embodiment of the present disclosure provides a preprocessing device, including:

[0075] One or more processing modules;

[0076] A storage module, on which one or more programs are stored. When the one or more programs are executed by the one or more processing modules, the one or more processing modules implement the preprocessing method of the computing task described in the second aspect of the embodiment of the present disclosure.

[0077] In an eighth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, it implements the preprocessing method for computing tasks described in the second aspect of the embodiments of the present disclosure.

[0078] In the preprocessing method for computing tasks provided by the embodiments of the present disclosure, an input tensor is divided into at least one block according to the data dimension and sparsity of the input tensor, and the number of processing units that make up a processing unit group is determined, so that a chip or system can divide the processing unit group according to the determined number. The total cache of multiple processing units in the processing unit group is adapted to the sparsity of the input tensor, and multiple processing units in the processing unit group can transmit data to each other, so that the processing unit group does not need to read data from the global cache when processing computing tasks. In the case where the sparsity change range of the input tensor is very large, a satisfactory cache data reuse rate can be achieved, cache resource waste can be avoided, and the computing performance can be improved.

[0079] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure, and do not constitute a limitation to the present disclosure. By describing the detailed exemplary embodiments with reference to the drawings, the above and other features and advantages will become more obvious to those skilled in the art. In the drawings:

[0081] Figure 1 is a flowchart of a scheduling method in an embodiment of the present disclosure;

[0082] Figure 2 is a schematic diagram of mapping multiple blocks to multiple processing unit groups;

[0083] Figure 3 is a flowchart of some steps in another scheduling method in an embodiment of the present disclosure;

[0084] Figure 4 is a schematic diagram of representing a tensor with a tensor network;

[0085] Figure 5 is a schematic diagram of a tensor calculation scenario;

[0086] Figure 6 is a flowchart of some steps in yet another scheduling method in an embodiment of the present disclosure;

[0087] Figure 7It is a flowchart of some steps in another scheduling method in the embodiments of the present disclosure;

[0088] Figure 8 It is a flowchart of some steps in another scheduling method in the embodiments of the present disclosure;

[0089] Figure 9 It is a flowchart of some steps in another scheduling method in the embodiments of the present disclosure;

[0090] Figure 10 It is a flowchart of some steps in another scheduling method in the embodiments of the present disclosure;

[0091] Figure 11 It is a flowchart of some steps in another scheduling method in the embodiments of the present disclosure;

[0092] Figure 12 It is a flowchart of a preprocessing method in the embodiments of the present disclosure;

[0093] Figure 13 It is a flowchart of some steps in another preprocessing method in the embodiments of the present disclosure;

[0094] Figure 14 It is a flowchart of a processing method in the embodiments of the present disclosure;

[0095] Figure 15 It is a flowchart of some steps in another processing method in the embodiments of the present disclosure;

[0096] Figure 16 It is a flowchart of some steps in yet another processing method in the embodiments of the present disclosure;

[0097] Figure 17 It is a flowchart of some steps in another processing method in the embodiments of the present disclosure;

[0098] Figure 18 It is a block diagram of the composition of a controller in the embodiments of the present disclosure;

[0099] Figure 19 It is a block diagram of the composition of a processing unit in the embodiments of the present disclosure;

[0100] Figure 20 It is a block diagram of the composition of an electronic device in the embodiments of the present disclosure;

[0101] Figure 21 It is a block diagram of the composition of a preprocessing device in the embodiments of the present disclosure. Detailed implementation manners

[0102] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the following provides descriptions of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0103] Without conflict, the various embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0104] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0105] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms "comprises" and / or "consists of" are used in this specification, it specifies the presence of the stated features, wholes, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups. "Connection" or "connected" and other similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0106] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless clearly defined herein.

[0107] Through research by the inventors of the present disclosure, it is found that when dealing with the multiplication of sparse tensors and low-rank tensors, the sparser the sparse tensor, the larger the storage space required to store the low-rank tensor to achieve a satisfactory cache data reuse rate. In some related technologies, the local cache capacity of the GPU core is fixed. When the non-zero element density of the sparse tensor is relatively large and the local cache capacity of the GPU core is relatively large, it is easy to cause waste of cache resources; when the non-zero element density of the sparse tensor is relatively small and the local cache capacity of the GPU core is relatively small, the cache data reuse rate will be reduced. Although the global cache of the GPU is large, the access latency is also large and it cannot be used as a local cache.

[0108] In view of this, in a first aspect, embodiments of the present disclosure provide a scheduling method for computing tasks, referring to Figure 1 , the scheduling method includes:

[0109] In step S100, at least one processing unit group is divided, and each processing unit group includes a plurality of processing units;

[0110] In step S200, at least one block of the first tensor is mapped to at least one of the processing unit groups, so that the plurality of processing units in the processing unit group process the computing tasks corresponding to the blocks; different processing unit groups correspond to different blocks; wherein, the processing unit has a cache, and the plurality of processing units in the same processing unit group can transmit data to each other.

[0111] In embodiments of the present disclosure, a plurality of processing elements (PEs) form a processing element array. In step S100, the processing units in the processing element array are divided into at least one processing unit group. In embodiments of the present disclosure, the processing element array can be a GPU, or a chip with a plurality of processing units, or a many-core system composed of chips with a plurality of processing units. Embodiments of the present disclosure do not make special limitations on this.

[0112] In embodiments of the present disclosure, the first tensor is an input tensor, and the computing task is the multiplication of the first tensor and the second tensor. When dividing the processing unit group in step S100, the number of processing units in the processing unit group is adapted to the sparsity of the first tensor. The sparser the first tensor, the more processing units that make up the processing unit group. In embodiments of the present disclosure, the processing unit has a cache, and the overall cache formed by the caches of the respective processing units in the processing unit group serves as the cache of the processing unit group. It should be noted that, relative to the global cache of the chip or system, the cache of the processing unit is a local cache. The number of processing units in the processing unit group is adapted to the sparsity of the first tensor, that is, the cache of the processing unit group can be dynamically adjusted according to the sparsity of the first tensor.

[0113] In embodiments of the present disclosure, no special limitations are made on the data dimensions of the first tensor and the second tensor. For example, the data dimension of the first tensor can be three-dimensional or more than three-dimensional. For example, the first tensor can be a three-dimensional tensor, a five-dimensional tensor, a ten-dimensional tensor; the data dimension of the second tensor can be three-dimensional or more than three-dimensional. For example, the second tensor can be a three-dimensional tensor, a five-dimensional tensor, a ten-dimensional tensor.

[0114] It should be noted that, in embodiments of the present disclosure, the number of processing units in the processing unit group can be input externally to the chip or system, or can be generated internally by the chip or system. Embodiments of the present disclosure do not make special limitations on this.

[0115] In the embodiment of the present disclosure, mapping the block (Tilling Box) of the first tensor to the processing unit group in step S200 includes loading the data of the second tensor corresponding to the block of the first tensor into the cache of the processing unit group, that is, the data of the second tensor is distributed in the caches of the respective processing units in the processing unit group; it also includes injecting the non-zero elements of the block of the first tensor into the processing units in the processing unit group. It should be noted that in the embodiment of the present disclosure, the non-zero elements represent the meaningful data in the tensor, and the corresponding zero elements represent the meaningless data, which is not limited to the numerical value 0.

[0116] In the embodiment of the present disclosure, the overall cache composed of the caches of the respective processing units in the processing unit group serves as the cache of the processing unit group, which is achieved by enabling multiple processing units in the same processing unit group to transfer data to each other. In the embodiment of the present disclosure, the computing tasks involved in the same non-zero element are completed by multiple processing units in the processing unit group. Multiple processing units in the same processing unit group can transfer intermediate calculation results to each other, and each processing unit can continue to perform calculations based on the received intermediate calculation results. Therefore, when the data of the second tensor is distributed in the caches of the respective processing units in the processing unit group in step S200, each processing unit reads data from its own cache without reading data from the global cache, and there is no need to transfer data between the processing units through the global cache either.

[0117] In the scheduling method of the computing task provided in the embodiment of the present disclosure, the processing unit group is divided according to the sparsity of the input tensor. The total cache of multiple processing units in the processing unit group is adapted to the sparsity of the input tensor. Multiple processing units in the processing unit group can transfer data to each other, so that the processing unit group does not need to read data from the global cache when processing the computing task. In the case where the sparsity change range of the input tensor is very large, a satisfactory cache data reuse rate can be achieved, cache resource waste can be avoided, and the operation performance can be improved.

[0118] In the embodiment of the present disclosure, the first tensor can be divided into at least one block. In step S100, one processing unit group can be divided, and the computing tasks involved in each block of the first tensor are processed by this processing unit group; or multiple processing unit groups can be divided, and each block of the first tensor is dynamically mapped to multiple processing unit groups, with each block corresponding to one processing unit group. Figure 2 FIG. is a schematic diagram for mapping multiple blocks to multiple processing unit groups. In Figure 2 it, every 4 processing units form a 2x2 processing unit group, Figure 2It shows a total of 4 processing unit groups, namely PE Group 1, PE Group 2, PE Group 3, and PE Group 4. Among them, block Box 35 is mapped to PE Group 1, block Box 36 is mapped to PE Group 2, block Box 33 is mapped to PE Group 3, and block Box 37 is mapped to PE Group 1.

[0119] In the embodiments of the present disclosure, after determining the processing unit group corresponding to the block, mapping the block of the first tensor to the processing unit group includes: loading the data of the second tensor corresponding to the block of the first tensor into the cache of the processing unit group; injecting the non-zero elements of the block of the first tensor into the processing units in the processing unit group.

[0120] Correspondingly, in some embodiments, referring to Figure 3 , step S200 includes:

[0121] In step S210, determine the target processing unit group, where the target processing unit group is the processing unit group corresponding to the block;

[0122] In step S220, load the data of the second tensor corresponding to the block into the caches of multiple processing units in the target processing unit group;

[0123] In step S230, inject the non-zero elements of the block into the processing units in the target processing unit group.

[0124] It should be noted that after executing step S220, the data of the second tensor is distributed in the caches of each processing unit in the processing unit group.

[0125] It should also be noted that when the sparsity of the first tensor is determined, each block of the first tensor has a similar sparsity. Therefore, when the computing tasks involved in different blocks are processed by the same processing unit group, it is also possible to ensure a satisfactory cache data reuse rate.

[0126] In the embodiments of the present disclosure, the second tensor is a low-rank tensor. As Figure 4 shown, the low-rank tensor W can be represented by a tensor network. Multiple factor tensors A, B, and C in the tensor network are contracted according to the connection relationship to obtain the low-rank tensor, where the data dimensions of the factor tensors A, B, and C are lower than the data dimension of the low-rank tensor W. Figure 4 Taking the low-rank tensor as a three-dimensional tensor as an example for illustration. In the embodiments of the present disclosure, the second tensor is not limited to a three-dimensional tensor.

[0127] In an embodiment of the present disclosure, when performing the multiplication of a first tensor and a second tensor, according to the operation rules of tensor multiplication, the blocks of the first tensor are correspondingly contracted to obtain a data segment in one of the factor tensors of the second tensor; in each factor tensor, the non-zero elements of the block correspond to a sub-segment in the data segment. It should be noted that the data segment refers to a set composed of partial elements of the factor tensor. As Figure 5 shown, the second tensor W is obtained by contracting the factor tensor A, the factor tensor B, and the factor tensor C. The block of the first tensor X corresponds to the data segment A n×r of the factor tensor A, the data segment B n×r of the factor tensor B, and the data segment C n×r of the factor tensor C. The non-zero element x 13,4,7 of the first tensor X corresponds to the sub-segments A(13, :), B(4, :), and C(7, :). In an embodiment of the present disclosure, the data segments of the factor tensors obtained by contracting the corresponding blocks are stored in the cache of the processing unit.

[0128] Correspondingly, in some embodiments, the second tensor is obtained by contracting multiple factor tensors. The data of the second tensor corresponding to the block is the data segment corresponding to the block in each of the factor tensors obtained by contracting the second tensor. Each of the data segments includes multiple sub-segments, and each non-zero element corresponds to one of the sub-segments of each of the data segments; referring to Figure 6 , step S220 includes:

[0129] In step S221, the multiple data segments corresponding to the block are loaded into the caches of multiple processing units in the target processing unit group.

[0130] In an embodiment of the present disclosure, the sub-segments of the same factor tensor may be stored in the cache of the same processing unit or in the caches of multiple processing units. The present disclosure does not make special limitations on this.

[0131] In an embodiment of the present disclosure, for multiple non-zero elements of the block, the multiple sub-segments corresponding to the same non-zero element may be stored in the cache of the same processing unit or in the caches of multiple processing units. In the scenario where the multiple sub-segments corresponding to the same non-zero element are stored in the caches of multiple processing units, the multiple sub-segments may be respectively stored in the caches of different processing units, or some sub-segments may be stored in the cache of the same processing unit. The present disclosure does not make special limitations on this.

[0132] It should be noted that when multiple factor tensors are contracted to obtain a second tensor, multiplying the first tensor by the second tensor is transformed into multiplying the first tensor by each of the factor tensors that are contracted to obtain the second tensor one by one. Multiple sub-fragments corresponding to the same non-zero element in the block are stored in the caches of multiple processing units. That is, the sub-fragments of each factor tensor corresponding to the same non-zero element are not stored in the cache of the same processing unit, which enables multiple processing units in the processing unit group to complete the calculation of multiplying the first tensor by each of the factor tensors that are contracted to obtain the second tensor one by one through relay. Each processing unit reads the corresponding sub-fragment from its own cache.

[0133] In the embodiments of the present disclosure, during the process of the processing units in the processing unit group processing the calculation of multiplying the first tensor by each of the factor tensors that are contracted to obtain the second tensor one by one, each factor tensor corresponds to a calculation stage. After the processing unit reads the sub-fragment corresponding to the current calculation stage in the cache and completes the calculation of the current calculation stage, it is necessary to transmit the calculation result to the processing unit storing the sub-fragment corresponding to the next calculation stage. In the embodiments of the present disclosure, the storage locations of each sub-fragment can be stored in each processing unit, so that the processing unit can determine the processing unit for executing the next calculation stage; or multiple sub-fragments can be loaded into the caches of multiple processing units according to a predetermined rule, and the processing unit can determine the processing unit for executing the next calculation stage according to the predetermined rule, thereby saving the storage resources of the processing unit.

[0134] Correspondingly, in some embodiments, referring to Figure 7 , step S221 includes:

[0135] In step S2211, determine the processing unit corresponding to each of the sub-fragments according to a predetermined rule;

[0136] In step S2212, load the sub-fragment into the cache of the processing unit corresponding to the sub-fragment.

[0137] The embodiments of the present disclosure do not make special limitations on the predetermined rule described in step S2211. As an optional implementation manner, the processing units in the processing unit group have numbers for identifying each processing unit; the predetermined rule is the algebraic relationship between the index of the sub-fragment, the index of the non-zero element corresponding to the sub-fragment, and the number of the processing unit in the processing unit group; referring to Figure 8 , step S2211 includes:

[0138] In step S2211a, determine the target number according to the index of the sub-fragment, the index of the non-zero element corresponding to the sub-fragment, and the algebraic relationship;

[0139] In step S2211b, determine the processing unit with the target number as the processing unit corresponding to the sub-fragment.

[0140] The embodiments of the present disclosure do not make special limitations on the algebraic relationship between the index of the sub - fragment, the index of the non - zero element corresponding to the sub - fragment, and the number of the processing units in the processing unit group. For example, the algebraic relationship can be taking the remainder, taking the integer, etc., so as to save the computing resources of the processing units.

[0141] In the embodiments of the present disclosure, the processing unit group can process the computing tasks involved in each non - zero element of the block in a serial manner, that is, injecting each non - zero element into the processing units in the processing unit group one by one; or can process the computing tasks involved in each non - zero element of the block in a parallel manner, that is, injecting multiple non - zero elements into the processing units in the processing unit group simultaneously. The embodiments of the present disclosure do not make special limitations on this. It should be noted that in the cache of the processing unit into which the non - zero element is injected, the sub - fragment corresponding to the non - zero element is stored.

[0142] Correspondingly, in some embodiments, referring to Figure 9 , step S230 includes:

[0143] In step S231, injecting at least one non - zero element to be injected into the processing unit storing the sub - fragment corresponding to the non - zero element to be injected; wherein, the non - zero element to be injected is one of the non - zero elements in the block.

[0144] In the embodiments of the present disclosure, when the processing unit group processes the computing tasks involved in non - zero elements in a serial manner or in a parallel manner, multiple processing units in the processing unit group process the computing tasks involved in each non - zero element according to the traversal mechanism. It should be noted that in the embodiments of the present disclosure, the traversal mechanism means that the computing tasks involved in non - zero elements can be divided into multiple consecutive computing stages, each processing unit in the processing unit group processes one computing stage among the multiple consecutive computing stages, multiple processing units in the processing unit group transmit intermediate computing results through a high - speed on - chip network, and the processing unit continues to process its corresponding computing stage based on the intermediate computing results calculated by the previous processing unit, so as to realize the traversal of the computing tasks involved in non - zero elements in the processing unit group and finally obtain the final result of the computing tasks involved in non - zero elements.

[0145] Correspondingly, in some embodiments, multiple sub - fragments corresponding to the non - zero element to be injected correspond one - to - one with multiple consecutive computing stages of the computing tasks corresponding to the non - zero element to be injected; sub - fragments corresponding to two consecutive computing stages are stored in the caches of different processing units; multiple processing units in the processing unit group storing multiple sub - fragments corresponding to the non - zero element to be injected can continue to process multiple consecutive computing stages of the computing tasks corresponding to the non - zero element to be injected.

[0146] It should be noted that in the embodiments of the present disclosure, the successive processing means that the processing unit receives the intermediate calculation result obtained by the previous processing unit, and processes its corresponding calculation stage based on the intermediate calculation result until all successive calculation stages of the calculation tasks involved in the non-zero elements are completed by multiple processing units in the processing unit group.

[0147] Correspondingly, in some embodiments, referring to Figure 10 , step S231 includes:

[0148] In step S2311, a plurality of the to-be-injected non-zero elements are simultaneously injected into a plurality of processing units respectively; wherein, sub-fragments of the same factor tensor corresponding to the plurality of to-be-injected non-zero elements are stored in the caches of different processing units.

[0149] It should be noted that in the embodiments of the present disclosure, the first tensor is multiplied by each of the factor tensors obtained by contraction one by one in a predetermined order. In step S2311, a plurality of non-zero elements are simultaneously injected into different processing units, and the sub-fragments of the same factor tensor corresponding to the simultaneously injected non-zero elements are stored in the caches of different processing units, which can ensure parallel performance.

[0150] In the embodiments of the present disclosure, the number of processing units in the processing unit group can be input from outside the chip or the system, or can be generated inside the chip or the system. The embodiments of the present disclosure do not make special limitations on this.

[0151] Correspondingly, in some embodiments, referring to Figure 11 , step S100 includes:

[0152] In step S110, a target number of processing units are divided into one of the processing unit groups to form the processing unit group, and the target number is determined according to the block size and the cache size of the processing unit.

[0153] It should be noted that the target number is transmitted by a device outside the chip or the system, or can be generated inside the chip or the system.

[0154] In the embodiments of the present disclosure, the multiple processing units forming the processing unit group can be continuous or discontinuous. The embodiments of the present disclosure do not make special limitations on this. It should be noted that in the embodiments of the present disclosure, the continuous processing units refer to the continuous processing units in terms of the data connection relationship. As an optional implementation manner, the multiple processing units forming the processing unit group are continuous target number of processing units, so as to reduce the latency of data transmission between processing units and further improve the operation performance.

[0155] Accordingly, in some embodiments, a continuous target number of processing units are divided into one of the processing unit groups to form the processing unit group.

[0156] As Figure 2 shown, in a scenario where the positional relationship of multiple processing units reflects a data connection relationship, the continuous processing units can also be processing units with a continuous topological relationship.

[0157] Further, in some embodiments, a target number of processing units that form a rectangular topology are divided into one processing unit group to form the processing unit group.

[0158] In a second aspect, an embodiment of the present disclosure provides a preprocessing method for a computing task. Referring to Figure 12 , the preprocessing method includes:

[0159] In step S300, according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor, the first tensor is divided into at least one block;

[0160] In step S400, according to the size of the block and the cache size of the processing unit, a target number is determined, where the target number is the number of the processing units that form one processing unit group; wherein, multiple processing units in the same processing unit group can transmit data to each other, and the processing unit group is used to process the computing task corresponding to the block.

[0161] In the embodiment of the present disclosure, the blocks obtained by dividing the first tensor with different data dimensions and different sparsities in step S300 can achieve a satisfactory cache data reuse rate, that is, the target cache data reuse rate. The target cache data reuse rates corresponding to different data dimensions and different sparsities can be the same or different, and the embodiment of the present disclosure does not make special limitations on this.

[0162] In the embodiment of the present disclosure, the processing unit has a cache, and the overall cache formed by the caches of the respective processing units in the processing unit group serves as the cache of the processing unit group. The number of processing units in the processing unit group is adapted to the sparsity of the first tensor, that is, the cache of the processing unit group can be dynamically adjusted according to the sparsity of the first tensor.

[0163] In the embodiment of the present disclosure, the overall cache formed by the caches of the respective processing units in the processing unit group serves as the cache of the processing unit group, which is achieved by enabling multiple processing units in the same processing unit group to transmit data to each other.

[0164] It should be noted that steps S300 to S400 can be executed inside the chip or system, or in a device external to the chip or system. The embodiments of the present disclosure do not make special limitations in this regard. In the scenario where steps S300 to S400 are executed in a device external to the chip or system, the preprocessing method further includes transferring each block to the inside of the chip or system, and transferring the target quantity to the inside of the chip or system.

[0165] In the preprocessing method of the computing task provided by the embodiments of the present disclosure, the input tensor is divided into at least one block according to the data dimension, sparsity, and target cache data reuse rate of the input tensor, and the number of processing units that make up the processing unit group is determined, so that the chip or system can divide the processing unit group according to the determined number, and the total cache of multiple processing units in the processing unit group is adapted to the sparsity of the input tensor. Therefore, in the case where the sparsity change range of the input tensor is very large, a satisfactory cache data reuse rate can be achieved, cache resource waste can be avoided, and the operation performance can be improved.

[0166] In some embodiments, referring to Figure 13 , step S300 includes:

[0167] In step S310, determine the size of the block according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor;

[0168] In step S320, divide the first tensor into at least one of the blocks according to the determined size of the block.

[0169] In some embodiments, referring to Figure 13 , step S400 includes:

[0170] In step S410, determine the storage space size required to store the data of the second tensor corresponding to the block according to the size of the block;

[0171] In step S420, determine the target quantity according to the determined storage space size required to store the data of the second tensor corresponding to the block and the cache size of the processing unit.

[0172] In a third aspect, the embodiments of the present disclosure provide a processing method for a computing task, which is applied to a processing unit. Referring to Figure 14 , the processing method includes:

[0173] In step S500, receive task data, where the task data corresponds to non-zero elements of a block of the first tensor; the block corresponds to a processing unit group to which the current processing unit belongs, and the processing unit group includes multiple processing units, and the processing unit group is used to process the computing task corresponding to the block;

[0174] In step S600, the data of the second tensor corresponding to the non-zero element is read from the cache according to the task data;

[0175] In step S700, calculations are performed according to the task data and the data of the second tensor;

[0176] In step S800, the calculation result is transmitted to a first target processing unit, where the first target processing unit is one of multiple processing units in the processing unit group.

[0177] In the embodiment of the present disclosure, the processing unit has a cache, and the overall cache formed by the caches of each processing unit in the processing unit group serves as the cache of the processing unit group. The number of processing units in the processing unit group is adapted to the sparsity of the first tensor, that is, the cache of the processing unit group can be dynamically adjusted according to the sparsity of the first tensor.

[0178] In the embodiment of the present disclosure, the computing tasks related to the same non-zero element are completed by multiple processing units in the processing unit group. Multiple processing units in the same processing unit group can transmit intermediate calculation results to each other, and each processing unit can continue to perform calculations based on the received intermediate calculation results. The task data can be a non-zero element or an intermediate calculation result, and the embodiment of the present disclosure does not make special limitations on this.

[0179] In the method for processing computing tasks provided in the embodiment of the present disclosure, multiple processing units in the processing unit group can transmit data to each other, realizing that the overall cache formed by the caches of each processing unit in the processing unit group serves as the cache of the processing unit group. Multiple processing units can relay to process the computing tasks related to the same non-zero element, and there is no need to read data from the global cache when processing computing tasks, so that the processing unit group can achieve a satisfactory cache data reuse rate in the case of a large change range of the sparsity of the input tensor, and can also avoid waste of cache resources and improve the operation performance.

[0180] In the embodiment of the present disclosure, the second tensor is a low-rank tensor. As Figure 4 shown, the low-rank tensor can be represented by a tensor network. Multiple factor tensors in the tensor network are contracted according to the connection relationship to obtain the low-rank tensor, where the data dimension of the factor tensor is lower than the data dimension of the low-rank tensor. Figure 4 Taking the low-rank tensor as a three-dimensional tensor as an example for illustration. In the embodiment of the present disclosure, the second tensor is not limited to a three-dimensional tensor.

[0181] In an embodiment of the present disclosure, when performing the multiplication of a first tensor and a second tensor, according to the operation rules of tensor multiplication, a block of the first tensor corresponds to a contraction to obtain a data segment in one of the factor tensors of the second tensor; in each factor tensor, the non-zero elements of the block correspond to a sub-segment in the data segment. It should be noted that the data segment refers to a set composed of partial elements of the factor tensor. As Figure 5 shown, the second tensor is obtained by contracting factor tensor A, factor tensor B, and factor tensor C. A block of the first tensor X corresponds to data segment A of factor tensor A n×r , data segment B of factor tensor B n×r , and data segment C of factor tensor C n×r . The non-zero element x of the first tensor X 13,4,7 corresponds to sub-segments A(13, :), B(4, :), and C(7, :). In an embodiment of the present disclosure, data segments of each factor tensor obtained by contracting the corresponding blocks are stored in the cache of the processing unit.

[0182] Correspondingly, in some embodiments, the second tensor is obtained by contracting multiple factor tensors. The block corresponds to a contraction to obtain a data segment in each of the factor tensors of the second tensor. Each data segment includes multiple sub-segments; the non-zero element corresponds to a sub-segment of each data segment; at least one sub-segment corresponding to the non-zero element is stored in the cache of the current processing unit; referring to Figure 15 , step S600 includes:

[0183] In step S610, determine a target sub-segment according to the task data. The target sub-segment is the sub-segment corresponding to the task data among the sub-segments corresponding to at least one non-zero element stored in the cache;

[0184] In step S620, read the target sub-segment from the cache. The target sub-segment is the data of the second tensor corresponding to the non-zero element.

[0185] In an embodiment of the present disclosure, the storage locations of each sub-segment can be stored in each processing unit, so that the processing unit can determine the processing unit for executing the next calculation stage; or multiple sub-segments can be loaded into the caches of multiple processing units according to a predetermined rule, and the processing unit can determine the processing unit for executing the next calculation stage according to the predetermined rule, thereby saving the storage resources of the processing unit.

[0186] Correspondingly, referring to Figure 16 , before step S800, the processing method further includes:

[0187] In step S910, determine the first target processing unit according to a predetermined rule; the first target processing unit stores a second target sub - fragment, and the second target sub - fragment is one of the multiple sub - fragments corresponding to the non - zero element.

[0188] As an alternative implementation, the processing units in the processing unit group have numbers for identifying each processing unit; the predetermined rule is the algebraic relationship among the index of the sub - fragment, the index of the non - zero element corresponding to the sub - fragment, and the number of the processing unit in the processing unit group; referring to Figure 17 , step S910 includes:

[0189] In step S911, determine a target number according to the index of the second target sub - fragment, the index of the non - zero element, and the algebraic relationship.

[0190] In step S912, determine the processing unit with the target number as the first target processing unit.

[0191] The embodiments of the present disclosure do not make special limitations on the algebraic relationship among the index of the sub - fragment, the index of the non - zero element corresponding to the sub - fragment, and the number of the processing unit in the processing unit group. For example, the algebraic relationship can be taking the remainder, taking the integer, etc., so as to save the computing resources of the processing unit.

[0192] In some embodiments, referring to Figure 16 , before step S800, the processing method further includes:

[0193] In step S920, determine whether the calculation result is the final result.

[0194] In step S930, when the calculation result is the final result, transmit the calculation result outside the chip; when the calculation result is not the final result, execute the step of transmitting the calculation result to the first target processing unit.

[0195] In the embodiments of the present disclosure, the task data received through step S500 can be a non - zero element or an intermediate calculation result, and the embodiments of the present disclosure do not make special limitations on this.

[0196] Correspondingly, in step S500, the task data injected into the current processing unit can be received, and the task data is the non - zero element.

[0197] Correspondingly, in step S500, the task data can also be received from a second target processing unit, and the task data is the calculation result sent by the second target processing unit; wherein, the second target processing unit is one of the multiple processing units in the processing unit group.

[0198] Fourth aspect, referring to Figure 18 , an embodiment of the present disclosure provides a controller, including:

[0199] One or more processing modules 101;

[0200] A storage module 102, on which one or more programs are stored. When the one or more programs are executed by the one or more processing modules 101, the one or more processing modules 101 implement the scheduling method of the computing tasks described in the first aspect of the embodiments of the present disclosure; and / or

[0201] The preprocessing method of the computing tasks described in the second aspect of the embodiments of the present disclosure.

[0202] Fifth aspect, referring to Figure 19 , an embodiment of the present disclosure provides a processing unit, including a computing unit 201 and a cache 202;

[0203] The computing unit 201 can read data from the cache 202 to implement the processing method of the computing tasks described in the third aspect of the embodiments of the present disclosure.

[0204] Sixth aspect, referring to Figure 20 , an embodiment of the present disclosure provides an electronic device, including a controller 100 and a plurality of processing units 200;

[0205] The controller 100 is the controller described in the fourth aspect of the embodiments of the present disclosure;

[0206] The processing unit 200 is the processing unit described in the fifth aspect of the embodiments of the present disclosure;

[0207] Wherein, the plurality of processing units 200 can transmit data to each other.

[0208] Seventh aspect, referring to Figure 21 , an embodiment of the present disclosure provides a preprocessing device, including:

[0209] One or more processing modules 301;

[0210] A storage module 302, on which one or more programs are stored. When the one or more programs are executed by the one or more processing modules 301, the one or more processing modules 301 implement the preprocessing method of the computing tasks described in the second aspect of the embodiments of the present disclosure.

[0211] Eighth aspect, an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, the preprocessing method of the computing tasks described in the second aspect of the embodiments of the present disclosure is implemented.

[0212] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components in cooperation. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0213] Example embodiments have been disclosed herein, and although specific terms have been employed, they are used only for and should be construed only as general illustrative meanings and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, features, characteristics, and / or elements described in connection with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in connection with other embodiments. Accordingly, those skilled in the art will understand that various forms and details may be changed without departing from the scope of the present disclosure as set forth by the appended claims.

Claims

1. A scheduling method for computing tasks, comprising: Dividing at least one processing unit group according to the sparsity of the first tensor, each processing unit group includes a plurality of processing units, the number of processing units in the processing unit group is adapted to the sparsity, and the total cache of the plurality of processing units in the processing unit group is adapted to the sparsity; Mapping at least one block of the first tensor to at least one of the processing unit groups, so that the data of the second tensor corresponding to the block of the first tensor is distributed in the caches of each processing unit in the processing unit group, and the plurality of processing units in the processing unit group process the computing tasks corresponding to the block; different processing unit groups correspond to different blocks; Wherein, the processing unit has a cache, and the plurality of processing units in the same processing unit group can transfer data to each other.

2. The scheduling method according to claim 1, wherein, The step of mapping at least one block of the first tensor to at least one of the processing unit groups includes: Determining a target processing unit group, the target processing unit group being the processing unit group corresponding to the block; Loading the data of the second tensor corresponding to the block into the caches of the plurality of processing units in the target processing unit group; Injecting the non-zero elements of the block into the processing units in the target processing unit group.

3. The scheduling method according to claim 2, wherein, The second tensor is obtained by contracting a plurality of factor tensors, the data of the second tensor corresponding to the block is the data segments corresponding to the block in each of the factor tensors that are contracted to obtain the second tensor, each data segment includes a plurality of sub-segments, and each non-zero element corresponds to one of the sub-segments of each data segment; The step of loading the data of the second tensor corresponding to the block into the caches of the processing units in the target processing unit group includes: Loading the plurality of data segments corresponding to the block into the caches of the plurality of processing units in the target processing unit group.

4. The scheduling method according to claim 3, wherein, The step of loading the plurality of data segments corresponding to the block into the caches of the plurality of processing units in the target processing unit group includes: Determining the processing unit corresponding to each sub-segment according to a predetermined rule; Loading the sub-segment into the cache of the processing unit corresponding to the sub-segment.

5. The scheduling method according to claim 4, wherein, The predetermined rule is the algebraic relationship between the index of the sub-segment, the index of the non-zero element corresponding to the sub-segment, and the number of the processing units in the processing unit group; The step of determining a plurality of target processing units according to a predetermined rule includes: Determining a target number according to the index of the sub-segment, the index of the non-zero element corresponding to the sub-segment, and the algebraic relationship; Determining the processing unit with the target number as the processing unit corresponding to the sub-segment.

6. The scheduling method according to any one of claims 3 to 5, wherein, The step of injecting the non-zero elements of the block into the processing units in the target processing unit group includes: Injecting at least one non-zero element to be injected into the processing unit storing the sub-segment corresponding to the non-zero element to be injected; wherein, the non-zero element to be injected is one of the non-zero elements in the block.

7. The scheduling method according to claim 6, wherein, The multiple sub - segments corresponding to the non - zero element to be injected correspond one - to - one with consecutive multiple computational stages of the computational task corresponding to the non - zero element to be injected; The sub - segments corresponding to two consecutive computational stages are stored in the caches of different processing units; Among the processing units in the processing unit group that store the multiple sub - segments corresponding to the non - zero element to be injected, these processing units can successively process consecutive multiple computational stages of the computational task corresponding to the non - zero element to be injected.

8. The scheduling method according to claim 7, wherein, The step of injecting at least one non - zero element to be injected into the processing unit that stores the sub - segment corresponding to the non - zero element to be injected includes: Simultaneously injecting multiple non - zero elements to be injected into multiple processing units respectively; wherein, the sub - segments of the same factor tensor corresponding to the multiple non - zero elements to be injected are stored in the caches of different processing units.

9. The scheduling method according to any one of claims 1 to 5, wherein, The step of dividing the processing unit group includes: Dividing a target number of processing units into one processing unit group to form the processing unit group, and the target number is determined according to the block size and the cache size of the processing unit.

10. The scheduling method according to claim 9, wherein, The step of dividing a target number of processing units into one processing unit group to form the processing unit group includes: Dividing consecutive target - number processing units into one processing unit group to form the processing unit group.

11. The scheduling method according to claim 10, wherein, The step of dividing consecutive target - number processing units into one processing unit group to form the processing unit group includes: Dividing a target number of processing units forming a rectangular topology into one processing unit group to form the processing unit group.

12. A pre - processing method for a computational task, including: Dividing the first tensor into at least one block according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor; Determining a target number according to the size of the block and the cache size of the processing unit, where the target number is the number of processing units that form one processing unit group, the number of processing units in the processing unit group is adapted to the sparsity, and the total cache of multiple processing units in the processing unit group is adapted to the sparsity; Among them, multiple processing units that form the same processing unit group can transmit data to each other, and the processing unit group is used to process the computational task corresponding to the block; The step of determining the target number according to the size of the block and the cache size of the processing unit includes: determining the storage space size required to store the data of the second tensor corresponding to the block according to the size of the block; determining the target number according to the determined storage space size required to store the data of the second tensor corresponding to the block and the cache size of the processing unit.

13. The preprocessing method according to claim 12, wherein, The step of dividing the first tensor into at least one block according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor includes: Determining the size of the block according to the data dimension of the first tensor, the target cache data reuse rate, and the sparsity of the first tensor; Dividing the first tensor into at least one such block according to the determined size of the block.

14. A method for processing a computing task, applied to a processing unit, comprising: Receiving task data, where the task data corresponds to non-zero elements of a block of a first tensor; The block corresponds to a processing unit group to which the current processing unit belongs. The processing unit group includes multiple processing units. The number of processing units in the processing unit group is adapted to the sparsity of the first tensor, and the total cache of the multiple processing units in the processing unit group is adapted to the sparsity. The processing unit group is used to process the computing task corresponding to the block; Reading, from the cache, data of a second tensor corresponding to the non-zero elements according to the task data; Performing a calculation according to the task data and the data of the second tensor; Transmitting the calculation result to a first target processing unit, where the first target processing unit is one of the multiple processing units in the processing unit group.

15. The processing method according to claim 14, wherein, The second tensor is obtained by contracting multiple factor tensors. The block corresponds to data segments in each of the factor tensors that are contracted to obtain the second tensor. Each of the data segments includes multiple sub-segments; the non-zero elements correspond to one of the sub-segments of each of the data segments; at least one sub-segment corresponding to the non-zero elements is stored in the cache of the current processing unit; The step of reading, from the cache, data of the second tensor corresponding to the non-zero elements according to the task data includes: Determining a target sub-segment according to the task data, where the target sub-segment is the sub-segment corresponding to the task data among at least one sub-segment corresponding to the non-zero elements stored in the cache; Reading the target sub-segment from the cache, where the target sub-segment is the data of the second tensor corresponding to the non-zero elements.

16. The processing method according to claim 15, wherein, Before the step of transmitting the calculation result to the first target processing unit, the processing method further includes: Determining the first target processing unit according to a predetermined rule; the first target processing unit stores a second target sub-segment, where the second target sub-segment is one of the multiple sub-segments corresponding to the non-zero elements.

17. The processing method according to claim 16, wherein, The predetermined rule is an algebraic relationship among the index of the sub-segment, the index of the non-zero element corresponding to the sub-segment, and the number of the processing unit in the processing unit group; The step of determining the first target processing unit according to a predetermined rule includes: Determining a target number according to the index of the second target sub-segment, the index of the non-zero element, and the algebraic relationship; Determining the processing unit with the number of the target number as the first target processing unit.

18. The processing method according to any one of claims 14 to 17, wherein, Before the step of transmitting the calculation result to the first target processing unit, the processing method further includes: Judging whether the calculation result is the final result; In the case where the calculation result is the final result, transmitting the calculation result outside the chip; In the case where the calculation result is not the final result, performing the step of transmitting the calculation result to the first target processing unit.

19. The processing method according to any one of claims 14 to 17, wherein, The step of receiving task data includes: Receiving the task data injected into the current processing unit, where the task data is the non-zero element.

20. The processing method according to any one of claims 14 to 17, wherein, The step of receiving task data includes: Receiving the task data from a second target processing unit, where the task data is a calculation result sent by the second target processing unit; wherein the second target processing unit is one of multiple processing units in the processing unit group.

21. A controller, comprising: One or more processing modules; A storage module storing one or more programs, which when executed by the one or more processing modules, cause the one or more processing modules to implement the scheduling method for the computing task according to any one of claims 1 to 11; and / or The preprocessing method for the computing task according to any one of claims 12 to 13.

22. A processing unit, comprising a computing unit and a cache; The computing unit is capable of reading data from the cache to implement the processing method for the computing task according to any one of claims 14 to 20.

23. An electronic device, comprising a controller and multiple processing units; The controller is the controller according to claim 21; The processing unit is the processing unit according to claim 22; Among them, Multiple of the processing units are capable of transmitting data to each other.

24. A preprocessing device, comprising: One or more processing modules; A storage module storing one or more programs, which when executed by the one or more processing modules, cause the one or more processing modules to implement the preprocessing method for the computing task according to any one of claims 12 to 13.

25. A computer-readable medium storing a computer program, which when executed by a processor, implements the preprocessing method for the computing task according to any one of claims 12 to 13.

Citation Information

Patent Citations

  • Sparse tensor calculation method and device, equipment and storage medium

    CN109857744A

  • Lower trigonometric equation parallel solving method for structural grid sparse matrix

    CN111079078A