Calculation method and device, equipment, storage medium and program product

By optimizing the operator segmentation algorithm in the slicing computing device, adjusting the number of slicing and number of cache blocks of the operator, the problems of low matrices or tensor calculation efficiency and waste of cache resources in the prior art are solved, and efficient calculation and resource utilization are achieved.

CN120104555APending Publication Date: 2025-06-06HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311654360.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently perform matrix or tensor computing in complex systems on chip, especially in the environment of multi-type computing engines and data pipelines, resulting in low computing efficiency and waste of cache resources.

Method used

By determining the number of slicing operators and the number of cache blocks in the slicing computing device, adjusting the length of the pipeline, optimizing the operator slicing algorithm, reducing the reuse of data transfer and computing resources, and improving computing efficiency and cache utilization.

Benefits of technology

It realizes improving the efficiency of matrix or tensor calculation in complex systems on chips, saving cache resources, and reducing computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104555A_ABST
    Figure CN120104555A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, a storage medium and a program product. In the method, a segmentation calculation device determines the segmentation number of at least one layer of circulation used for executing an operator and the number of needed cache blocks, and the operator is a unit for executing an instruction. Further, the segmentation calculation device determines, in a pipeline including an operator, an adjustment point for adjusting the length of the pipeline. Further, the segmentation calculation device adjusts at least one of the number of segmentation portions and the number of cache blocks based on the adjustment point. Therefore, according to the embodiment of the invention, the calculation efficiency can be improved, and cache resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application generally relate to the field of computers, and more specifically, to a method, apparatus, device, computer-readable storage medium, and computer program product for computing. Background Art

[0002] For processors such as graphic processing units (GPUs) and network processing units (NPUs), such as tensor accelerators, their on-chip systems are very complex. They may have different types of computing engines (Computation Engines) such as scalar, vector, and matrix calculations (Cube), as well as internal caches (Local Memory) and data pipelines (Pipes). Efficiently running data processing such as matrix or tensor calculations on them is a complex and extremely challenging task, and the specific processing method needs to be optimized. Summary of the invention

[0003] The embodiments of the present application provide a technical solution for data processing, which can split the data processing of matrix or tensor operations, for example, to improve the operation efficiency.

[0004] In the first aspect, a data processing method is provided. The method can be executed by a split computing device. Unless otherwise specified, the split computing device in the embodiment of the present application can refer to the split computing device itself (for example, implemented as a terminal device, a server device), or a component in the split computing device (for example, a processor, a chip, or a chip system, etc.), or a logic module or software that can realize all or part of the functions of the split computing device. The following is described as an example in which the execution subject is a split computing device. In this method, the split computing device determines the number of splits and the number of cache blocks required for executing at least one layer of loop of an operator, and the operator is a unit for executing instructions. Furthermore, the split computing device determines an adjustment point for adjusting the length of the pipeline in a pipeline containing the operator. Furthermore, the split computing device adjusts at least one of the number of splits and the number of cache blocks based on the adjustment point. In this way, the embodiments of the present application can determine the reasonable initial state of the operator splitting algorithm by determining the number of splits and the required number of cache blocks for executing at least one layer of loop of the operator, obtain the most time-consuming and adjustment-required stage of the operator by selecting an adjustment point in the pipeline containing the operator, and adjust at least one of the number of splits and the number of cache blocks based on this adjustment point, optimize the number of operator splits and the required cache space by step-by-step tuning, and ultimately improve computing efficiency and save cache resources.

[0005] In some implementations, at least one loop of an operator corresponds to at least one axis. The axis includes at least one of an independent axis and a repeated axis. An independent axis corresponds to a loop that does not repeat data transfer, and a repeated axis corresponds to a loop that repeats data transfer. In this way, the embodiments of the present application can reasonably classify loops in data processing, such as tensor data processing of large amounts of data, in the form of independent axes and repeated axes, which is conducive to separate processing in optimization.

[0006] In some implementations, the splitting calculation device obtains a calculation graph, and the calculation graph includes a mathematical calculation relationship and a scheduling pipeline between operators, and the scheduling pipeline corresponds to at least one layer of loops that execute the operator, wherein the mathematical calculation relationship and the scheduling pipeline between the operators of the calculation graph are used to determine the number of splits and the number of cache blocks required for executing at least one layer of loops of the operator. In this way, the embodiments of the present application can intuitively describe the logical relationship and scheduling order between operators in the form of a calculation graph, which is beneficial to the analysis of developers, and also accurately describes the coupling mode between operators, thereby providing accurate input for determining the number of splits and the number of cache blocks required for at least one layer of loops that execute the operator, determining a good initial state for the optimization of operator splitting, and ultimately improving the optimization efficiency.

[0007] In some implementations, the computation graph includes a directed acyclic graph, the operators correspond to the nodes of the directed acyclic graph, the mathematical computation relationships between the operators correspond to the topological connection relationships and node types of the directed acyclic graph, the node types correspond to the data handling operators and the data computation operators, the scheduling pipelines correspond to the attributes of the nodes, and obtaining the computation graph includes: obtaining the topological connection relationships, node types, and node attributes of the directed acyclic graph. In this way, the embodiments of the present application can accurately describe the topological logical relationships and node types between nodes using a directed acyclic graph, thereby providing accurate input for determining the number of splits and the number of cache blocks required for at least one layer of loops used to execute the operator, determining a good initial state for the optimization of the operator split, and ultimately improving the optimization efficiency.

[0008] In some implementations, the cache includes one or more of a swappable cache, a fixed cache, and a hybrid cache. The swappable cache is used to determine the cached data in a first-in-first-out manner, and the fixed cache is used to continuously cache specific data. The hybrid cache uses a swappable cache on the outer axis for data access and a fixed cache on the inner axis for data access, the outer axis corresponds to the outer loop after the split loop, and the inner axis corresponds to the inner loop after the split loop. In this way, the embodiments of the present application can use a variety of different cache methods to store data, and can use different cache management methods for caching for different data features, so that the number of cache hits can be increased for different data features, thereby improving cache usage efficiency.

[0009] In some implementations, in a swappable cache, the number of cache hits is determined by the number of slices of the repeat axis and the number of slices of the free axis in the repeat axis. In this way, the embodiment of the present application can obtain the number of cache hits through theoretical calculations based on the number of slices of the repeat axis and the number of slices of the free axis in the repeat axis without actual board testing, and perform optimization processing, thereby improving optimization efficiency.

[0010] In some implementations, a fixed cache is used under the condition that the number of cache blocks is less than the minimum cycle period for repeated data access, and the number of hits of the fixed cache is determined according to the number of data blocks segmented by the operator, the number of data block repetitions, and the number of cache blocks. In this way, the embodiments of the present application use a fixed cache method to ensure that part of the data is always stored in the cache and can be hit under the condition that the number of cache blocks is less than the minimum cycle period for repeated data access, which can improve the cache hit rate.

[0011] In some implementations, the split computing device determines the number of splits and the required number of cache blocks for executing at least one layer of loop of the operator, including: the split computing device determines the number of splits and the required number of cache blocks for executing at least one layer of loop of the operator without exceeding the cache capacity limit. In this way, the embodiment of the present application can determine a suitable initial state for the number of splits and cache adjustment of the operator, and the initial value of the required number of cache blocks does not exceed the capacity limit of the cache space, providing a good foundation for the subsequent optimization of operator splitting.

[0012] In some implementations, under the condition that the cache capacity limit is not exceeded, the splitting calculation device determines the number of splits and the required number of cache blocks for executing at least one layer of loop of the operator, including: the splitting calculation device determines the cache occupancy of at least one cache and determines the excess cache according to the cache occupancy. Further, the splitting calculation device determines the benefit of the cache occupancy of the excess cache caused by increasing the number of splits of the axis. Further, the splitting calculation device determines the cost caused by increasing the number of splits of the axis. Further, the splitting calculation device determines the cost performance based on the benefit and the cost. Further, the splitting calculation device determines the specific axis to increase the number of splits according to the maximum value of the cost performance. Further, the splitting calculation device determines the number of splits for executing at least one layer of loop of the operator according to the increased number of splits of the specific axis. Further, the splitting calculation device determines the required number of cache blocks according to the number of splits for executing at least one layer of loop of the operator. In this way, when the cache of an operator exceeds the capacity limit, the embodiment of the present application can adjust the specific axis with the best cost-effectiveness for the over-limit cache and the cost-effectiveness obtained by adjusting each axis, increase the number of splits, and thus reduce the number of cache blocks required, until the cache of the operator is no longer within the capacity limit. In this way, the appropriate initial state can be determined for the number of splits and cache adjustment of the operator, which is beneficial to the subsequent optimization process of operator splitting.

[0013] In some implementations, under the condition that the cache capacity limit is not exceeded, the split calculation device determines the number of splits and the number of cache blocks required for executing at least one layer of loop of the operator, and further includes: the split calculation device determines the step size of increasing the number of splits according to the minimum value of the cache occupancy rate of the over-limit cache affected by the specific axis. Furthermore, the split calculation device determines the number of splits after the increase of the specific axis according to the step size of increasing the number of splits. In this way, the embodiment of the present application can determine a reasonable value of the step size of increasing the number of splits, optimize the adjustment process of the number of splits, and finally determine a suitable initial state for the adjustment of the number of splits and cache of the operator.

[0014] In some implementations, the splitting calculation device determines the adjustment point for adjusting the length of the pipeline in the pipeline containing the operator, including: the splitting calculation device removes the gap in the pipeline containing the operator to obtain the pipeline histogram. Furthermore, the splitting calculation device obtains the longest item among multiple items in the histogram. Furthermore, the splitting calculation device selects the longest time-consuming stage in the longest item as the adjustment point. In this way, the embodiment of the present application can find the operator that consumes the most time and needs to be adjusted the most by finding the longest time-consuming stage of the longest item in the histogram, and reasonably set the adjustment point, that is, the position of the operator that needs to be adjusted. By selecting the appropriate operator for adjustment, the overall optimization efficiency of the pipeline is improved.

[0015] In some implementations, the slicing computing device adjusts at least one of the number of slicing shares or the number of cache blocks, including one or more of the following: the slicing computing device reduces the number of slicing shares of the data handling operator, the slicing computing device increases the number of cache blocks of the data handling operator, or the slicing computing device reduces the number of slicing shares of the data computing operator. In this way, the embodiments of the present application flexibly adjust the operator in a variety of ways, ultimately achieving the purpose of "less slicing, less moving", and comprehensively optimizing the slicing of the operator.

[0016] In some implementations, the splitting computing device reduces the number of splits of the data handling operator, including one or more of the following: the splitting computing device reduces the number of splits of the repeated axis of the data handling operator, the splitting computing device adjusts the number of cache blocks at no cost under the condition that the cache capacity limit is not met, the splitting computing device adjusts the number of cache blocks at a cost under the condition that the cache capacity limit is not met, or the splitting computing device increases the number of splits of the data handling operator's own axis under the condition that the cache capacity limit is not met. In this way, the embodiment of the present application can reduce repeated handling by reducing the number of splits of the repeated axis of the data handling operator; then adjust the number of buffer blocks in a costly and costly manner to reduce cache occupancy and meet the cache capacity limit; and can also reduce cache occupancy by increasing the number of splits of the free axis to meet the cache capacity limit. Finally, the number of splits of the data handling operator is reduced in a comprehensive manner to improve the computing efficiency of the operator.

[0017] In some implementations, the splitting calculation device reduces the number of splits of the repetitive axis of the data handling operator, including: under the condition that the data handling operator has multiple repetitive axes, the splitting calculation device selects the specific repetitive axis that has the greatest impact on the amount of data to be repeatedly transported. Furthermore, the splitting calculation device adjusts according to a step size of 1 to reduce the number of splits of the specific repetitive axis, and the number of splits of the adjusted specific repetitive axis satisfies the requirement of being rounded down twice. In this way, the embodiment of the present application can select the specific repetitive axis that has the greatest impact on the amount of data to be repeatedly transported, that is, the key repetitive axis for adjustment, and make fine adjustments with a step size of 1. The number of splits of the adjusted specific repetitive axis satisfies the requirement of being rounded down twice, that is, it can be divided by the overall dimension without remainder, thereby ensuring the integrity of the data operation. The embodiment of the present application uses the above-mentioned comprehensive optimization method to reduce the number of splits of the repetitive axis of the data handling operator and improve the calculation efficiency of the operator.

[0018] In some implementations, the split computing device adjusts the number of cache blocks at no cost, including: the split computing device calculates a cost-effectiveness list and an adjustment step size of at least one cache. Furthermore, the split computing device selects an adjustable cache according to the condition that the number of cache blocks is greater than a lower limit. Furthermore, the split computing device reduces the number of cache blocks of the selected cache according to the adjustment step size. In this way, the embodiment of the present application can select a suitable cache and reduce the number of cache blocks without paying any price, thereby satisfying the cache space constraint as much as possible and improving the splitting efficiency of the operator.

[0019] In some implementations, the split computing device adjusts the number of cache blocks at a cost, including: The split computing device selects an adjustable cache according to the condition that the number of cache blocks is greater than a lower limit. Furthermore, the split computing device adjusts the number of cache blocks of the adjustable cache with a small cost under the condition of a surplus. Furthermore, the split computing device adjusts the number of cache blocks of the adjustable cache with a high cost-performance ratio under the condition of no surplus. In this way, the number of cache blocks can be adjusted at a cost in an optimized manner. For example, the embodiment of the present application can further adjust the number of cache blocks at a cost in an optimized manner after using the costless adjustment method. Moreover, adjustments are made successively under the conditions of surplus and no surplus, minimizing the cost and obtaining benefits, and the number of cache blocks is adjusted at a cost in a comprehensive manner, as little as possible within the capacity limit of the cache, to optimize the splitting of the operator.

[0020] In some implementations, the slicing computing device increases the number of slicing shares of the own axis of the data handling operator under the condition that the cache capacity limit is not met, including: the slicing computing device freezes the repeated axis of the data handling operator under the condition that the cache capacity limit is not met. Then, the slicing computing device determines the cache occupancy rate of at least one cache and determines the over-limit cache according to the cache occupancy rate. Then, the slicing computing device determines the benefit of the cache occupancy rate of the over-limit cache caused by increasing the number of slicing shares of the own axis. Then, the slicing computing device determines the cost caused by increasing the number of slicing shares of the own axis. Then, the slicing computing device determines the cost performance based on the benefit and the cost. Then, the slicing computing device determines the specific own axis for increasing the number of slicing shares according to the maximum value of the cost performance. In this way, the embodiment of the present application can avoid the influence of adjusting the repeated axis on the own axis by freezing the repeated axis, calculate the cost performance of the adjustment of each own axis by the over-limit cache, and determine the specific own axis for final adjustment and increase of the number of slicing shares based on the highest cost performance. In this way, the optimal own axis is selected for adjustment in a comprehensive manner, and the number of splits of the own axis of the data handling operator is increased without meeting the cache capacity limit, thereby ultimately improving the performance of operator splitting.

[0021] In some implementations, the slicing computing device increases the number of cache blocks of the transport operator, including: the slicing computing device determines the transport operator with the most repeated transport. Then, the slicing computing device increases the number of cache blocks of the transport operator. Then, the slicing computing device freezes the transport operator's own axis under the condition that the cache capacity limit is not met, and increases the number of slicing portions of the repeated axis of the transport operator. In this way, the embodiment of the present application increases the cache hit rate of the operator by increasing the number of cache blocks of the transport operator for the node with the most repeated transport, and freezes the transport operator's own axis and increases the number of slicing portions of the repeated axis of the transport operator for the cache capacity exceeding the limit that may be caused by increasing the number of cache blocks. Finally, the number of cache blocks of the transport operator is increased without exceeding the capacity limit, thereby improving the cache hit rate of the transport operator and improving the execution efficiency of the transport operator.

[0022] In some implementations, the splitting computing device reduces the number of splits of the data computing operator, including: the splitting computing device determines the data computing operator with the longest task execution time. Then, the splitting computing device reduces the number of splits of the inner axis of the own axis of the data computing operator with the longest execution time. Then, the splitting computing device adjusts the number of cache blocks at no cost under the condition that the cache capacity limit is not met. Then, the splitting computing device also adjusts the number of cache blocks at a cost under the condition that the cache capacity limit is not met. In this way, the embodiment of the present application can reduce the number of splits of the own axis of the data computing operator with the longest execution time, and solve the problem of cache capacity overrun by adjusting the number of cache blocks at no cost and at a cost in turn under the condition that the cache capacity may exceed the limit. Finally, the number of splits of the data computing operator is reduced by a comprehensive optimization method to achieve "less cutting" and improve the computing efficiency of the data computing operator.

[0023] In a second aspect, a data processing device is provided, which may be the split computing device in the above method embodiment, or a chip arranged in the split computing device. The data processing device includes: a split number and cache block number determination module, which is used to determine the split number and the required number of cache blocks for executing at least one layer of loop of an operator, where the operator is a unit for executing instructions. The data processing device also includes: an adjustment point determination module, which is used to determine an adjustment point for adjusting the length of the pipeline in a pipeline containing the operator to reduce the length of the pipeline. The data processing device also includes: an adjustment module, which is used to adjust at least one of the split number or the number of cache blocks based on the adjustment point. In this way, the embodiments of the present application can determine the reasonable initial state of the operator splitting algorithm by determining the number of splits and the required number of cache blocks for executing at least one layer of loop of the operator, obtain the most time-consuming and adjustment-required stage of the operator by selecting an adjustment point in the pipeline containing the operator, and adjust at least one of the number of splits and the number of cache blocks based on this adjustment point, optimize the number of operator splits and the required cache space by step-by-step tuning, and ultimately improve computing efficiency and save cache resources.

[0024] In a third aspect, a device is provided, which may be the split computing device in the above method embodiment, or a chip set in the split computing device. The device includes a processor and a memory. The memory is used to store a computer program or instruction, and when the processor runs the computer program or instruction, the split computing device executes the method executed by the split computing device in the above method embodiment.

[0025] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the method performed by the split computing device in the above aspects is implemented.

[0026] In a fifth aspect, a computer program product is provided, the computer program product comprising: a computer program code, when the computer program code is run, the method performed by the split computing device in the above aspects is executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1A A computing system in which embodiments of the present application can be implemented.

[0028] Figure 1B Schematic diagram of matrix segmentation calculation.

[0029] Figure 2A The flowchart of the method of segmentation calculation in the embodiment of the present application is shown in FIG.

[0030] Figure 2B This is a schematic diagram of the statement axis decomposition in the embodiment of the present application.

[0031] Figure 2C Schematic diagram of a directed acyclic graph in an embodiment of the present application.

[0032] Figure 2D A schematic diagram of the hardware architecture in an embodiment of the present application.

[0033] Figure 2E It is a schematic diagram of the hardware architecture in another embodiment of the present application.

[0034] Figure 3A A schematic diagram of a swappable cache in an embodiment of the present application.

[0035] Figure 3B A schematic diagram of a fixed cache in an embodiment of the present application.

[0036] Figure 3C A schematic diagram of a hybrid cache in an embodiment of the present application.

[0037] Figure 3D A schematic diagram of a hybrid cache in an embodiment of the present application.

[0038] Figure 4 A schematic diagram of the segmentation calculation in an embodiment of the present application.

[0039] Figure 5 This is a schematic diagram of setting the segmentation axis and the number of segments in an embodiment of the present application.

[0040] Fig. 6A It is a schematic diagram of the pipeline in the embodiment of the present application.

[0041] Figure 6B Schematic diagram of a pipeline histogram in an embodiment of the present application.

[0042] Figure 7 This is a flow chart for reducing the number of parts for a transport task in an embodiment of the present application.

[0043] Figure 8 The flowchart of the costless cache adjustment and the costly cache adjustment in the embodiment of the present application.

[0044] Fig. 9 It is a schematic diagram of a costless adjustment cache and a costly adjustment cache in another embodiment of the present application.

[0045] Fig.10 This is a flowchart for increasing the number of cache blocks for a transport task in an embodiment of the present application.

[0046] Fig.11 A flowchart for reducing the number of parts for a computing task in an embodiment of the present application.

[0047] Fig.12 It is a block diagram of the segmentation calculation device in the embodiment of the present application.

[0048] Fig.13 It is a structural diagram of the split computing device in the embodiment of the present application.

[0049] Fig.14 It is a schematic diagram of the structure of a partitioned computing device cluster in an embodiment of the present application.

[0050] Fig.15 A schematic diagram of the structure of a cluster of split computing devices connected via a network in an embodiment of the present application. DETAILED DESCRIPTION

[0051] Figure 1A A computing system in which the embodiments of the present application can be implemented is provided in the computing system 100, which includes a memory 101, a cache 103, and an execution unit 105. The data in the memory 101 is transferred to the execution unit 105 through the cache 103 to perform calculations such as data segmentation.

[0052] For processors of tensor accelerators such as GPU and NPU, their system on chip (System on Chip, SOC) is very complex, and there may be different types of computing engines (Computation Engine) such as scalar (Scalar), vector (Vector), matrix (Cube), and internal cache (Local Memory) and data pipeline (Pipeline) such as cache. Efficiently running data processing such as tensor calculation on it is a complex and challenging task. Tensor can be a mapping processing method for data such as scalar, vector, and matrix, or it can be a multi-dimensional data representation method. Tensor can also include scalar, vector, and matrix, which is not limited in this disclosure. On the one hand, since the computing engine and the data pipeline are independent of each other, they can be asynchronously parallel in the entire operator calculation process; on the other hand, the cache capacity and bandwidth of each level are different. If you want to form an efficient asynchronous concurrent pipeline calculation (data moving in + calculation + data moving out), you must consider splitting the tensor data. In addition, the internal cache capacity is relatively small relative to the off-chip storage (Global Memory, such as DDR, HBM, etc.), and objectively it is also required to split the input tensor data. For example, in the matrix multiplication Matmul calculation C = Matmul (A, B), the left and right hand matrices are usually split (tiling), and then the data blocks are read and calculated in a certain traversal order. Figure 1B In the above example, the second row in matrix A multiplied by the second column in matrix B obtains the second row and the second column in matrix C, and the second row in matrix A multiplied by the third column in matrix B obtains the second row and the third column in matrix C. The above matrix segmentation calculation method needs to be optimized.

[0053] At present, the prior art consists of a sampling module, a performance evaluation module and related operator executable files to form a tuning module. The tuning module can randomly sample from a set of operator tiling parameter configurations prepared in advance, and then execute on the board to obtain performance feedback, and select the configuration of the first performance as the tiling parameter of the final operator for execution. This solution selects the optimal configuration item from a known configuration set through random sampling + boarding. First, this technology does not generate a tiling strategy (i.e., operator parameter configuration item), but manually prepares the strategy in advance, and can only be selected from an existing configuration set during tuning. Secondly, random sampling is not combined with any prior or heuristic rules, and the efficiency is low. Finally, performance estimation is done through board feedback, which involves using configuration items to compile and run operators, and tuning takes a long time.

[0054] The above solution is time-consuming, requires customized development, has a large development workload, requires performance optimization, and requires on-board testing for performance evaluation, which is inefficient. Therefore, it is necessary to find a more optimized data segmentation method. The disclosed embodiment proposes a general segmentation strategy calculation method that runs fast (millisecond level), can support different hardware platforms, and supports different types of operators.

[0055] The disclosed embodiment adopts a heuristic method to select a specific axis to increase the number of slicing shares according to the cost performance, and then determine the number of slicing shares and the number of cached data blocks, and determine whether to split according to the priority and cost value of the axis. Instead of using black box methods such as meta-heuristic methods or AI methods, the speed of tiling strategy calculation can be greatly improved, and the performance is more stable. Abstracting different hardware architectures into a unified expression (computation and pipeline), it can support heterogeneous architectures (scalar / cube / vector) and hierarchical storage (L2 / L1 / L0 / Unified Buffer / Fix Data Pipe, etc.), and the method can be applied to different tensor acceleration processors such as GPU and NPU. The algorithm does not depend on the specific operator calculation process and is universal. The calculation graph and its node attributes are used to describe the mathematical calculation relationship and scheduling pipeline of the operator. The disclosed embodiment also uses a pipeline histogram to evaluate performance, rather than an ordinary end-to-end (End to End, E2E) bubble pipeline or board verification, which reduces the cost of performance evaluation. The embodiments of the present disclosure can be used in an operator development software stack or tool chain, and can be applied to any development components of platform operators involving tensor calculations.

[0056] Figure 2A The flowchart of the method for splitting calculation in the embodiment of the present application is shown in FIG. 201. The number of splits and the number of cache blocks required for executing at least one layer of loop of the operator are determined, and the operator is a unit for executing instructions. In 203, an adjustment point for adjusting the length of the pipeline including the operator is determined. In 205, at least one of the number of splits and the number of cache blocks is adjusted based on the adjustment point. The number of cache blocks is the number of cache blocks used in the cache space when the operator is executed. In this way, the embodiment of the present application can optimize and adjust the operator, determine the reasonable initial state of the operator splitting algorithm by determining the number of splits and the number of cache blocks required for executing at least one layer of loop of the operator, obtain the most time-consuming and adjustment-required stage of the operator by selecting an adjustment point in the pipeline including the operator, and adjust at least one of the number of splits and the number of cache blocks based on the adjustment point, optimize the number of operator splits and the required cache space by step-by-step tuning, and finally improve the computing efficiency and save cache resources.

[0057] The general operator tiling calculation method can have a unified form of input. The tiling calculation can obtain the specific implementation of the operator as the input for optimizing the operator tiling. The specific implementation of the operator includes the mathematical calculation relationship of the operator and the scheduling pipeline (also known as the code skeleton). The mathematical calculation relationship determines the type and front-end dependency of the high-level application program interface (API) or instruction-level API that appears in the operator, while the scheduling pipeline defines the pipeline execution of data transfer and calculation in the hardware chip when the operator is executed. Figure 2B Schematic diagram of the statement axis in the embodiment of the present application. In the embodiment of the present disclosure, "axis" corresponds to the loop in the operator implementation, such as the for loop statement. And "axis" can include "own axis" and "repeat axis". "Own axis" corresponds to the loop that does not repeat the data, and "repeat axis" corresponds to the loop that repeats the data. For example, Figure 1B In the matrix multiplication in the calculation of C(2,2), the second row of matrix A is read from A in a loop, and the data is not moved repeatedly, which corresponds to the "own axis". In the subsequent calculation of C(2,3), the second row of matrix A is read and moved repeatedly, which corresponds to the "repeated axis". Take the matrix multiplication (Matmul) operator of C(m,n)=A(m,k)×B(k,n) as an example, Figure 2BThree different operator implementations are shown. In embodiment 210, 211 corresponds to an axis splitting, where the M axis is split into M0 and M1 axes, that is, the loop M is split into two loops of M0 and M1; the N axis is split into N0 and N1 axes, that is, the loop N is split into two loops of N0 and N1. The statements are at the same level, that is, copyA(), copyB(), loadA(), loadB(), and mad(A,B) statements are at the same level. The "axes" in the subsequent sections all correspond to "loops", which will not be repeated in this disclosure. For example, 213 corresponds to an axis without splitting, where the M0 and N0 axes are not split and rearranged in order. 215 corresponds to an axis with splitting, where the statements are at different levels, that is, copyA() and copyB() are on one layer, and loadA(), loadB(), and mad(A,B) statements are on another layer. The copy() statement can transfer data from global memory such as external RAM to L1 cache (Layer1 cache, first-level cache), while the load() statement can transfer data from L1 cache to L0 cache (Layer 0 cache, a level 0 cache that is closer to the engine and has a faster access speed than L1 cache). Among them, the five APIs copyA, copyB, loadA, loadB, and mad describe the mathematical calculation relationship of Matmul, while the different arrangement orders, number of occurrences, and different coverage situations of the five loop sub-axes m0, m1, n0, n1, and k0 describe the scheduling pipeline calculation. "Sub-axis" can be a specific implementation of "axis", and is also used to represent loops. For example, it can be a sub-axis obtained by splitting the axis, or it can be the "axis" itself. The splitting strategy can refer to determining the number of splits for each loop sub-axis (equivalent to the split size or the number of executions of the loop, such as Figure 2B The three code architectures 211, 213, and 215 constitute a complete operator execution code. In this way, the embodiment of the present application can appropriately classify the operator execution code, and use the selected code architecture as the input of the operator optimization, which is conducive to determining the initial state of the operator segmentation optimization, and ultimately improve the operator optimization efficiency.

[0058] In the disclosed embodiment, a directed acyclic graph (DAG) is used to express operator implementation, mathematical calculation relationships are mapped to the topological connection relationships and different node types of the graph, and the scheduling pipeline can be described by node attributes. Figure 2CIt is a schematic diagram of the directed acyclic graph in the embodiment of the present application, and specifically shows the property definition diagram of the calculation graph implemented by the Matmul calculation kernel and its scheduling pipeline. The operator API therein is divided into two types, namely, the transport API and the calculation API, which are represented by different filled shapes. The transport API can be an operator for data transport, such as Copy, Load, FixpipeTrans, etc., and the calculation API can be an operator for data calculation, such as mad, etc. The processing characteristics of the transport API and the calculation API are different, and they can be optimized in different ways later, thereby improving the effect of comprehensive optimization. The scheduling pipeline corresponds to the properties of the operator, such as [K / 16, N / 16, 16, 16] corresponding to the loop. For example, in embodiment 220, the input 221 containing the sequential structure passes through the CopyA operator 223, the LoadA operator 225, the mad operator 227, the Fixpipetrans operator 229, and the output 0, etc. The directed acyclic graph DAG is used as the input of the operator segmentation optimization algorithm, and at the same time Figure 2B One of the three code architectures is selected, which is also used as the input of the operator segmentation optimization algorithm. The output of the optimization algorithm is: the numerical value of the specific number of segments of each axis m0, n0, m1, n1, etc., and the cache depth in the cache, or the number of cache blocks. In this way, the embodiment of the present application can use an intuitive directed acyclic graph to characterize the relationship between the operators, which is conducive to intuitive design, and accurately describes the coupling method between operators, thereby providing accurate input for determining the number of segments and the required number of cache blocks for at least one layer of loop used to execute the operator, determining the initial state of optimization for the optimization of operator segmentation, and ultimately improving the optimization efficiency for optimization.

[0059] In the disclosed embodiment, the acceleration chip for calculating tensors is a heterogeneous system on chip (System OnChip, SoC), including a scalar engine for executing instructions and scalar calculations, a vector engine for performing vector type calculations, and a Cube engine for performing matrix type calculations. In the disclosed embodiment, the dimension of the tensor is higher than that of the vector (one-dimensional) and the matrix (two-dimensional). Tensors, vectors, and matrices are calculated using different engines to improve computing efficiency. There will also be multi-level cache buffers (such as L2 / L1 / L0 Buffer) in the chip for temporary storage of data, as well as pipelines for data transmission, etc. Figure 2DSchematic diagram of the hardware architecture in the embodiment of the present application. In embodiment 240, the system on chip (SOC) 242 has a global memory (Global Memory) 243, various caches and a computing engine. For example, the cache may have L1_input cache 245, L0a_input cache 247, L0_output cache 251, L1_output cache 257, etc., and the computing engine includes a scalar engine (scalar engine) 253, a vector engine (vector engine) 255, a matrix engine (cube engine) 249, etc. Caches are connected to each other, or caches and engines are connected by pipes. For example, the global memory (Global Memory) 243 and the L0a_input cache 247 are connected by pipe 0 259. The operator optimization method in the embodiment of the present disclosure can be applied to Figure 2D The hardware architecture is shown.

[0060] Figure 2E Schematic diagram of the hardware architecture in another embodiment of the present application. In embodiment 260, the system on chip (SOC) 262 includes: a global memory (Global Memory) 261, a computing core 0 263 and a computing core 1 279. The computing core 0 263 and the computing core 1 279 contain various caches and computing engines. For example, the cache of the computing core 0 263 may include L1_input cache 265, L0a_input cache 267, L0_output cache 271, L1a_output cache 273, etc., and the engine may include a matrix engine 269. The cache of the computing core 1 279 may include L1b_output cache 285, and the engine may include a scalar engine (scalarengine) 281, a vector engine (vectorengine) 283. The operator optimization method in the embodiment of the present disclosure may also be applicable to Figure 2E Those skilled in the art will appreciate that the operator optimization method in the embodiments of the present disclosure may also be applicable to other hardware architectures, such as other hardware architectures used in GPUs and NPUs, and the present disclosure does not limit this.

[0061] exist Figure 2D and Figure 2E When running operator calculations on a hardware architecture, the factors that affect its performance may include the following four points: data handling, scalar engine calculation amount, number of instructions, and calculation instruction efficiency.

[0062] In the embodiment of the present disclosure, for data transport, Figure 2D , Figure 2EAsynchronous data transmission pipes such as pipe0 and pipe1 are responsible for data transfer within the chip. The factors that affect the transfer performance are the transfer instruction efficiency and the amount of repeated data transfer. If not handled properly, it may cause memory bound, that is, performance bottlenecks due to limited cache capacity. Generally speaking, when a transfer instruction is called to process a large amount of data, the number of single instruction repetitions can improve instruction efficiency, because the total amount of data remains unchanged, and the number of instruction calls will be reduced accordingly (that is, the header overhead of instruction execution is reduced); for a certain calculation process, the amount of data that is repeatedly transferred is related to the number of divisions of the sub-axis (SubAxis) to which it belongs, as shown in the following code:

[0063]

[0064] In the disclosed embodiment, axis n is the repeat axis of data A, and the number of its cuts N determines the number of repeated transfers of data A. Axis m and axis k are the own axes of data A, and the number of their cuts M and K determine the amount of data transferred each time. Assuming that the total amount of data A is H, the total amount of repeated transfer data is N*H, and the number of repeated transfers is M*K*N. Therefore, by controlling the number of cuts N of the repeat axis n, the amount of repeated transfer data can be reduced.

[0065] In the disclosed embodiment, regarding the amount of scalar engine calculations: if there are a large number of scalar calculations in the operator code running on the system on chip (SoC), there will be a scalar bound problem, that is, a performance bottleneck caused by limited scalar engine resources. Generally, scalar calculations are mostly related to branch condition judgments, loop judgments, cache operation judgments, etc. If the number of split blocks is small and the block size is large, the number of loops will be reduced accordingly. Under the condition that the total cache capacity is constant, the number of buffer blocks in the cache space will also be reduced, and the scalar calculations will be reduced accordingly.

[0066] In the disclosed embodiment, there is a startup header overhead for the instruction execution in terms of the number of instructions. If the number of instructions to complete the same amount of computation is small, the header overhead will be reduced accordingly. In addition, a large number of instructions also makes it easier to generate instruction cache misses, thereby affecting the efficiency of instruction fetching. Therefore, reducing the number of cycles also helps to reduce the number of instructions.

[0067] In the embodiments of the present disclosure, regarding the computing instruction efficiency, compute bound will occur in some scenarios, that is, a performance bottleneck caused by limited computing resources, such as in scenarios with a small amount of data and a large amount of computing. However, in some cases, it may be caused by problems with the computing instruction efficiency, and the principle is the same as the problem explained in the "number of instructions".

[0068] By analyzing the above four factors that affect pipeline computing efficiency, namely data handling, scalar engine computing class, instruction quantity and computing instruction efficiency, we can obtain the key points of the operator segmentation strategy in the tensor accelerator on-chip system: (1) reduce repeated data handling (2) improve instruction efficiency (including computing and handling). Furthermore, the heuristic rule used in the embodiment of the present disclosure is derived: under the condition of meeting the buffer capacity in the core, reduce the number of segments and reduce repeated handling, that is, the "less cutting, less moving" principle.

[0069] In the disclosed embodiment, for cache space, the tiling strategy can solve two problems: 1. Make full use of the intra-core hierarchical cache while ensuring that the data does not exceed the cache capacity limit; 2. Reduce duplicate data transfer. In the disclosed embodiment, repeated transfer is due to the existence of non-own duplicate sub-axes in the data transfer operator of the transfer API. For example, the tensor shape and sub-axes of the transfer API node A are [M=4, N=2], and the other axes are L and K. The scheduling code is as follows:

[0070]

[0071] Axes M0, M1, and N can be the own sub-axes of node A, and axes L and K can be repeated sub-axes of node A. For example, the splitting strategy is to split M0 and M1 twice, each time into 2 parts, and split L, K, and N once, each time into 2 parts. Therefore, node A is split into 8 parts (called 8 slice elements), numbered from 0 to 7, and repeated 32 times in total. The "repetition" times all include the first time. The repetition pattern is as follows, and the numbers are slice numbers.

[0072] 0 1 2 3||0 1 2 3||4 5 6 7||4 5 6 7||0 1 2 3||0 1 2 3||4 5 6 7||4 5 67

[0073] The corresponding blocks of M1 and N loops are 0, 1, 2, 3. The loop at M0 corresponds to the distinction between 0, 1, 2, 3 and 4, 5, 6, 7. In the multiple loops, the upper layer of the M1 and N loops is the repeat axis K, so 0, 1, 2, 3 are looped twice. Then it goes to the own axis M0 and switches to 4, 5, 6, 7. At this time, 4, 5, 6, 7 also loop twice because of the repeat axis K. Therefore, the loop structure of 0 1 2 3||0 1 2 3||45 6 7||4 5 6 7|| appears.

[0074] In the disclosed embodiment, the number of executions of the data transfer operator under the repeated sub-axis, such as the transfer API, can be reduced by reducing the split score of the repeated sub-axis. Alternatively, the cache space can be divided on each level of the cache in the core and a cache mechanism can be established to cache the data that needs to be transferred repeatedly. Once the accessed data hits the cache, the purpose of reducing repeated transfers can be achieved.

[0075] For example, the Matmul operator must first copy the A matrix from the global memory (GM) to L1 and then load it to L0a. Since the copyA node may have repeated sub-axis N and needs to be moved repeatedly, if the cache block is managed by configuration, and a cache space is established on the L1 Buffer, the number of times data is copied from the GM can be reduced. The cache management of the cache space can be implemented in a variety of ways, such as cyclic cache, fixed cache and hybrid cache. The cyclic cache can also be called a "cyclic cache". According to different data characteristics, the embodiments of the present disclosure can adopt different cache management methods for caching, so that the number of cache hits can be increased according to different data characteristics, and the cache usage efficiency can be improved.

[0076] Figure 3A Schematic diagram of a swappable cache in an embodiment of the present application. In embodiment 300, for a matrix 305 of M rows and K columns, a queue with a queue length of 4 can be used for caching. The queue adopts a first-in first-out (FIFO) structure. For example, in queue 310, the cached data is 1, 2, 3, and 4. After the first shift, the cached data in queue 315 is 2, 3, 4, and 5. After the second shift, the cached data in queue 320 is 3, 4, 5, and 6.

[0077] Figure 3A The swappable queue in is similar to a circular queue. The number of hits to access each data block is the same, so we only need to analyze the number of hits for one of the data blocks and multiply it by the number of data blocks. Now we can only consider data block 0 in the above repeated transport cycle pattern: Node A is contained by two repetition axes, that is, there are 2 repetition levels, from inside to outside.

[0078] In the embodiment of the present disclosure, in repetition level 1, at the current repetition level, data block 0 appears 1 time and repeats again, with 4 data blocks (including itself) appearing in between. For example, for 0 1 2 3||0 1 2 3, 0 is repeated twice, with 4 data in between. 4 is recorded as the cycle period T1 at the current repetition level, T1 = the product of all the self-axis segmentation numbers in the current repetition axis; a total of 2 repetitions occur, recorded as the number of repetitions R1 = 2, R1 = the number of segments of the current repetition axis.

[0079] In the embodiment of the present disclosure, in repetition level 2, at the current repetition level, data block 0 appears once and then repeats again, with 8 data blocks (including itself) appearing in the middle. It is recorded as the cycle period T2 at the current repetition level, T2 = the product of all the self-axis segmentation numbers in the current repetition axis; it is repeated 2 times in total, recorded as the repetition number R2, R2 = the number of segments of the current repetition axis.

[0080] In the embodiment of the present disclosure, given that the number of cache blocks in the cache space is X, the pseudo code of the algorithm for calculating the number of cache hits is as follows:

[0081]

[0082] In this way, the embodiments of the present application do not need to pass actual board testing. Through the number of slices of the repeated axis and the number of slices of the free axis within the repeated axis, the number of cache hits can be obtained through theoretical calculation, and optimization processing can be performed, thereby improving optimization efficiency.

[0083] Figure 3B This is a schematic diagram of a fixed cache in an embodiment of the present application. In the aforementioned cyclic cache mode, if the number of cache blocks in the cache space is less than the minimum cycle period for repeated data access, the number of cache hits is zero. In this case, you can consider using a fixed cache to ensure that a portion of the data can hit the cache when it is accessed and is not cyclically swapped out. In embodiment 330, for a matrix 335 of M rows and K columns, a queue with a queue length of 4 is used for data caching. The queue 340 has fixed cache data 1, 2, 3, and 4, and the cached data content does not change, which is a fixed cache. Assuming that the number of segmented data blocks for each operator (node) is P, and each data block is repeated Q times, and the number of cache blocks in a given cache space is X, the number of cache hits is In this way, under the condition that the number of cache blocks is less than the minimum cycle period for repeated data access, the embodiments of the present application use a fixed cache method to ensure that part of the data is always stored in the cache and can be hit, thereby improving the cache hit rate.

[0084] In the embodiments of the present disclosure, under certain conditions, due to reasons such as the access order of tensor data, it is necessary to use cyclic caches on some outer axes and fixed caches on the remaining inner axes, so as to increase the number of cache hits in a limited cache space. Figure 3C and Figure 3D For example, for the matmul node, Figure 3C In , matrix 355 is multiplied by matrix 360 to obtain matrix 365. Split the M axis of matrix A into two sub-axes, M1 and M0, as Figure 3D As shown, the calculation order is as follows: For the A matrix 375 (corresponding to Figure 3C Assume that the number of cache blocks in the cache space is 6, the M0 axis is cyclic, and the M1 / K axis is fixed. The actual cache behavior is as follows: Figure 3C As shown. The number of cache blocks in the cache space is 6. First, cache data blocks 1, 2, 3, 4, 5, 6, and then cache 9, 10, 11, 12, 13, 14. Then, the cached data blocks can be hit, and the amount of data moved from the global memory is reduced. Data blocks 7, 8, 15, and 16 that have not entered the cache need to be moved from the global memory. It can be seen that the performance of the method of using the cache space to cache data that needs to be moved repeatedly and then reduce repeated movement mainly depends on the number of cache blocks in the cache space. The more cache blocks there are, the greater the probability of being hit by the cache. However, the large number of cache blocks means that the buffer capacity occupied by the cache space is large. Under the limit that the operator segmentation strategy ensures that the data block does not exceed the buffer capacity, the data segmentation score will be increased, and the amount of repeated data movement will also potentially increase, such as increasing the segmentation score on the repeated axis. Reducing the split score of the repeated axis and increasing the number of cache blocks in the cache space are contradictory operations, and the best performance balance point needs to be found. Therefore, the operator split strategy needs to output the number of cache blocks in the cache space established on each level of buffer in the core. In this way, the performance balance is optimized.

[0085] The disclosed embodiment uses "less cutting, less moving" as a heuristic rule, takes minimizing the maximum pipeline task execution time as the optimization goal, and continuously iterates the optimization goal until there is no performance improvement. The algorithm mainly handles two types of tasks, namely repeated moving tasks and computing tasks. That is, if there is no good optimization method for general data moving tasks, the principle of less cutting is already beneficial to the efficiency of moving instructions. For repeated moving tasks, two optimization methods, "less cutting of repeated axes" and "increasing the number of cache blocks", will be tried. For computing tasks, "less cutting of internal axes" is tried to increase the amount of instruction processing data and reduce the total number of instruction executions. When the performance cannot be further improved, simply optimize the repeated moving tasks, and try to reduce the amount of repeated moving data as long as the performance does not decrease. Although the end-to-end time of the longest pipeline has not been reduced in theory, the performance evaluation standard adopted by the algorithm is the simple accumulation of the execution time on each pipeline (for example, using a histogram), and does not consider the gaps (bubbles) caused by the mutual dependence between tasks, and there is an evaluation error. Therefore, the pipeline where the repeated moving task is located may be the busiest in reality, and this "harmless" post-processing logic needs to be executed. The overall process of the operator splitting calculation algorithm is as follows Figure 4 As shown, the heuristic rule of "less cutting, less moving" is specifically implemented.

[0086] In embodiment 400, operator computation graph 405 corresponds to Figure 2C . At 410, the splitting axis and the number of splits are set. The splitting axis corresponds to the loop, and the number of splits corresponds to the number of loops. In 415, the splitting calculation device selects an adjustment point. At 420, the splitting calculation device reduces the number of splits of the transport task. At 425, the splitting calculation device increases the number of cache blocks of the transport task. At 430, the splitting calculation device reduces the number of splits of the computing task. At 435, the splitting calculation device performs optimization. At 440, the splitting calculation device determines whether performance improvement is still needed. If performance improvement is needed, the splitting calculation device returns to 415 to select an adjustment point. If performance improvement is not needed, the splitting calculation device performs 445 to reduce repeated transport. At 450, the splitting calculation device determines whether the performance does not decrease. If the performance does not decrease, it returns to 445 to reduce repeated transport and continue optimization. Otherwise, optimization can no longer be performed to obtain the optimal tiling strategy. In this way, the embodiments of the present application can determine the reasonable initial state of the operator splitting algorithm by determining the number of splits and the required number of cache blocks for executing at least one layer of loop of the operator, obtain the most time-consuming and adjustment-required stage of the operator by selecting an adjustment point in the pipeline containing the operator, and adjust at least one of the number of splits and the number of cache blocks based on this adjustment point, optimize the number of operator splits and the required cache space by step-by-step tuning, and ultimately improve computing efficiency and save cache resources.

[0087] In the following examples, Figure 4 The implementation of the five modules in the algorithm flow is "setting the splitting axis and the number of splits", "selecting the adjustment point", "reducing the number of splits for the transport task", "increasing the number of cache blocks for the transport task", and "reducing the number of splits for the computing task".

[0088] Figure 5 Schematic diagram of calculating the split axis and the number of splits in the embodiment of the present application. Figure 4 In "Setting the Splitting Axis and the Number of Splits" 410, a parameter M / N / K for the number of splits that each axis can work on is calculated, that is, the number of cycles that each loop can work on. Under this parameter M / N / K for the number of splits, the cache block can be accommodated in the global memory such as the Random Access Memory (RAM) and caches at all levels, ensuring that it will not exceed the buffer tolerance and the system can work normally. This setting serves as the basis for subsequent operator split optimization. Moreover, in the subsequent optimization, if the buffer capacity is exceeded, the axis splitting function in embodiment 500 will be called to reduce the amount of data transported each time to ensure that the buffer capacity limit is not exceeded.

[0089] 505 corresponds to the initial state of the splitting strategy. In the initial state, the used buffer may exceed the capacity limit of the buffer buffer and cause overflow, which needs to be adjusted. At 510, the splitting calculation device evaluates the occupancy in the buffer buffer. The buffer occupancy evaluation returns the occupancy rate of each buffer, as shown in Table 1. At 515, the splitting calculation device determines whether the constraint of the buffer capacity limit is met. The occupancy rate of all buffers is not higher than 100% to meet the constraint. The data that does not meet the constraint in Table 1 are the buffer occupancy rates of L1 and L0C.

[0090] Table 1. Cache occupancy at each level

[0091] L1 200% L0A 80% L0B 100% L0C 300%

[0092] If all buffers are satisfied, then end 520. If not satisfied, the slicing calculation device calculates the benefit-cost ratio, that is, executes 525 to set the priority of the sub-axis according to the association between the sub-axis and the buffer, and 530 to find the sub-axis with the lowest slicing cost in the highest priority sub-axis. In the embodiment of the present disclosure, the sub-axis can be a repeated axis or an independent axis. In the embodiment of the present disclosure, calculating the benefit-cost ratio can include: calculating the slicing benefit, calculating the slicing cost, and calculating the cost performance.

[0093] First, calculate the benefits of splitting. The benefits of splitting different sub-axes in terms of buffer occupancy are different. For example, continuing to split the A axis may reduce the occupancy rate of a certain buffer from 300% to 180%, and continuing to split the B axis may only reduce the occupancy rate of this buffer from 300% to 250%. The continued splitting action on the B axis is called "inefficient splitting". This is because a certain buffer is occupied by different data such as tensors, and the tile size of different axes controls the buffer occupancy of different tensors. In order to avoid "inefficient splitting", it is necessary to quantitatively calculate the benefit value that can be brought by splitting a certain axis. The calculation method is shown in Table 2.

[0094] Table 2 Benefits of axis segmentation

[0095] Occupancy rate A when cut=x Occupancy rate B when cut = 2x Return A-max(B,1)if A>1 buffer0 3.0 1.8 1.2 Buffer1 1.2 0.8 0.2 Buffer2 0.9 0.9 0.0 Buffer3 1.0 0.9 0.0 1.4

[0096] Then, the splitting cost is calculated. If some axes have the same priority, such as the M axis and the N axis in the above embodiment, the cost of splitting on these axes can be examined. The splitting cost of an axis is defined as the sum of the number of splits added to all nodes after increasing the number of splits. The splitting cost of each axis is different, which is related to the relationship between the axis and each application program interface API and the number of splits. In the embodiment of the present disclosure, the current splitting strategy is shown in Table 3.

[0097] Table 3 The cost of axis segmentation

[0098] M K N Number of parts CopyDataA 10 2 20 CopyDataB 2 2 4 LoadDataA 10 2 20 LoadDataB 2 2 4 Mmad 10 2 2 40 Fixpipe 10 2 20

[0099] For M and N, which have the highest current splitting priority, their splitting costs are 20+20+40+20=100 and 4+4+40+20=68. For example, since the splitting cost of the N axis is smaller than that of the M axis, the N axis with the smallest splitting cost can be selected to increase its number of splits.

[0100] Finally, calculate the cost-effectiveness R = slicing benefit / slicing cost, select the axis with the highest R as the specific axis, and increase its slicing number.

[0101] After the split calculation device calculates the benefit-cost ratio, the step length of the increase in the number of splits is determined at 535. Because the split cost of the axis is related to the current number of splits of each axis, the step length of the increase in the number of splits should not be too large, otherwise it may go too far along the optimal adjustment direction of the current point and be far away from the global optimal point. On the other hand, considering the algorithm execution efficiency, the step length should not be too small. The current step length setting strategy is:

[0102] a. Find the current occupancy rate of all over-limit buffers affected by the current sub-axis, take the minimum value, and record it as X: In the above embodiment, the over-limit buffers related to the N-axis are L1 (200%) and L0C (300%), so X = 2;

[0103] b. Execute pseudocode: if X>=2, then newCut=cut*2, else newCut=max(cut*X,cut+1).

[0104] After determining the increase step size of the number of splits and adjusting the number of splits, return to the buffer occupancy evaluation 510. If the buffer constraint is satisfied, then the acquisition of the appropriate number of splits and the corresponding number of cache blocks is terminated, otherwise the adjustment continues. In this way, the embodiment of the present application can adjust the specific axis with the best cost-effectiveness for the over-limit cache and the cost-effectiveness obtained by adjusting each axis under the condition that the operator's cache exceeds the capacity limit, increase the number of splits and thereby reduce the number of cache blocks required, until the operator's cache is not within the capacity limit. In this way, the appropriate initial state can be determined for the adjustment of the operator's number of splits and cache, which is beneficial to the subsequent optimization process of operator splitting.

[0105] The present disclosure can support Figure 4 The optimization goal of the algorithm is to minimize the execution time of the most time-consuming task on the busiest pipeline, and it is necessary to iteratively process the adjustment point tasks on the pipeline. The main process of the algorithm can be determined by the selection and analysis of the adjustment points. In this embodiment, the pipeline histogram is used to estimate the operator performance under a specific Tiling strategy.

[0106] exist Figure 2D , Figure 2E In the embodiment of the disclosure, two different hardware architectures are described. In the embodiment of the present disclosure, they are abstracted as a "stage", which means that no matter it is computing or data handling, they are all stages in the asynchronous pipeline parallel computing of operators. The computing engine can be represented as computing stages such as "cube", "vector", and "scalar"; the data handling pipeline can be represented as "pipe0", "pipe1"... multiple pipeline stages. This abstraction ignores the composition relationship between the hardware and looks at the different devices of the system on chip from a "flattened" perspective. With the above hardware abstraction, it is possible to draw a pipeline structure for a given segmentation strategy, such as Fig. 6AAs shown. The horizontal axis in embodiment 600 represents the calculation time, and the vertical axis represents the related hardware "stage". In the pipeline cube_MTE2 605, there are operators A1 610, A1 615, B1 617, etc., and in the pipeline cube_MTE1 620, there are operators A2 625, A2 630, etc. There is a logical sequence relationship 635 between operator A1 610 and operator A2 625. Since calculating such a pipeline requires analyzing the dependencies between tasks, when the split score is large and the number of tasks is large, the time-consuming analysis will seriously affect the efficiency of the algorithm. Therefore, this scheme adopts the pipeline histogram to estimate the performance, which intuitively removes the gaps (bubbles) between tasks at each stage. This calculation method can avoid analyzing the dependencies between tasks and is faster, such as Figure 6B As shown. In the histogram 640, pipelines are used to represent the operator execution time after removing the gaps between tasks. Pipeline pipe1 645 includes operators A1 650, A1 655 and B1 675, and pipeline pipe3&pipe4 includes operators A2 665 and A2 670. Pipeline pipe1 645 corresponds to pipeline cube_MTE2 605, and operators A1 650, A1655 and B1 675 correspond to operators A1 610, A1 615 and B1 617, respectively. Pipeline pipe3&pipe4 660 corresponds to pipeline pipe3&pipe4 660, and operators A2 665 and A2 670 correspond to operators A2 665 and A2 670, respectively. In the histogram 640 for removing gaps between tasks, the longest pipeline pipe1 645 can be selected, and then the most time-consuming operator B1 675 of pipeline pipe1 645 can be selected as the adjustment point. In this way, the embodiment of the present application can find the most time-consuming operator that needs to be adjusted by finding the longest stage of the longest item in the histogram, and reasonably set the adjustment point. By selecting the appropriate operator for adjustment, the overall optimization efficiency of the pipeline is improved.

[0107] Figure 7 This is a flow chart for reducing the number of parts for a transport task in an embodiment of the present application.

[0108] In the disclosed embodiment, the corresponding adjustment point for entering the processing flow is a transfer task and there is repeated data transfer. First, reduce the splitting score of the relevant repeated axis at 710. If there are multiple repeated axes, select the repeated axis that affects the largest amount of repeated data transfer. The adjustment step is 1, and the two rounding ups are satisfied at the same time, that is, cut = ceil (dim / ceil (dim / cut)), that is, cut can divide dim, so as to meet the integrity of data transfer. At this time, if the buffer constraint is satisfied at 715, the adjustment is successful and returns to 720, otherwise, the number of cache blocks is adjusted at no cost at 725. Because the number of cache blocks affects the buffer space occupancy, the more the number, the larger the occupied space. If the buffer exceeds the limit after adjusting the splitting score, the number of cache blocks can be appropriately reduced. At this time, the buffer constraint judgment is performed again at 730. If it still exceeds the limit, try "increasing the splitting score of the own axis" 745 and "adjusting the number of cache blocks with cost" 750 respectively. The way to increase the number of splits of the own axis by 740 is to freeze all duplicate axes and then call Figure 5 The slicing function of the illustrated embodiment. "No cost" and "cost" are both evaluations of the operator slicing optimization scheme on whether this adjustment will incur additional overhead. If no additional overhead is incurred, it is "no cost", and if additional overhead is incurred, it is "costly". First make the "no cost" adjustment and optimization, and there is no need to balance all aspects at this time. Then make the "cost" adjustment and optimization to balance the additional overhead. The "cost" can be comprehensively calculated by the operator slicing optimization scheme. In this way, the embodiment of the present application combines the costless adjustment and the costly adjustment to optimize the operator slicing. The adjustment of the number of cache blocks with and without cost is relatively complicated, and the processing flow in the dotted box is expanded as follows: Figure 8 The processing flow is shown.

[0109] When reducing the number of repeated axis splits causes buffer overflow, it is necessary to reduce the number of cache blocks in the cache space related to the overflow buffer. The specific processing details include the following steps:

[0110] First, the split computing device finds all cache spaces involved in the overflow buffer, and iteratively attempts to reduce the number of cache blocks in the cache space until the buffer no longer overflows or is reduced to the lower limit of the number of cache blocks in each cache space (such as minimax buffer block num=2);

[0111] Then, the splitting calculation device calculates the cost-performance list of each cache space. Among them, cost-performance = benefit / cost. Benefit refers to the size of buffer space saved by one-step adjustment. Cost refers to the impact on the cache effect, that is, the actual increase in the amount of repeated data transfer.

[0112] Then, the split calculation device sets the adjustment step size: 1) cyclic step = C j -C i (i, j is the difference between the two cycle periods), where C 0 =0; or 2) fixed step = 1.

[0113] Figure 8 This is a flowchart of costless cache adjustment and costly cache adjustment in the embodiment of the present application, corresponding to Figure 7 The dotted box part.

[0114] In the disclosed embodiment, the costless adjustment process is as follows: After initialization 805, at 810, select the adjustable cache space, that is, select the cache space with the initial number of cache blocks greater than the lower limit. At 815, determine whether there is no candidate cache space. If there is no candidate adjustable cache space, fail and exit 820. First, at 825, adjust each cache space costlessly, give priority to the one with high benefit, and skip if there is no cache space. At 830, if the buffer is determined to be satisfied, exit successfully at 835, otherwise continue to execute the subsequent process.

[0115] In the disclosed embodiment, the cost adjustment process is as follows: at 840, the candidate adjustable cache spaces are screened, that is, the queues whose current number of cache blocks is greater than the lower limit are screened. If there is no candidate adjustable cache space, it fails and exits at 850. Then, the cost performance of each cache space currently adjusted is obtained from the cost performance list. If the benefits of all cache spaces are less than remain, the cache space with the highest cost performance is adjusted at 860 (if the cost performance is the same, the one with higher benefit is selected), and remain is updated, and the aforementioned "find all cache spaces involved in the overflow buffer" is returned. If there are some cache spaces with over-benefit, that is, the benefit is greater than or equal to remain, the cache space with the lowest cost is selected for adjustment at 865, and then the process returns successfully.

[0116] In this way, the embodiment of the present application combines costless adjustment with costly adjustment to optimize operator segmentation, minimize costs, increase benefits, and achieve better optimization effects.

[0117] Fig. 9 It is a schematic diagram of a costless adjustment cache and a costly adjustment cache in another embodiment of the present application.

[0118] In embodiment 900, it is assumed that there are two cache spaces A and B, the cache mode of A is swappable, and the cache mode of B is fixed, SliceA=2, that is, the cache block size of A is 2, SliceB=10, that is, the cache block size of B is 10, the initial number of cache blocks is 12, the minimum number of cache blocks required is 2, and a loop pattern is adopted. A and B are both divided into: 0123|0123|4567|4567||0123|0123|4567|4567, and the adjustment target is target=remain=60, that is, the number of cache overflows is 60. C0=0, C1=4, that is, in 0123|0123, the minimum repetition period of 0 is 4. C2=8, that is, in 0123|……|4567|, 01234567 is non-repetitive, and there are 8 blocks in total.

[0119] CA=2 is the number of repetitions of the maximum repetition cycle, which is the product of all outer axes of the outermost repetition axis. CB=4 is the number of repetitions of a cache slice, that is, the product of all axes.

[0120] The first row of the table in embodiment 900 is the number of blocks that are reduced successively, and the second and third rows are the benefits and costs generated if adjustments are made in cache space A. The fourth and fifth rows are the benefits and costs generated if adjustments are made in cache space B. For example, the circle in 905 is the actual position adjusted for reducing the cache capacity in each step. That is, the optimization result is: cache space A is adjusted in steps 1, 2, 3, 4, and 6, and cache space B is adjusted in steps 5 and 7.

[0121] For cache space B, since it is a fixed cache and 01234567 only needs to occupy 8 blocks, the number of cache blocks of B decreases from 12. The adjustments in the first four steps (circles numbered 1-4) are free of cost, and each time will generate a net benefit of 10 bytes, or 1 block.

[0122] The gain from the adjustment in step 6 of B is 10 bytes, but since one block in 01234567 is no longer cached, a cache miss may occur, and the cost is 10 bytes*(C B -1) times. The "-1" here is because even if there are more than 8 cache blocks, 01234567 are all cached, there is no cache miss, and there is also a process of initial data moving in, which has a cost of 1. Therefore, the actual cost after a cache miss is C B -1.

[0123] For cache space A, A uses swappable cache. Since 01234567 only takes up 8 cache blocks out of 12, reducing 1-4 blocks from 12 blocks is free. Reducing the 5th block incurs a cost of 8*(C A -1). Since the swappable cache is used, reducing 5 blocks and reducing 6-8 blocks are the same cache miss, and the subsequent reduction of 6-8 blocks will not incur an additional cost compared to reducing the 5th block, so the cost is 0. The 5th step of cache space A adjusts 1, 2, 3, and 4 blocks in total, and the 7th step adjusts 5, 6, 7, and 8 blocks in total. Therefore, one adjustment of cache space A is to reduce 4 blocks.

[0124] In the disclosed embodiment, after the first five adjustment steps (steps 1-4 adjust cache space A, and step 5 adjusts swap space B), the cost-free adjustment benefit is obtained. After five adjustments, the remaining overflow amount remains = 60-48 = 12 (the reduced overflow amount 48 = 10*4in A+2*4in B, that is, the total benefit is 48).

[0125] For the sixth adjustment, select the adjustment position in cache space A and B. For A, step size step = C2-C1 = 4, benefit is 2*4 = 8, cost is 16*2 = 32, and cost performance is 0.25. For B, step = 1, benefit is 10, cost is 10*4 = 40, and cost performance is 0.25; because the cost performance is the same and B has a higher benefit, B is selected for adjustment, corresponding to the circle with sequence number 6, remain = 12-10 = 2.

[0126] For the seventh adjustment, choose between A and B. At this time, A's over-benefit = 8-2>0, B's over-benefit = 10-2>0, both are over-adjustments, satisfying the requirement that the cache will not overflow, so we no longer consider the cost-effectiveness, but look for the one with the lowest cost. costA = 32, costB = 40, A has the lowest cost, so we choose to adjust A, corresponding to the circle number 7.

[0127] In the embodiment of the present disclosure, the method of embodiment 900 is adopted. For each adjustment, a comparison can be made between cache spaces A and B each time, and a cache space with low cost and high benefit can be selected for adjustment, thereby optimizing the adjustment of the cache space, saving the use of the cache space, and improving the calculation efficiency of the operator. It can be understood by those skilled in the art that the above method can also be applied to the optimization selection of three or more cache spaces, and the present disclosure does not limit this.

[0128] Fig.10 This is a flowchart of increasing the number of cache blocks for a transport task in an embodiment of the present application. Embodiment 1000 corresponds to Figure 4 of 425.

[0129] The cache space can temporarily cache data that needs to be repeatedly moved. When the data is accessed again, if the cache hit is enough, the amount of movement can be reduced. Therefore, this embodiment can increase the number of cache blocks in a certain cache space. At 1010, the splitting calculation device first finds the operator or node with the largest amount of repeated movement, and then selects a cache space with a larger cache block size to adjust the number of its cache blocks at 1015. There are two possibilities for the adjustment step size. In the cyclic mode, it is the minimum adjustment amount that can increase the cache hit, and in the fixed cache mode, it is 1. Then at 1020, it is determined whether the cache exceeds the limit and overflows. If it overflows and the constraints are not met, then at 1025, the split score of the axis that does not affect the cache hit is further modified by freezing the node's own axis and then calling Figure 5 The block function. In this way, for the nodes with the most repeated transport, the number of cache blocks of the transport operator can be increased, and the step size can be adjusted according to different cache modes to increase the cache hit rate of the operator. In order to prevent the capacity from exceeding the limit due to increasing the number of cache blocks, the transport operator's own axis can be frozen and the number of splits of the repeated axis of the transport operator can be increased for fine-tuning. Finally, the number of cache blocks of the transport operator can be increased without exceeding the capacity limit, thereby improving the cache hit rate of the transport operator and the execution efficiency of the transport operator.

[0130] Fig.11 This is a flowchart of reducing the number of computing task splits in an embodiment of the present application. Embodiment 1100 corresponds to Figure 4 430 in the computation bound operator. In the computation bound operator, by reducing the number of repeated axis splits, “fewer splits” are implemented, which reduces the loop instruction overhead and improves the efficiency of this operator.

[0131] For computing tasks, reducing their split fractions can effectively improve the efficiency of computing instructions. Therefore, the innermost sub-axis of its own sub-axis is selected for splitting, and the step size is adjusted to 1, while satisfying the unchanged rounding up twice, that is, cut = ceil (dim / ceil (dim / cut)). Specifically: at 1105, reduce the number of splits of the computing task node. At 1110, sort the computing tasks by execution time from long to short. At 1115, determine the API node with the longest computing task, that is, the operator with the longest computing task. At 1120, prioritize reducing the number of splits of the inner axis of the own axis from the inside to the outside. Then at 1125, determine whether the buffer exceeds the limit or overflows, or satisfies the buffer constraint. If overflow, execute at 1135 Figure 7 The costless cache space adjustment + the costly cache space adjustment in Figure 7The process in the dotted box. At 1140, if the buffer constraint is satisfied, 1145 is successfully returned, otherwise 1150 is returned as failure. In this way, the number of computing task partitions can be reduced efficiently. The disclosed embodiment reduces the number of partitions of the data computing operator with the longest execution time, and solves the problem of cache capacity excess by adjusting the number of cache blocks at no cost and at a cost in turn under the condition that the cache capacity may exceed the limit. Finally, a comprehensive optimization method is used to reduce the number of partitions of the data computing operator to achieve "less partitioning" and improve the computing efficiency of the data computing operator.

[0132] Fig.12 1 is a block diagram of a split calculation device 1200 in an embodiment of the present application. In some example embodiments, the device capable of executing method 200 may include a device for executing each step of method 200. The device may be implemented in any suitable form. For example, the device may be implemented in a circuit or a software module. Fig.12 Will combine Figure 2A to describe.

[0133] The split calculation device 1200 includes a split number and cache block number determination module 1210 , an adjustment point determination module 1220 , and an adjustment module 1230 .

[0134] The split number and cache block number determination module 1210 is configured to determine the split number and the required cache block number for executing at least one layer of loop of an operator, where the operator is a unit for executing instructions. The adjustment point determination module 1220 is configured to determine an adjustment point for adjusting the length of the pipeline including the operator. The adjustment module 1230 is configured to adjust at least one of the split number and the cache block number based on the adjustment point.

[0135] Fig.13 1 shows a schematic block diagram of an example computing device 1300 that can be used to implement an example implementation of the present disclosure. Fig.13 As shown, the split computing device 1300 includes: a bus 1302, a processor 1304, a memory 1306, and a communication interface 1308. The processor 1304, the memory 1306, and the communication interface 1308 communicate through the bus 1302. The split computing device 1300 can be a server, a storage device with computing capabilities, or a terminal device. It should be understood that the present disclosure does not limit the number of processors and memories in the split computing device 1300.

[0136] The bus 1302 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.13 The bus 1302 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 1302 may include a path for transmitting information between various components of the split computing device 1300 (for example, the memory 1306, the processor 1304, and the communication interface 1308).

[0137] The processor 1304 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0138] The memory 1306 may include a volatile memory, such as a random access memory (RAM). The processor 1304 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0139] The memory 1306 stores executable program codes, and the processor 1304 executes the executable program codes to respectively implement the functions of one or more of the aforementioned split number and cache block number determination module 1210, the adjustment point determination module 1220, and the adjustment module 1230, thereby implementing the operator splitting method according to the embodiment of the present disclosure. That is, the memory 1306 stores instructions for executing the splitting calculation method according to the embodiment of the present disclosure.

[0140] The communication interface 1308 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the split computing device 1300 and other devices or communication networks.

[0141] The disclosed embodiments also provide a computing device cluster, such as a database system. Fig.14A schematic block diagram of an example computing device cluster 1400 that can be used to implement an exemplary implementation of the present disclosure is shown. The computing device cluster 1400 includes at least one computing device. The computing device can be a server, such as a central server, an edge server, a storage device with computing capabilities, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop, a laptop, or a smart phone.

[0142] like Fig.14 As shown, the computing device cluster 1400 includes at least one split computing device 1300. The memory 1306 in one or more split computing devices 1300 in the computing device cluster 1400 may store the same instructions for executing the operator split computing method according to the embodiment of the present disclosure.

[0143] In some possible implementations, the memory 1306 of one or more split computing devices 1300 in the computing device cluster 1400 may also store partial instructions for executing the distributed lock method according to the embodiment of the present disclosure. In other words, the combination of one or more split computing devices 1300 can jointly execute the instructions for executing the distributed lock method according to the embodiment of the present disclosure.

[0144] It should be noted that the memory 1306 in different split computing devices 1300 in the computing device cluster 1400 may store different instructions. That is, the instructions stored in the memory 1306 in different split computing devices 1300 may implement the functions of one or more of the aforementioned split number and cache block number determination module 1210, adjustment point determination module 1220, and adjustment module 1230. In this way, the operator splitting method may be implemented in a distributed cluster and parallel computing manner to improve optimization efficiency.

[0145] In some possible implementations, one or more computing devices in the computing device cluster 1400 may be connected via a network, which may be a wide area network or a local area network. Fig.15 A possible implementation is shown. Fig.15As shown, the three computing devices 1300A, 1300B and 1300C are connected via a network. Specifically, the network is connected via the communication interface in each computing device. In this type of possible implementation, the memory 1306 in the computing device 1300A stores instructions for the functions of the module 1210 for determining the number of split shares and the number of cache blocks, the module 1220 for determining the adjustment point, and the module 1230 for adjusting. The memory 1306 in the computing device 1300B stores instructions for the functions of the module 1210 for determining the number of split shares and the number of cache blocks, the module 1220 for determining the adjustment point, and the module 1230 for adjusting. The memory 1306 in the computing device 1300C stores instructions for the functions of the module 1210 for determining the number of split shares and the number of cache blocks, the module 1220 for determining the adjustment point, and the module 1230 for adjusting. The split computing task can be completed in a distributed manner.

[0146] It should be understood that Fig.15 The functions of the computing device 1300A shown in FIG. 1 may also be completed by multiple computing devices 1300. Similarly, the functions of the computing device 1300B may also be completed by multiple computing devices 1300, and the functions of the computing device 1300C may also be completed by multiple computing devices 1300.

[0147] The present disclosure also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Fig.13 and Fig.14 The connection mode of the computing device cluster 1400. The difference is that the memory 1306 in one or more computing devices 1300 in the computing device cluster may store the same instructions for executing the method of splitting calculation according to the embodiment of the present disclosure.

[0148] In some possible implementations, the memory 1306 of one or more computing devices 1300 in the computing device cluster may also store partial instructions for executing the split calculation according to the embodiment of the present disclosure. In other words, the combination of one or more computing devices 1300 may jointly execute the instructions for executing the split calculation according to the embodiment of the present disclosure.

[0149] The embodiments of the present disclosure also provide a computer program product including instructions, which may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method and function of any of the above embodiments.

[0150] The embodiments of the present disclosure also provide a computer-readable storage medium, which may be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the method and function of any of the above embodiments.

[0151] In general, various embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be performed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as, by way of non-limiting example, hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof.

[0152] The present disclosure also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process / method as described above with reference to the accompanying drawings. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.

[0153] The computer program code for realizing the method of the present disclosure can be written in one or more programming languages. These computer program codes can be provided to the processor of general-purpose computer, special-purpose computer or other programmable data processing device, so that the program code, when being executed by computer or other programmable data processing device, causes the function / operation specified in flow chart and / or block diagram to be implemented. The program code can be executed completely on computer, partly on computer, as independent software package, partly on computer and partly on remote computer or completely on remote computer or server.

[0154] In the context of the present disclosure, computer program codes or related data may be carried by any appropriate carrier to enable a device, apparatus or processor to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, and the like. Examples of signals may include electrical, optical, radio, acoustic or other forms of propagation signals, such as carrier waves, infrared signals, and the like.

[0155] A computer-readable medium may be any tangible medium containing or storing a program for or related to an instruction execution system, apparatus or device, or a data storage device such as a data center containing one or more available media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any suitable combination thereof. More detailed examples of computer-readable storage media include an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0156] In addition, although the operation of the method of the present disclosure is described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flow chart can change the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of a device described above can be further divided into being embodied by multiple devices.

[0157] The various implementations of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those of ordinary skill in the art that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the various embodiments of the present invention. Without departing from the scope and spirit of the various implementations described, many modifications and changes are obvious to those of ordinary skill in the art. The choice of terms used in this article is intended to explain the principles of each implementation, practical application or improvement of technology in the market, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A data processing method, include: Determining the number of slices and the number of cache blocks required for executing at least one layer of loop of an operator, wherein the operator is a unit for executing instructions; Determining, in a pipeline including the operator, an adjustment point for adjusting a length of the pipeline; as well as Based on the adjustment point, at least one of the number of splits and the number of cache blocks is adjusted.

2. The method according to claim 1, wherein at least one layer of circulation of the operator corresponds to at least one axis, The shaft includes at least one of the following: Own axis, corresponding to a loop that does not repeat the movement of data; or The repeating axis corresponds to a cycle of repeated movement of data.

3. The method according to claim 1 or 2, further comprising: include: Obtain a computation graph, the computation graph including a mathematical computation relationship between the operators and a scheduling pipeline, the scheduling pipeline corresponding to at least one layer of loops of the execution operator, wherein: The mathematical calculation relationship and scheduling pipeline between the operators of the calculation graph are used to determine the number of divisions of at least one layer of loop used to execute the operator and the required number of cache blocks.

4. The method according to claim 3, in, The computation graph includes a directed acyclic graph, and the operators correspond to nodes of the directed acyclic graph. The mathematical calculation relationship between the operators corresponds to the topological connection relationship and node type of the directed acyclic graph, and the node type corresponds to the data handling operator and the data calculation operator. The scheduling pipeline corresponds to the attributes of the node, The obtaining of the computation graph comprises: The topological connection relationship of the directed acyclic graph, the node type and the node attribute are obtained.

5. The method according to claim 2, wherein the cache comprises at least one of the following: A swappable cache is used to determine cached data in a first-in, first-out manner; A fixed cache, which is used to cache specific data persistently; or A hybrid cache uses a swappable cache on an outer axis for accessing data of the hybrid cache, and uses a fixed cache on an inner axis for accessing data of the hybrid cache, wherein the outer axis corresponds to an outer loop after splitting the loop, and the inner axis corresponds to an inner loop after splitting the loop.

6. The method according to claim 5, wherein In a swappable cache, the number of cache hits is determined by the number of slices of the repeat axis and the number of slices of the own axis within the repeat axis.

7. The method according to claim 5, further comprising: include: The fixed cache is used under the condition that the number of cache blocks is less than the minimum cycle period of repeated data access. The hit count of the fixed cache is determined according to the number of data blocks segmented by the operator, the number of data block repetitions, and the number of cache blocks.

8. The method according to claim 1 or 2, wherein the determining of the number of slices of at least one layer of loop of the execution operator and the number of cache blocks required include: Under the condition that the cache capacity limit is not exceeded, the number of slices of at least one layer of loop executing the operator and the number of cache blocks required are determined.

9. The method according to claim 8, wherein the number of slices of at least one layer of loop of the execution operator and the number of cache blocks required are determined without exceeding the cache capacity limit. include: Determining a cache occupancy rate for at least one cache and determining an excess cache based on the cache occupancy rate; Determine the benefit of increasing the number of shards of the axis on the cache occupancy of the over-limit cache; Determining the cost of increasing the number of slices of the axis; Determining cost-effectiveness based on the benefits and the costs; Determining a specific axis to increase the number of cuts according to the maximum value of the performance / price ratio; The number of slices of at least one layer of loop of the execution operator and the required number of cache blocks are determined according to the increased number of slices of the specific axis.

10. The method according to claim 9, wherein the step of determining the number of slices of at least one layer of loop for executing the operator and the number of cache blocks required without exceeding the cache capacity limit is also include: Determine the step size for increasing the number of splits according to the minimum value of the cache occupancy rate of the over-limit cache affected by the specific axis; as well as The increased number of slices of the specific axis is determined according to the step size of increasing the number of slices.

11. The method according to claim 1 or 2, wherein the step of determining an adjustment point for adjusting the length of the pipeline including the operator include: removing gaps in a pipeline including the operator to obtain a pipeline histogram; Obtain the longest item among multiple items in the histogram; The most time-consuming phase in the longest item is selected as the adjustment point.

12. The method according to claim 4, wherein adjusting at least one of the number of splits or the number of cache blocks comprises at least one of the following: Reducing the number of splits of the data handling operator; Increase the number of cache blocks of the data handling operator; or Reduce the number of splits of the data calculation operator.

13. The method according to claim 12, wherein reducing the number of splits of the data handling operator comprises at least one of the following: Reducing the number of divisions of the repeated axis of the data handling operator; Adjusting the number of cache blocks at no cost under the condition that the cache capacity limit is not met; Adjusting the number of cache blocks at a cost under the condition that the cache capacity limit is not met; or The number of splits of the own axis of the data handling operator is increased under the condition that the cache capacity limit is not met.

14. The method according to claim 13, wherein the reducing the number of splits of the repeated axis of the data handling operator include: Under the condition that the data handling operator has multiple repetitive axes, selecting a specific repetitive axis that has the greatest impact on the amount of data handled repeatedly; The number of segments of the specific repeating axis is reduced by adjusting the number of segments according to a step size of 1, and the number of segments of the specific repeating axis after adjustment satisfies the requirement of being rounded down twice and remaining unchanged.

15. The method according to claim 13, wherein the cost-free adjustment of the number of cache blocks include: Calculate a cost-performance list and adjustment step size of at least one cache; Selecting an adjustable cache according to the condition that the number of cache blocks is greater than a lower limit; as well as Reduce the number of cache blocks in the selected cache by the adjustment step size.

16. The method according to claim 15, wherein the amount of cache blocks is adjusted at a cost. include: Selecting an adjustable cache according to the condition that the number of cache blocks is greater than a lower limit; Under the condition of surplus, adjusting the number of cache blocks of the adjustable cache with low cost; Under the condition of no surplus, the number of cache blocks of the adjustable cache with high cost performance is adjusted.

17. The method according to claim 13, wherein the number of splits of the own axis of the data handling operator is increased under the condition that the cache capacity limit is not met. include: Freezing the repetitive axis of the data handling operator under the condition that the cache capacity limit is not met; Determining a cache occupancy rate for at least one cache and determining an excess cache based on the cache occupancy rate; Determine the benefit of increasing the number of shards of the own axis on the cache occupancy rate of the over-limit cache; Determining the cost of increasing the number of slices of the own axis; Determining cost-effectiveness based on the benefits and the costs; as well as The specific own axis for increasing the number of cuts is determined based on the maximum value of the cost-performance ratio.

18. The method according to claim 12, wherein the increasing the number of cache blocks of the transport operator include: Determine the transport operator with the most repeated transports; Increasing the number of cache blocks of the transport operator; as well as Under the condition that the cache capacity limit is not met, the own axis of the transport operator is frozen, and the number of divisions of the repeated axis of the transport operator is increased.

19. The method according to claim 12, wherein the reducing the number of splits of the data computing operator include: Determine the data calculation operator with the longest task execution time; Reducing the number of divisions of the inner axis of the own axis of the data calculation operator with the longest execution time; Adjusting the number of cache blocks at no cost under the condition that the cache capacity limit is not met; as well as The number of cache blocks is adjusted at a cost under the condition that the cache capacity limit is not met.

20. A data processing device, include: A module for determining the number of splits and the number of cache blocks, used to determine the number of splits and the number of cache blocks required for executing at least one layer of loop of an operator, wherein the operator is a unit for executing instructions; An adjustment point determination module, used to determine, in a pipeline including the operator, an adjustment point for adjusting the length of the pipeline, so as to reduce the length of the pipeline; as well as An adjustment module is used to adjust at least one of the number of splits or the number of cache blocks based on the adjustment point.

21. An electronic device, include: A processor, and a memory storing instructions, wherein when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 19.

22. A computer-readable storage medium storing instructions, wherein when the instructions are executed by an electronic device, the electronic device executes the method according to any one of claims 1 to 19.

23. A computer program product, comprising instructions, which, when executed by an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 19.