Processing Method, Device and Computer Readable Storage Medium of Computation Graph
By inserting global aggregation and global reduction operators into the calculation graph, using shared memory of multi-level hardware architecture for tensor segmentation, the problem of high communication bandwidth occupancy in distributed parallel computing is solved, and the computing performance is improved.
Patent Information
- Application Number
- CN202411784634.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-05
AI Technical Summary
In distributed parallel computing, the high communication bandwidth occupancy and low inference throughput of matrix multiplication tasks, especially when data exchange and synchronization are performed on multi-level GPU architectures.
By inserting appropriate levels of global aggregation and global reduction operators into the calculation graph, tensor slicing is used to utilize shared memory of multi-level hardware architecture to reduce communication bandwidth usage and data traffic.
Improves the throughput of model inference and reduces latency, reduces communication overhead and data handling.
Smart Images

Figure CN119624748B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chip technology, and more particularly, to a method and apparatus for processing a computation graph and a computer-readable storage medium. Background Art
[0002] With the development of the high-performance computing field, large-scale data processing is usually achieved by means of distributed parallel computing. Distributed parallel computing means that a large computing task is decomposed into multiple smaller subtasks, and then these subtasks are assigned to different computing units (such as CPUs, GPUs, TPUs, etc.) for parallel execution. Distributed parallel computing usually requires communication between computing units to exchange data and synchronize, and a computing unit (such as a graphics processing unit (GPU)) adopts a multi-level hardware architecture to achieve parallel processing of computing and communication. For a matrix multiplication computing task, the shared axis (i.e., the K axis) of the left matrix (M, K) and the right matrix (K, N) is usually sliced to obtain multiple matrix multiplication subtasks, and these matrix multiplication subtasks are assigned to a GPU with a multi-level architecture for parallel execution, which requires complex communication and data merging operations, resulting in high communication bandwidth occupancy and low inference throughput. Summary of the Invention
[0003] This application provides a method and apparatus for processing a computation graph and a computer-readable storage medium to reduce communication bandwidth occupancy and data communication volume, thereby improving the computing performance of a chip system.
[0004] First aspect, a method for processing a computation graph is provided, which is applied to a chip system. The chip system includes multiple computing units, and each computing unit includes multiple computing nodes and a shared memory among the multiple computing nodes. The processing method includes: obtaining an original computation graph of a model and determining a first operator and a second operator from each operator in the original computation graph, where the first operator and the second operator are two adjacent matrix multiplication operators in the original computation graph, and the second operator is a successor operator of the first operator; determining a splitting method for the first operator and the second operator according to a first number of computing units and a second number of computing nodes, and performing tensor splitting on the first operator and the second operator according to the corresponding splitting method, where the splitting method for the first operator is row splitting of an input matrix or column splitting of a weight matrix, the splitting method for the second operator is row splitting and column splitting of a weight matrix, the first number is an integer multiple of the number of computing units in the chip system; the second number is an integer multiple of the number of computing nodes in each computing unit; inserting a first-level global aggregation operator before the second operator and inserting a second-level global reduction operator and a first-level global aggregation operator after the second operator to obtain multiple target computation graphs, where the computing nodes are used to execute the target computation graphs, the first level is an operation among computing nodes within a computing unit, and the second level is an operation among the same computing nodes in different computing units.
[0005] In a possible implementation, the first-level global aggregation operator is used to perform an aggregation operation on the computation results of multiple computing nodes on each computing unit, and the second-level global reduction operator is used to perform a reduction operation on the computation results of the corresponding computing nodes on each computing unit.
[0006] In a possible implementation, the first-level global aggregation operator performs the aggregation operation by using the first shared memory.
[0007] In a possible implementation, the first shared memory includes multiple storage addresses corresponding to the multiple computing nodes one by one, the multiple storage addresses are consecutive in sequence, and a first computing node among the multiple computing nodes stores the one-to-one correspondence between the multiple computing nodes and the multiple storage addresses. The first-level global aggregation operator performing the aggregation operation by using the first shared memory includes: according to the correspondence, the multiple computing nodes respectively store multiple data into the corresponding multiple storage addresses; the first computing node among the multiple computing nodes combines the multiple storage addresses as the target storage address of the aggregation result corresponding to the aggregation operation, and sends the target storage address to other computing nodes; the multiple computing nodes read the aggregation result according to the target storage address.
[0008] In a possible implementation, the first shared memory includes a plurality of storage addresses corresponding to the plurality of computing nodes one by one, and the plurality of computing nodes store the one-to-one correspondence between the plurality of computing nodes and the plurality of storage addresses. The first-level global aggregation operator uses the first shared memory to perform an aggregation operation, including: according to the correspondence, the plurality of computing nodes respectively store a plurality of data to the plurality of storage addresses; the plurality of computing nodes merge the plurality of storage addresses as the target storage address of the aggregation result corresponding to the aggregation operation, and update their respective corresponding storage addresses to the target storage address; according to the target storage address, the plurality of computing nodes respectively read the aggregation result.
[0009] In a possible implementation, the chip system includes a data transmission unit, the plurality of computing units are connected through the data transmission unit, the data transmission unit includes a second shared memory, and the second-level global reduction operator uses the second shared memory to perform a reduction operation.
[0010] In a possible implementation, the second-level global reduction operator uses the second shared memory to perform a reduction operation, including: the first computing unit among the plurality of computing units reads the computing results corresponding to the same computing node in the plurality of computing units from the second shared memory; determines a reduction result according to the computing results respectively corresponding to the plurality of computing units; stores the reduction result to the second shared memory.
[0011] In a possible implementation, determining the partitioning methods of the first operator and the second operator according to the first number of computing units and the second number of computing nodes and performing tensor partitioning on the first operator and the second operator according to the corresponding partitioning methods includes: determining a third number according to the first number of computing units and the second number of computing nodes, where the third number is the product of the first number and the second number; performing row partitioning on the input matrix of the first operator according to the third number to obtain third number of input sub-matrices or performing column partitioning on the weight matrix of the first operator to obtain third number of weight sub-matrices; performing row partitioning on the weight matrix of the second operator according to the first number and performing column partitioning on the weight matrix of the second operator according to the second number to obtain third number of weight sub-matrices.
[0012] In a possible implementation, row-splitting the weight matrix of a second operator according to a first number and column-splitting the weight matrix of the second operator according to a second number to obtain a third number of weight sub-matrices includes: row-splitting the weight matrix of the second operator according to the first number to obtain a first number of intermediate weight sub-matrices; column-splitting the first number of intermediate weight sub-matrices according to the second number to obtain a third number of weight sub-matrices; or, column-splitting the weight matrix of the second operator according to the second number to obtain a second number of intermediate weight sub-matrices; row-splitting the second number of intermediate weight sub-matrices according to the first number to obtain a third number of weight sub-matrices. In a possible implementation, the first operator includes a plurality of juxtaposed third operators. The input matrices of the plurality of third operators are the same, and the plurality of third operators each have their own weight matrix, and the splitting methods of the plurality of third operators are the same.
[0013] In a second aspect, a processing device for a computational graph is provided, which is applied to a chip system. The chip system includes a plurality of computing units, and each computing unit includes a plurality of computing nodes and a shared memory between the plurality of computing nodes. The processing device includes: an operator determination module, configured to obtain an original computational graph of a model and determine a first operator and a second operator from each operator in the original computational graph, where the first operator and the second operator are two adjacent matrix multiplication operators in the original computational graph, and the second operator is a successor operator of the first operator; a tensor splitting module, configured to determine splitting methods of the first operator and the second operator according to a first number and a second number and perform tensor splitting on the first operator and the second operator according to the corresponding splitting methods, where the splitting method of the first operator is row-splitting of the input matrix or column-splitting of the weight matrix, the splitting method of the second operator is row-splitting and column-splitting of the weight matrix, the first number is an integer multiple of the number of computing units in the chip system; the second number is an integer multiple of the number of computing nodes in each computing unit; a computational graph processing module, configured to insert a first-level global aggregation operator before the second operator and insert a second-level global reduction operator and a first-level global aggregation operator after the second operator to obtain a plurality of target computational graphs, where the computing nodes are used to execute the target computational graphs, the first level is an operation between computing nodes within a computing unit, and the second level is an operation between the same computing nodes in different computing units.
[0014] In a third aspect, a computer-readable storage medium is provided, on which program code for executing the method according to the first aspect or any implementation manner of the first aspect is stored.
[0015] In a fourth aspect, a computer program product is provided, which includes program code for executing the method according to the first aspect or any implementation manner of the first aspect.
[0016] In the method and apparatus for processing a computational graph according to the embodiments of the present application, the tensor slicing method of the original model and the insertion of communication operators at appropriate levels are determined by using the hardware information of a multi-level hardware architecture. Since the communication operators use the shared memory of the multi-level architecture to implement corresponding operations, the communication bandwidth occupancy and data communication volume can be reduced, the throughput of model inference can be improved, and the latency of model inference can be reduced.
[0017] Furthermore, the global aggregation operation at the first level uses the first shared memory of multiple computing nodes to perform the aggregation operation between multiple computing nodes. The multiple computing nodes read the corresponding storage addresses to implement the aggregation operation between multiple computing nodes, without the need for communication between the computing nodes, reducing data transfer and communication overhead.
[0018] Furthermore, the global reduction operation at the second level uses the second shared memory of multiple computing units to perform the reduction operation on the corresponding computing nodes in multiple computing units, which can eliminate the need for communication between the computing units, reducing communication overhead and data communication volume. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the following described drawings are only some embodiments of the present application, and other drawings obtained by those of ordinary skill in the art based on these drawings fall within the scope of the present application.
[0020] Figure 1 is an example diagram of the hardware structure of a multi-level architecture provided by an embodiment of the present application.
[0021] Figure 2 is a schematic flowchart of the method for processing a computational graph provided by an embodiment of the present application.
[0022] Figure 3 is a schematic diagram of the original model adopting a traditional slicing strategy in related technologies.
[0023] Figure 4 is a schematic diagram of the original computational graph being sliced to obtain a target computational graph provided by an embodiment of the present application.
[0024] Figure 5 is a schematic flowchart of step S220 in the method for processing a computational graph provided by an embodiment of the present application.
[0025] Figure 6 is a schematic flowchart of step S530 provided by an embodiment of the present application.
[0026] Figure 7 is another schematic flowchart of step S530 provided by an embodiment of the present application.
[0027] Figure 8 It is a schematic flowchart of performing an aggregation operation using the shared memory provided by an embodiment of the present application.
[0028] Figure 9 It is another schematic flowchart of performing an aggregation operation using the shared memory provided by an embodiment of the present application.
[0029] Figure 10 It is a schematic flowchart of performing a reduction operation using the shared memory provided by an embodiment of the present application;
[0030] Figure 11 It is a schematic structural diagram of a processing device for a computational graph provided by an embodiment of the present application. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained based on the embodiments in the present application belong to the scope protected by the present application.
[0032] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0033] It should be understood that the specific embodiments described below are only used to explain the present application and are not used to limit the present application.
[0034] Figure 1 It shows a schematic diagram of a multi-level hardware architecture provided by an embodiment of the present application. As Figure 1 shown, the chip system 10 (chip) includes multiple computing units, and each computing unit includes multiple computing nodes. The embodiments of the present application do not specifically limit the number of computing units included in the chip system 10 and the number of computing nodes included in each computing unit. Figure 1Eight computing units are shown, and each computing unit includes four computing nodes. However, the number of computing units included in the chip system 10 in the embodiments of the present application and the number of computing nodes included in each computing unit are not limited thereto. The chip system 10 may include a greater or smaller number of computing units, and each computing unit may include a greater or smaller number of computing nodes. Each computing node may include a memory for storing data corresponding to the computing node. The computing node may also have a computing function. In some cases, a computing node may be referred to as a cluster. The computing unit here may be, for example, a chip supporting compute-in-memory (CIM). In some cases, a chip supporting compute-in-memory may exist in the form of a die. A die can be understood as a chip that is cut from a wafer and contains a complete circuit but has not been packaged yet. The chip system here may be, for example, a GPU. The chip system mentioned in the embodiments of the present application may be an independent chip or a part of a certain chip. In the related art, in order to improve the computing efficiency, a neural network model is usually split into multiple original models, and then these original models are assigned to different computing units (such as CPUs, GPUs, TPUs, etc.) for parallel execution, and each original model is executed on a chip system.
[0035] The original model includes a large number of matrix multiplication computing tasks. For the computing tasks of the original model, they will be split into multiple smaller subtasks, and these subtasks will be assigned to different computing nodes in different computing units for parallel execution. During the splitting process, the shared axis (i.e., the K axis) of the input matrix (M, K) and the weight matrix (K, N) is usually sliced to obtain multiple matrix multiplication subtasks, and these matrix multiplication subtasks are assigned to different computing nodes for parallel execution. In order to obtain the correct computing result, a multi-level global reduction operator needs to be inserted to communicate and compute the computing results on each computing node. For example, if the encoder (decoder) of a multi-layer attention (Transformer) model is used as the original model and assigned to a GPU card, and the original model is executed on the GPU card with a multi-level architecture by slicing the K axis of the matrix, then 8 computing units are required to perform global reduction operations between the computing nodes within their respective units to obtain 8 first reduction results, and then perform global reduction operations between the 8 computing units to obtain the reduction result as the final computing result. The data communication volume is large, and the overall communication operation logic is relatively complex.
[0036] To solve the above problems, a method for processing a computational graph is proposed. The hardware information of a multi-level hardware architecture is used to determine the tensor slicing method of the original model and insert communication operators at appropriate levels. Since the communication operators utilize the shared memory of the multi-level architecture to implement corresponding operations, it is possible to reduce the communication bandwidth occupancy and data communication volume, improve the throughput of model inference, and reduce the latency of model inference. It should be understood that the original model mentioned in the embodiments of the present application can be any neural network model.
[0037] Figure 2 FIG. is a schematic flowchart of the method for processing a computational graph provided by an embodiment of the present application. As Figure 2 shown, the method for processing the computational graph includes the following steps.
[0038] In step S210, the original computational graph of the model is obtained, and a first operator and a second operator are determined from each operator in the original computational graph, where the first operator and the second operator are two adjacent matrix multiplication operators in the original computational graph, and the second operator is the successor operator of the first operator.
[0039] In this embodiment, the chip system is used to execute the original computational graph of the model, such as the encoder (decoder) of a Transformer model. The chip system includes multiple computing units, each computing unit includes multiple computing nodes and shared memory between the multiple computing nodes, and the computing unit is used to execute each operator of the original model. The first operator and the second operator are determined by traversing each operator of the original model (see Figure 4 ), the first operator and the second operator are respectively matrix multiplication operators in the original model, and the second operator is the successor operator of the first operator. The first operator and the second operator are two adjacent matrix multiplication operators, and there can be any number of other element-wise operators or operators unrelated to slicing between the first operator and the second operator.
[0040] In a possible implementation manner, the first operator can be a single operator or include multiple parallel third operators, where the input matrices of each third operator are the same, and each third operator has its own corresponding weight matrix.
[0041] In step S220, the slicing method of the first operator and the second operator is determined according to the first number of computing units and the second number of computing nodes, and the first operator and the second operator are tensor sliced according to the corresponding slicing method.
[0042] In this embodiment, the splitting method of the first operator is row splitting of the input matrix or column splitting of the weight matrix. The first operator can be column splitting of the weight matrix or row splitting of the input matrix. If the number of rows of the input matrix is small and the split is less than the minimum number of rows that can be executed by the computing node, then column splitting of the weight matrix is performed. The first number is an integer multiple of the number of computing units in the chip system, and the second number is an integer multiple of the number of computing nodes in each computing unit.
[0043] In this embodiment, an example is given where the first number is the number of computing units in the chip system and the second number is the number of computing nodes in each computing unit, but it is not limited to this. Then, the number of matrix splits of the first operator is the same as the number of computing nodes in the chip system. For example, if the chip system includes 8 computing units and each computing unit includes 4 computing nodes, then the first number is 8, the second number is 4, and the weight matrix of the first operator is split into 32 pieces of weight data, or the input matrix of the first operator is split into 32 pieces of input data.
[0044] The splitting method of the second operator is row splitting and column splitting of the weight matrix. The second operator splits both the rows and columns of the weight matrix. For example, the number of row splits of the weight matrix of the second operator is the first number, and the number of column splits of the weight matrix of the second operator is the second number. For example, if the chip system includes 8 computing units and each computing unit includes 4 computing nodes, then the first number is 8, the second number is 4. First, the weight matrix of the second operator is row-split to obtain 8 pieces of intermediate weight data, and then each piece of intermediate weight data is column-split to obtain their respective 4 pieces of weight data. Or, the weight matrix of the second operator is column-split first to obtain 4 pieces of intermediate weight data, and then each piece of intermediate weight data is row-split to obtain their respective 8 pieces of weight data, finally obtaining 32 pieces of weight data.
[0045] Optionally, in some embodiments, refer to Figure 5 , the tensor splitting of the first operator and the second operator according to the corresponding splitting method includes the following steps.
[0046] In step S510, a third number is determined according to the first number of computing units and the second number of computing nodes, where the third number is the product of the first number and the second number.
[0047] In this embodiment, the first number is the number of computing units in the chip system, and the second number is the number of computing nodes in each computing unit. The third number is the number of computing nodes in the chip system. For example, if the chip system includes 8 computing units and each computing unit includes 4 computing nodes, then the first number is 8, the second number is 4, and the third number is 32.
[0048] In step S520, the input matrix of the first operator is row - sliced according to the third number to obtain the third number of input sub - matrices, or the weight matrix of the first operator is column - sliced according to the third number to obtain the third number of weight sub - matrices.
[0049] In this embodiment, the input matrix of the first operator is row - sliced to obtain the third number of input sub - matrices, or the weight matrix of the first operator is column - sliced to obtain the third number of weight sub - matrices to implement tensor slicing of the first operator.
[0050] In step S530, the weight matrix of the second operator is row - sliced according to the first number and column - sliced according to the second number to obtain the third number of weight sub - matrices.
[0051] In this embodiment, step S530 includes the following steps.
[0052] In step S610, the weight matrix of the second operator is row - sliced according to the first number to obtain the first number of intermediate weight sub - matrices.
[0053] In step S620, the first number of intermediate weight sub - matrices are column - sliced according to the second number to obtain the third number of weight sub - matrices.
[0054] In this embodiment, the weight matrix of the second operator can be first row - sliced and then column - sliced. The number of row - slices is the same as the first number, and the number of column - slices is the same as the second number. For example, if the chip system includes 8 computing units and each computing unit includes 4 computing nodes, then the first number is 8 and the second number is 4. The weight matrix of the second operator is first row - sliced to obtain 8 pieces of intermediate weight data, and then each piece of intermediate weight data is column - sliced to obtain its corresponding 4 pieces of weight data.
[0055] Optionally, in some embodiments, step S530 includes the following steps.
[0056] In step S710, the weight matrix of the second operator is column - sliced according to the second number to obtain the second number of intermediate weight sub - matrices;
[0057] In step S720, the second number of intermediate weight sub - matrices are row - sliced according to the first number to obtain the third number of weight sub - matrices.
[0058] In this embodiment, the weight matrix of the second operator can be first column-split and then row-split. The number of row-splits is the same as the first number, and the number of column-splits is the same as the second number. For example, the chip system includes 8 computing units, and each computing unit includes 4 computing nodes. Then the first number is 8, and the second number is 4. The weight matrix of the second operator is first column-split to obtain 4 pieces of intermediate weight data, and then each piece of intermediate weight data is row-split to obtain 8 pieces of weight data corresponding to each, and finally 32 pieces of weight data are obtained.
[0059] In step S230, a first-level global aggregation operator is inserted before the second operator, and a second-level global reduction operator and a first-level global aggregation operator are inserted after the second operator to obtain multiple target computation graphs. Among them, the computing nodes are used to execute the target computation graphs. The first level is the operation between the computing nodes within the computing unit, and the second level is the operation between the same computing nodes in different computing units.
[0060] In this embodiment, the original computation graph undergoes tensor splitting and insertion of communication operators at different levels to obtain the target computation graphs executed on each computing node. The first-level global aggregation operator is used to aggregate the computation results of the second number of computing nodes on each computing unit. The second-level global reduction operator is used to reduce the computation results of the corresponding computing nodes on each computing unit. Refer to Figure 4 , a first-level global aggregation operator is inserted before the second operator, that is, the computation results of multiple computing nodes in each computing unit are aggregated before the second operator. A second-level global reduction operator and a first-level global aggregation operator are inserted after the second operator, that is, the computation results of the computing nodes with the same number in different computing units for the second operator are globally reduced, and then each computing unit aggregates the data of multiple computing nodes inside it to obtain the final computation result.
[0061] Refer to Figure 4, on the chip system, the input matrix of the first operator is [32, 8192], and the weight matrix is [8192, 28672]; the weight matrix of the second operator is [32, 28672]. After splitting the first operator and the second operator according to the above splitting method, they are allocated to each computing node. On each computing node, the weight matrix of the first operator is [8192, 896], and the weight matrix of the second operator is [3584, 2048]. On each computing node, the respective target computation graph is executed. The output matrix of the first operator is [32, 896], and after a series of element-wise operator calculations, the output matrix remains [32, 896]. Since the first operator performs column splitting on the weight matrix, the calculation results can be aggregated to obtain the final result. Since each computing unit includes the shared memory of 4 computing nodes, therefore, a first-level global aggregation operator (all-gather operator) is inserted before the second operator, that is, an aggregation operation is performed between the computing nodes. In this way, the input matrix of the second operator on each computing node is [32, 3584]; the weight matrix of the second operator is [3584, 2048], and the output matrix of the second operator is [32, 2048]; since the weight matrix of the second operator performs row splitting, a global reduction operator needs to be inserted into the calculation results to obtain the correct result. Therefore, a second-level global reduction operator is inserted after the second operator, that is, a reduction operation is performed between the computing units. For example, the calculation results of 8 computing nodes 0 are globally reduced, the calculation results of 8 computing nodes 1 are globally reduced, the calculation results of 8 computing nodes 2 are globally reduced, and the calculation results of 8 computing nodes 3 are globally reduced. In this way, the matrix on each computing node is [32, 2048]; then, the data of 4 computing nodes in each computing unit are globally aggregated to obtain the final result [32, 8192] of the original computation graph.
[0062] Optionally, in some embodiments, the first-level global aggregation operator performs the aggregation operation using the shared memory. The shared memory includes a plurality of storage addresses corresponding to the plurality of computing nodes one by one. The plurality of storage addresses are consecutive in order, and the first computing node among the plurality of computing nodes stores the one-to-one correspondence between the plurality of computing nodes and the plurality of storage addresses. As Figure 8 shown, the first-level global aggregation operator performing the aggregation operation using the shared memory includes the following steps.
[0063] In step S810, according to the correspondence, the plurality of computing nodes respectively store a plurality of data to the corresponding plurality of storage addresses.
[0064] In this embodiment, the multiple pieces of data are the data corresponding one by one to the multiple computing nodes mentioned above. The multiple computing nodes respectively store their corresponding data in the corresponding storage addresses. The multiple storage addresses corresponding to the multiple computing nodes in the shared memory are consecutive in order. Exemplarily, the data corresponding to computing node 0, computing node 1, computing node 2, and computing node 3 are respectively denoted as d0, d1, d2, and d3. These four pieces of data are sequentially stored in the shared memory, and the sizes of these four pieces of data are the same. The starting address addr_id of each piece of data in the shared memory is determined according to the base address base_addr, the id parameter (program_id, 0 ≤ program_id ≤ 3) corresponding to each computing node, and the size data_size of the data, that is, the starting address addrs_id of each piece of data = base_addr + program_id * data_size. These four pieces of data are spliced according to the splitting dimension of the first computing operator before the aggregation operator. Exemplarily, if the first computing operator splits according to the row dimension and distributes the computing tasks to multiple computing nodes for computing, then the aggregation result is spliced according to the row dimension; if the first computing operator splits according to the column dimension and distributes the computing tasks to multiple computing nodes for computing, then the aggregation result is spliced according to the column dimension.
[0065] In step S820, the first computing node among the multiple computing nodes combines the multiple storage addresses as the target storage address of the aggregation result corresponding to the aggregation operation, and sends the target storage address to other computing nodes.
[0066] In this embodiment, the first computing node combines the multiple storage addresses as the target storage address of the aggregation result corresponding to the aggregation operation, and sends the target storage address corresponding to the aggregation result to other computing nodes. Exemplarily, computing node 0 combines the storage addresses of the four pieces of data into a target storage address, and sends this target storage address to computing node 1, computing node 2, and computing node 3. This target storage address corresponds to the storage location of the aggregation result.
[0067] In step S830, the multiple computing nodes read the aggregation result according to the target storage address.
[0068] In this embodiment, the aggregation result is the data obtained by splicing the data corresponding to the multiple computing nodes. The multiple computing nodes can read the aggregation result according to the target storage address. Exemplarily, computing node 0, computing node 1, computing node 2, and computing node 3 can read the aggregation result according to this target storage address.
[0069] This application stores data in the shared memory between computing nodes and stores the data at a predetermined storage address. By adjusting the storage address for the computing nodes to read the data, a gather operation is achieved. This process does not involve actual communication, which can reduce data movement and communication overhead.
[0070] Optionally, in some embodiments, the shared memory includes a plurality of storage addresses corresponding one-to-one to the plurality of computing nodes, and the plurality of computing nodes store the one-to-one correspondence between the plurality of computing nodes and the plurality of storage addresses.
[0071] See Figure 9 , the first-level global gather operator using the first shared memory to perform a gather operation includes the following steps.
[0072] In step S910, according to the correspondence, the plurality of computing nodes respectively store a plurality of data at a plurality of storage addresses.
[0073] In step S920, the plurality of computing nodes merge the plurality of storage addresses as the target storage address of the gather result corresponding to the gather operation, and update their respective corresponding storage addresses to the target storage address.
[0074] In step S930, according to the target storage address, the plurality of computing nodes respectively read the gather result.
[0075] Step S910 is the same as step S810, and step S930 is the same as step S830, which will not be elaborated here.
[0076] In step S920, each computing node stores the one-to-one correspondence between the plurality of computing nodes and the plurality of storage addresses. Each computing node merges the plurality of storage addresses as the target storage address of the gather result corresponding to the gather operation, and updates its own corresponding storage address to the target storage address. The target storage address corresponds to the storage location of the gather result, which can reduce the interaction of storage addresses between computing nodes. Exemplarily, computing nodes 0, 1, 2, and 3 all merge the storage addresses of 4 pieces of data into one target storage address, and update their respective corresponding storage addresses to the target storage address.
[0077] Optionally, in some embodiments, the chip system includes a data transmission unit. The plurality of computing units are connected through the data transmission unit. The data transmission unit includes a second shared memory. The second-level global reduction operator uses the second shared memory to perform a reduction operation.
[0078] See Figure 10, the global reduction operator at the second level performs a reduction operation using the second shared memory, which includes the following steps.
[0079] In step S1010, the first computing unit among the multiple computing units reads the computation results corresponding to the same computing node among the multiple computing units from the second shared memory.
[0080] In this embodiment, these multiple data are the data corresponding to the multiple computing units one by one as mentioned above. The data corresponding to the multiple computing units one by one are stored in the second shared memory.
[0081] In step S1020, a reduction result is determined according to the computation results corresponding to the multiple computing units respectively. After the first computing unit reads multiple data from the second shared memory, it can determine the reduction result corresponding to the reduction operation according to the multiple data. If the reduction operation is a summation operation, the reduction result can be the sum of the multiple data.
[0082] In step S1030, the reduction result is stored into the second shared memory. After obtaining the reduction result, the first computing unit can store the reduction result into the second shared memory.
[0083] Through the above steps, the reduction operation among multiple computing units in the chip system is completed by using the shared memory in the chip system. From the above description, it can be seen that during the execution of the reduction operation, the first computing unit can obtain the data corresponding to the other computing units through the second shared memory, rather than obtaining the data corresponding to the other computing units by communicating with the other computing units. In addition, the first computing unit enables the corresponding computing nodes in the other computing units to obtain the result of the reduction operation by writing the result of the reduction operation into the second shared memory, without sending the result of the reduction operation to the other computing units by communicating with the other computing units. Therefore, using the shared memory for the reduction operation can reduce the latency and communication overhead of the reduction operation.
[0084] The method for processing a computational graph provided by an embodiment of the present application determines the tensor slicing method of the original model and inserts appropriate communication operators by using the hardware information of a multi-level hardware architecture. Since the communication operator uses the shared memory of the multi-level architecture to implement corresponding operations, it can reduce the communication bandwidth occupancy and data communication volume, improve the throughput of model inference, and reduce the latency of model inference.
[0085] Furthermore, the global aggregation operation at the first level performs an aggregation operation among multiple computing nodes by using the first shared memory of the multiple computing nodes. The multiple computing nodes read the corresponding storage addresses to implement the aggregation operation among the multiple computing nodes, without communicating among the computing nodes, reducing data transfer and communication overhead.
[0086] Furthermore, the global reduction operation at the second level uses the second shared memory of multiple computing units to perform reduction operations on the corresponding computing nodes in the multiple computing units, which can avoid communication between computing units and reduce communication overhead and data traffic.
[0087] An embodiment of the present application also provides a processing device 110 for a computation graph, which is applied to the chip system 10 described in the above embodiment. The chip system includes multiple computing units, and each computing unit includes multiple computing nodes and a shared memory between the multiple computing nodes.
[0088] See Figure 11 , the processing device 110 includes an operator determination module 111, a tensor splitting module 112, and a computation graph processing module 113.
[0089] Among them, the operator determination module 111 is used to obtain the original computation graph of the model and determine a first operator and a second operator from each operator in the original computation graph. Among them, the first operator and the second operator are two adjacent matrix multiplication operators in the original computation graph, and the second operator is the successor operator of the first operator.
[0090] The tensor splitting module 112 is used to determine the splitting methods of the first operator and the second operator according to the first number and the second number, and perform tensor splitting on the first operator and the second operator according to the corresponding splitting methods. Among them, the splitting method of the first operator is row splitting of the input matrix or column splitting of the weight matrix, the splitting method of the second operator is row splitting and column splitting of the weight matrix, the first number is an integer multiple of the number of computing units in the chip system; the second number is an integer multiple of the number of computing nodes in each computing unit.
[0091] The computation graph processing module 113 is used to insert a first-level global aggregation operator before the second operator, and insert a second-level global reduction operator and a first-level global aggregation operator after the second operator to obtain multiple target computation graphs. Among them, the computing nodes are used to execute the target computation graphs. The first level is the operation between computing nodes within a computing unit, and the second level is the operation between the same computing nodes in different computing units.
[0092] In some possible implementation manners, the first-level global aggregation operator is used to perform an aggregation operation on the computation results of multiple computing nodes on each computing unit, and the second-level global reduction operator is used to perform a reduction operation on the computation results of the corresponding computing nodes on each computing unit.
[0093] In some possible implementation manners, the first-level global aggregation operator uses the shared memory between computing nodes to perform the aggregation operation.
[0094] In some possible implementation manners, the tensor splitting module 112 is specifically configured to: determine a third number according to a first number of computing units and a second number of computing nodes, where the third number is the product of the first number and the second number; perform row splitting on the input matrix of the first operator according to the third number to obtain the third number of input sub-matrices, or perform column splitting on the weight matrix of the first operator according to the third number to obtain the third number of weight sub-matrices; perform row splitting on the weight matrix of the second operator according to the first number and perform column splitting on the weight matrix of the second operator according to the second number to obtain the third number of weight sub-matrices.
[0095] In some possible implementation manners, the tensor splitting module 112 is specifically configured to: perform row splitting on the weight matrix of the second operator according to the first number to obtain the first number of intermediate weight sub-matrices; perform column splitting on the first number of intermediate weight sub-matrices according to the second number to obtain the third number of weight sub-matrices; or, perform column splitting on the weight matrix of the second operator according to the second number to obtain the second number of intermediate weight sub-matrices; perform row splitting on the second number of intermediate weight sub-matrices according to the first number to obtain the third number of weight sub-matrices.
[0096] In some possible implementation manners, the first operator includes a plurality of parallel third operators, the input matrices of the plurality of third operators are the same, and each of the plurality of third operators has its own weight matrix, and the splitting manners of the plurality of third operators are the same.
[0097] An embodiment of the present application further provides a computer-readable storage medium. Program code is stored on the computer-readable storage medium, and the program code can be used to execute the method for processing data in any of the foregoing embodiments.
[0098] An embodiment of the present application further provides a computer program product. The computer program product includes program code for executing the method for processing data in any of the foregoing embodiments.
[0099] It should be understood that in the embodiments of the present application, determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0100] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0101] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and back associated objects.
[0102] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0103] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0104] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0105] In addition, the functional units in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0106] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be read by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0107] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for processing a computational graph, characterized in that Applied to a chip system, the chip system includes a plurality of computing units and a data transmission unit. The plurality of computing units are connected through the data transmission unit. Each computing unit includes a plurality of computing nodes and a first shared memory between the plurality of computing nodes. The data transmission unit includes a second shared memory. The processing method includes: Obtain the original computation graph of the model and determine a first operator and a second operator from the respective operators of the original computation graph, where the first operator and the second operator are two adjacent matrix multiplication operators in the original computation graph, and the second operator is the successor operator of the first operator; Determine the partitioning methods of the first operator and the second operator according to the first number of computing units and the second number of computing nodes, and perform tensor partitioning on the first operator and the second operator according to the corresponding partitioning methods. The partitioning method of the first operator is row partitioning of the input matrix or column partitioning of the weight matrix. The partitioning method of the second operator is row partitioning and column partitioning of the weight matrix. The first number is an integer multiple of the number of computing units in the chip system; the second number is an integer multiple of the number of computing nodes in each computing unit; Insert a first-level global aggregation operator before the second operator and insert a second-level global reduction operator and a first-level global aggregation operator after the second operator to obtain a plurality of target computation graphs. The computing nodes are used to execute the target computation graphs. The first level is the operation between the computing nodes within the computing unit, and the second level is the operation between the same computing nodes in different computing units. The first-level global aggregation operator uses the first shared memory to perform an aggregation operation on the computation results of the plurality of computing nodes on each computing unit. The second-level global reduction operator uses the second shared memory to perform a reduction operation on the computation results of the corresponding computing nodes on each computing unit.
2. The processing method according to claim 1, wherein The shared memory includes a plurality of storage addresses corresponding to the plurality of computing nodes one by one. The plurality of storage addresses are consecutive in sequence, and the first computing node among the plurality of computing nodes stores the one-to-one correspondence between the plurality of computing nodes and the plurality of storage addresses. The first-level global aggregation operator using the first shared memory to perform the aggregation operation includes: According to the correspondence, the plurality of computing nodes respectively store a plurality of data to the corresponding plurality of storage addresses; The first computing node among the plurality of computing nodes combines the plurality of storage addresses as the target storage address of the aggregation result corresponding to the aggregation operation, and sends the target storage address to other computing nodes; The plurality of computing nodes read the aggregation result according to the target storage address.
3. The processing method according to claim 1, characterized in that, The shared memory includes a plurality of storage addresses corresponding to the plurality of computing nodes one by one, and the plurality of computing nodes store the one-to-one correspondence between the plurality of computing nodes and the plurality of storage addresses. The first-level global aggregation operator using the first shared memory to perform the aggregation operation includes: According to the correspondence, the plurality of computing nodes respectively store a plurality of data to the plurality of storage addresses; The multiple computing nodes merge multiple storage addresses as the target storage address of the aggregation result corresponding to the aggregation operation, and update their respective corresponding storage addresses to the target storage address; According to the target storage address, the multiple computing nodes respectively read the aggregation result.
4. The processing method according to claim 1, characterized in that, The global reduction operator at the second level uses the second shared memory to perform a reduction operation, including: The first computing unit among the multiple computing units reads the computing results corresponding to the same computing node among the multiple computing units from the second shared memory; Determine the reduction result according to the computing results corresponding to the multiple computing units respectively; Store the reduction result in the second shared memory.
5. The processing method according to claim 1, characterized in that, Determining the splitting methods of the first operator and the second operator according to the first number of computing units and the second number of computing nodes, and performing tensor splitting on the first operator and the second operator according to the corresponding splitting methods includes: Determine a third number according to the first number of computing units and the second number of computing nodes, where the third number is the product of the first number and the second number; Perform row splitting on the input matrix of the first operator according to the third number to obtain third number of input sub-matrices, or perform column splitting on the weight matrix of the first operator to obtain third number of weight sub-matrices; Perform row splitting on the weight matrix of the second operator according to the first number and perform column splitting on the weight matrix of the second operator according to the second number to obtain third number of weight sub-matrices.
6. The processing method according to claim 5, characterized in that, Performing row splitting on the weight matrix of the second operator according to the first number and performing column splitting on the weight matrix of the second operator according to the second number to obtain third number of weight sub-matrices includes: Perform row splitting on the weight matrix of the second operator according to the first number to obtain first number of intermediate weight sub-matrices; Perform column splitting on the first number of intermediate weight sub-matrices according to the second number to obtain third number of weight sub-matrices; Or, Perform column splitting on the weight matrix of the second operator according to the second number to obtain second number of intermediate weight sub-matrices; Perform row splitting on the second number of intermediate weight sub-matrices according to the first number to obtain third number of weight sub-matrices.
7. The processing method according to claim 1, characterized in that, The first operator includes multiple parallel third operators. The input matrices of the multiple third operators are the same, and each of the multiple third operators has its own weight matrix. The splitting methods of the multiple third operators are the same.
8. A processing device for a computational graph, characterized in that, Applied to a chip system, the chip system includes multiple computing units and a data transmission unit. The multiple computing units are connected through the data transmission unit. Each computing unit includes multiple computing nodes and a first shared memory between the multiple computing nodes. The data transmission unit includes a second shared memory. The processing device includes: An operator determination module, configured to obtain an original computation graph of a model and determine a first operator and a second operator from each operator in the original computation graph, where the first operator and the second operator are two adjacent matrix multiplication operators in the original computation graph, and the second operator is a successor operator of the first operator; A tensor splitting module, which is used to determine the splitting methods of a first operator and a second operator according to a first number and a second number, and perform tensor splitting on the first operator and the second operator according to the corresponding splitting methods. Among them, the splitting method of the first operator is row splitting of an input matrix or column splitting of a weight matrix, the splitting method of the second operator is row splitting and column splitting of the weight matrix, the first number is an integer multiple of the number of computing units in the chip system; the second number is an integer multiple of the number of computing nodes in each computing unit; A computation graph processing module, which is used to insert a first-level global aggregation operator before the second operator and insert a second-level global reduction operator and a first-level global aggregation operator after the second operator to obtain multiple target computation graphs. Among them, the computing nodes are used to execute the target computation graphs. The first level is the operation among the computing nodes within a computing unit, and the second level is the operation among the same computing nodes in different computing units; The first-level global aggregation operator uses a first shared memory to perform an aggregation operation on the computation results of multiple computing nodes on each computing unit, and the second-level global reduction operator uses a second shared memory to perform a reduction operation on the computation results of the corresponding computing nodes on each computing unit.
9. A computer-readable storage medium, characterized in that, Program code for executing the method according to any one of claims 1-7 is stored on the computer-readable storage medium.
Citation Information
Patent Citations
Computing method based on multiple bare chips and related equipment
CN117827419A
Data processing method and related device
CN119003957A