Large model compiling method and device based on near memory computing architecture

By splitting the operator weight parameters of the large model according to the number of tiles in the near-storage computing architecture and polling and passing them, the data communication delay problem between the calculation kernels is solved, and the efficient deployment and calculation of the large model in the multi-machine, multi-card, and multi-die system is realized, and the inference process is accelerated.

CN120371308APending Publication Date: 2025-07-25BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359145.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize near-existence computing architecture to optimize the allocation of computing resources in large models, resulting in data communication delay and cache consistency problems between computing cores, affecting computing efficiency.

Method used

The operator weight parameters of the large model are split according to the number of tiles in the chip unit, and polling and passing them between adjacent tiles to ensure the shape continuity of the input and output of the operator, and realize parallel calculations, and finally merge the results in the chip unit, and deploy them in the multi-machine, multi-card, and multi-die system.

Benefits of technology

It significantly improves the operator computing efficiency, accelerates the inference process of large models, and improves the universality and adaptability of chip systems in near-access computing architectures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371308A_ABST
    Figure CN120371308A_ABST
Patent Text Reader

Abstract

The invention discloses a large model compiling method and device based on a near memory computing architecture, the method and the device are applied to intelligent equipment, the intelligent equipment comprises deployed single chip units, and each chip unit comprises a plurality of tiles; the compiling method of the single chip unit comprises the following steps: in an operator compiling stage, splitting a dimension N of an operator weight parameter split by a large model according to the number of tiles in the chip unit; constraining the shape of the input and output sensor of the operator at the same time; performing polling transmission on the input tensor of the operator between adjacent tiles; after parallel calculation on all tiles in the chip unit is completed, obtaining an output result of an operator on each tile; the output results of all tiles are combined along the dimension of the tilen; and finally, converting and outputting the data according to a data arrangement rule in the chip unit. The method has the beneficial effects that the operator calculation efficiency is remarkably improved, and the reasoning process of a large model is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a large model compilation method and device based on a near-memory computing architecture. Background Art

[0002] The full name of a large model is a large language model (LLM, Large Language Model). Large-scale language models have the characteristics of a large number of parameters and a large amount of computation, and usually require more computing resources for inference models, such as computing resources and memory resources. For problems with a large amount of computation, parallel computing through multiple computing cores can be used to continue to accelerate. However, as the number of computing cores increases, data communication between the computing cores will lead to increased latency and cache coherence problems.

[0003] The near-memory computing architecture uses multi-core computing, and a local memory is configured for each computing core. Each computing core can directly access the local memory. There is shared memory between multiple computing cores, and each core can also access the local memory of other cores. At the same time, for the problems of large amount of computation and large number of parameters in large models, in order to minimize the latency caused by data transmission between cores and make the most of the characteristics of near-memory computing, that is, the memory access time depends on the position of the memory relative to the core, the speed of the core accessing the local memory is faster than that of the non-local memory. For this characteristic, how to map large models to this near-memory computing architecture is an urgent problem to be solved currently. Summary of the Invention

[0004] Aiming at the technical defects of the prior art, the purpose of the embodiments of the present invention is to provide a large model compilation method and device based on a near-memory computing architecture.

[0005] To achieve the above purpose, in a first aspect, the embodiments of the present invention provide a large model compilation method based on a near-memory computing architecture, which is applied to an intelligent device. The intelligent device includes a deployed single-chip unit, where the chip unit is represented by die, each chip unit contains multiple tiles, the tile represents the smallest allocation unit, and the smallest allocation unit contains a certain number of logic units, computing units, and storage units. The compilation method of the single-chip unit includes:

[0006] In the compilation stage of the operator, the dimension N of the operator weight parameters split from the large model is split according to the number of tiles in the chip unit; where the first dimension of the shape of the weight parameters corresponding to the operator must be tile_num, and the remaining dimensions are the parameter shapes allocated to each tile.

[0007] Constrain the shapes of the input and output tensors of the operator simultaneously, requiring that the first dimension of the input and output tensors must be tile_num to ensure the continuity of the operations between operators;

[0008] Poll and transfer the input tensor of the operator between adjacent tiles to ensure that the input tensor data of each tile can appear on other tiles and implement the complete calculation process of the operator;

[0009] After completing the parallel calculation on all tiles within the chip unit, obtain the output results of the operator on each tile;

[0010] Then merge the output results of all tiles along the tile_num dimension to obtain the output data with tile_num as the first dimension; where tile_num represents the number of tiles in a chip unit;

[0011] Finally, convert the tile_num dimension and other dimensions to the original dimension N through the data layout rule within the chip unit to generate the final output result.

[0012] As a preferred implementation manner of this application, when splitting, if the weight parameters cannot be evenly distributed to each tile in dimension N of the original parameters, it is necessary to first pad dimension N of the original parameters to ensure that they can be evenly distributed to each tile within the chip unit after splitting.

[0013] As a specific implementation manner of this application, the generation of the final output result specifically includes:

[0014] Preprocessing and conversion of weight parameters, rearranging the weight parameters in the original layout to a layout adapted to the chip architecture;

[0015] Unification of the input and output shapes of the operator;

[0016] Fusion of multiple operators;

[0017] Support for dynamic shapes of operators.

[0018] As a specific implementation manner of this application, the unification of the input and output shapes of the operator specifically includes:

[0019] Determine the execution shape of each operator by statically analyzing the dependency relationships of the operators in the network, and the input and output shapes of the operators in the entire network can be unified by changing the shape of the weight parameters or inserting data processing operators.

[0020] As a specific implementation manner of this application, the fusion of multiple operators specifically includes:

[0021] Based on the dependency relationship and data flow, adjacent operators are fused into one operator to avoid redundant storage and loading of intermediate data, improve the inference efficiency and reduce the chip storage.

[0022] As a preferred implementation manner of the present application, the intelligent device further includes a deployed multi-machine multi-card multi-die system. When applied to the multi-machine multi-card multi-die system, the method includes:

[0023] The total number of computational structural blocks of the entire large model is evenly divided into pp_num parts, that is, the computational graph of the large model is decomposed into pp_num subgraphs, and these subgraphs are called pipeline parallel subgraphs; where pp_num represents the number of stages included when the large model inference task is decomposed into multiple stages;

[0024] According to the tp_num parameter, the operators in each pipeline parallel subgraph are further split into tp_num tp subgraphs, finally forming a total of pp_num * tp_num subgraphs; where tp_num represents the number of devices involved when distributing a single tensor to multiple devices for execution, and the device is a chip or a chip unit;

[0025] According to the number of chips and the number of chip units in the multi-machine multi-card multi-die system to be deployed, a chip_id and a die_id are assigned to each of the pp_num * tp_num subgraphs;

[0026] According to the chip_id and die_id of each subgraph and the pipeline parallel and tensor parallel logics, necessary communication operators are inserted into each subgraph to ensure the correct running logic of the overall subgraphs; where the pipeline parallel means decomposing the large model inference task into multiple stages, and then executing these stages on different computing units or chip units, and connecting the stages in sequence to form a pipeline; the tensor parallel means that during the inference of the large model, the calculation of a single tensor is distributed to multiple computing units or chip units for execution;

[0027] During the compilation stage, running an inference once according to the running logics of the pipeline parallel and tensor parallel can complete the compilation of the entire large model; where the chip unit executes the compilation method of the single chip unit described above.

[0028] Second aspect, an embodiment of the present invention further provides a large model compilation device based on a near-memory computing architecture, which is applied to an intelligent device. The intelligent device includes a deployed single-chip unit, where the chip unit is represented by a die, and each chip unit includes multiple tiles. The tile represents the smallest allocation unit, and the smallest allocation unit includes a certain number of logic units, computing units, and storage units. The compilation device includes:

[0029] A splitting module, configured to split the dimension N of the operator weight parameters obtained by splitting the large model according to the number of tiles within the chip unit during the compilation stage of the operator. Among them, the first dimension of the shape of the weight parameters corresponding to the operator must be tile_num, and the remaining dimensions are the parameter shapes allocated to each tile respectively;

[0030] A constraint module, configured to simultaneously constrain the shapes of the input and output tensors of the operator, and require that the first dimension of the input and output tensors must be tile_num, so as to ensure the continuity of the operations between operators;

[0031] A polling module, configured to poll and transfer the input tensor of the operator between adjacent tiles to ensure that the input tensor data of each tile can appear on other tiles, and implement the complete calculation process of the operator;

[0032] A processing module, configured to:

[0033] After completing the parallel calculation on all tiles within the chip unit, obtain the output results of the operator on each tile;

[0034] Then merge the output results of all tiles along the tile_num dimension to obtain the output data with tile_num as the first dimension. Among them, tile_num represents the number of tiles in a chip unit;

[0035] Finally, convert the tile_num dimension and other dimensions into the original dimension N according to the data arrangement rule within the chip unit to generate the final output result.

[0036] Third aspect, an embodiment of the present invention further provides a large model compilation device based on a near-memory computing architecture, which is applied to an intelligent device. The intelligent device further includes a deployed multi-machine multi-card multi-die system. When applied to a multi-machine multi-card multi-die system, the device includes:

[0037] A decomposition module, configured to:

[0038] The total number of computational structural blocks of the entire large model is evenly divided into pp_num parts, and the computational graph of the large model is decomposed into pp_num subgraphs, which are called pipeline parallel subgraphs; where pp_num represents the number of stages included when the large model inference task is decomposed into multiple stages;

[0039] According to the tp_num parameter, the operators in each pipeline parallel subgraph are further split into tp_num tp subgraphs, finally forming a total of pp_num * tp_num subgraphs; where tp_num represents the number of devices involved when distributing a single tensor to multiple devices for execution, and the device is a chip or a chip unit;

[0040] The allocation module is used to allocate a chip_id and a die_id to each of the pp_num * tp_num subgraphs according to the number of chips and the number of chip units in the multi-machine, multi-card, and multi-die system to be deployed;

[0041] The insertion module is used to insert necessary communication operators for each subgraph according to the chip_id and die_id of each subgraph and the pipeline parallel and tensor parallel logics to ensure the correct running logic of the overall subgraphs; where the pipeline parallel means decomposing the large model inference task into multiple stages, and then letting these stages be executed on different computing units or chip units, and the stages are connected in sequence to form a pipeline; the tensor parallel means that during the inference of the large model, the calculation of a single tensor is distributed to multiple computing units or chip units for execution;

[0042] The compilation module is used to complete the compilation of the entire large model by running an inference again according to the running logics of the pipeline parallel and tensor parallel during the compilation stage; where the chip unit executes the compilation method of the single chip unit described above.

[0043] The technical solution provided by the embodiments of the present invention realizes the purpose of only transferring data between adjacent kernels by splitting the calculation of the operators of the large model on a single chip unit, distributing the calculation of the entire operator to multiple tiles for parallel execution, and polling and transferring the input tensors of the operators between adjacent tiles, significantly improving the operator calculation efficiency, thereby accelerating the inference process of the large model and realizing the mapping process at the tile level;

[0044] Meanwhile, for models with a large parameter scale, when the hardware resources of a single-chip unit cannot fully accommodate the model, the model is mapped to a multi-machine, multi-card, multi-die system for calculation, enabling the mapping of large models to a near-memory computing architecture chip system with any number of chips and chip units. This significantly enhances the versatility and adaptability of the near-memory computing architecture chip system, as well as the efficient mapping and deployment of large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art.

[0046] Figure 1 is a flowchart of a large model compilation method based on a near-memory computing architecture provided by an embodiment of the present invention;

[0047] Figure 2 is a calculation demonstration diagram of multiple tiles within a single-chip unit of a gemm operator based on a near-memory computing architecture provided by an embodiment of the present invention;

[0048] Figure 3 is the deployment process of a large model with four computational building blocks in a multi-machine, multi-card, multi-die system provided by an embodiment of the present invention;

[0049] Figure 4 is the compilation process of a certain sub-graph provided by an embodiment of the present invention;

[0050] Figure 5 is a structural block diagram of a large model compilation device based on a near-memory computing architecture provided by an embodiment of the present invention;

[0051] Figure 6 is a structural block diagram of another large model compilation device based on a near-memory computing architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0054] Throughout the specification, references to "an embodiment", "embodiments", "an example", or "examples" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "in an embodiment", "in embodiments", "an example", or "examples" that appear throughout the specification do not necessarily refer to the same embodiment or example. Additionally, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples.

[0055] Chip: Refers to a chip. A chip contains multiple dies and can also be understood as a card in a multi-machine, multi-card, multi-die system.

[0056] Die: Refers to a single chip unit, which is the core part of the chip. A die has multiple tiles.

[0057] Tile: Refers to the smallest repeating unit within a single chip unit, i.e., a die. It contains a certain number of logic units, computing units, and storage units, and is also the smallest allocation unit visible for large model compilation.

[0058] Multi-card, multi-die system: In the near-memory computing chip architecture for large models in the present invention, it means that a machine has multiple chips, each chip has multiple dies, and each die has multiple tiles.

[0059] Multi-machine, multi-card, multi-die system: Can be understood as a machine that can deploy multiple cards, and each card can have one die or multiple dies. A card can be understood as a chip.

[0060] Single-card, single-die: Refers to a card that has only one die.

[0061] Shape: Without special explanation in this text, it represents the dimension (shape) of the operator output, input, or weight parameter.

[0062] Gemm: Refers to a general matrix multiplication operator commonly used in large models.

[0063] Rope: Refers to a rotary position encoding operator commonly used in large models.

[0064] Chip_num: The number of chips.

[0065] Tensor: Represents a tensor.

[0066] Die_num: The number of dies in a chip.

[0067] Tile_num: The number of tiles in a die.

[0068] Block: Specifically refers to the repeated structure in the large model network structure in the present invention.

[0069] Block_num: Represents the number of repeated structures in the large model in the present invention.

[0070] Pipeline parallelism: In the present invention, it means decomposing the inference task of the large model into multiple stages, and then allowing these stages to be executed on different computing units (chips or dies). Each stage is connected in sequence to form a pipeline.

[0071] Tensor parallelism: In the present invention, during the inference of the large model, it means distributing the calculation of a single tensor to multiple chips or dies for execution, thereby breaking through the storage and computing power limitations of a single device.

[0072] Pp_num: In the present invention, it refers to the number of stages included when decomposing the inference task of the large model into multiple stages.

[0073] Tp_num: In the present invention, it refers to the number of devices involved when distributing a single tensor to multiple devices for execution. Here, the device can be a chip or a die.

[0074] It should be noted that, unless otherwise specified, the technical terms in this embodiment have the usual meanings understood by those skilled in the art.

[0075] Please refer to Figure 1 , A large model compilation method based on a near-memory computing architecture provided by an embodiment of the present invention is applied to an intelligent device. The intelligent device includes a deployed single-chip unit. Here, the chip unit is represented by a die, and each chip unit contains multiple tiles. The tile represents the smallest allocation unit, and the smallest allocation unit contains a certain number of logic units, computing units, and storage units. The compilation method of the single-chip unit includes:

[0076] S101, In the compilation stage of the operator, split the dimension N of the operator weight parameters split from the large model according to the number of tiles in the chip unit; among them, the first dimension of the shape of the weight parameters corresponding to the operator must be tile_num, and the remaining dimensions correspond to the parameter shapes allocated to each tile.

[0077] S102, Simultaneously constrain the shapes of the input and output tensors of the operator, and require that the first dimension of the input and output tensors must be tile_num, so as to ensure the continuity of the operations between operators.

[0078] S103. Poll and transfer the input tensor of the operator among adjacent tiles to ensure that the input tensor data of each tile can appear on other tiles, and implement the complete calculation process of the operator;

[0079] S104. After completing the parallel calculation on all tiles within the chip unit, obtain the output result of the operator on each tile;

[0080] S105. Then, merge the output results of all tiles along the dimension of tile_num to obtain the output data with tile_num as the first dimension; where tile_num represents the number of tiles in a chip unit;

[0081] S106. Finally, convert the tile_num dimension and other dimensions to the original dimension N through the data layout rule within the chip unit to generate the final output result.

[0082] For convenience of description, in the specific embodiment, the chip unit is represented by die, and the chip is represented by Chip.

[0083] During implementation, when splitting, if the weight parameter cannot be evenly distributed to each tile in dimension N, the dimension N of the original parameter needs to be padded first to ensure that it can be evenly distributed to each tile within the chip unit after splitting.

[0084] Specifically, split the dimension N (usually the output dimension, but not limited to this) of the operator weight parameters such as gemm and rope according to the number of tiles in the die; if the parameter cannot be evenly distributed to each tile in dimension N, the dimension N of the original parameter needs to be padded first to ensure that it can be evenly distributed to each tile within the die after splitting. In this case, the first dimension of the shape of the weight parameter of the corresponding operator must be tile_num, indicating that the operator will perform parallel calculation on tile_num tiles, and the remaining dimensions correspond to the parameter shape distributed to each tile. These distributed parameters will be stored in the local memory of each tile.

[0085] In this embodiment, since the parameters of the tiles allocated to each die are different, but the shapes are the same, in order to complete the complete calculation process of the operator on all tiles within a single die, it is necessary to poll and transfer the input tensor of the operator (i.e., the Feature eigenvalue) between adjacent tiles. This method can ensure that the input tensor data of each tile can appear on other tiles. Therefore, the complete calculation process of the operator can be realized through the polling of the input tensor between adjacent tiles. In this way, the transfer of eigenvalues between tiles and the calculation process of tiles can be parallelized, thus significantly improving the calculation efficiency of the operator.

[0086] Refer to Figure 2 , taking the gemm operator as an example, based on the calculation demonstration diagram of multiple tiles within a single die of the gemm operator in the near-memory computing architecture, a complete gemm operator is split into multiple tiles within a single die for parallel calculation, and the characteristics of the near-memory computing architecture are fully utilized. The data transfer speed between tiles depends on the relative positions between tiles, and the input tensor only needs to be transferred between adjacent tiles, greatly accelerating the parallel speed of the operator.

[0087] In another embodiment, for a model with extremely large parameter scale, it needs to be distributedly deployed on a multi-machine, multi-card, and multi-die system based on the near-memory computing architecture. For this purpose, we propose a compilation method that can flexibly map a large model to a chip system with any number of Chips and Dies. That is, the intelligent device further includes a deployed multi-card and multi-chip unit system. When applied to the multi-card and multi-chip unit system, the specific implementation is as follows:

[0088] The total number of computational structural blocks of the entire large model is evenly divided into pp_num parts, that is, the computational graph of the large model is decomposed into pp_num subgraphs, and these subgraphs are called pipeline parallel subgraphs. Among them, pp_num represents the number of stages included when the inference task of the large model is decomposed into multiple stages.

[0089] According to the tp_num parameter, the operators in each pipeline parallel subgraph are further split into tp_num tp subgraphs, finally forming a total of pp_num * tp_num subgraphs. Among them, tp_num represents the number of devices involved when a single tensor is distributed to multiple devices for execution, and the device is a chip or a chip unit.

[0090] According to the number of chips and the number of chip units in the multi-machine, multi-card, and multi-die system to be deployed, a chip_id and a die_id are assigned to each of the pp_num * tp_num subgraphs, thereby indicating which die on which chip the subgraph runs.

[0091] According to the chip_id and die_id of each sub-graph and the pipeline parallel and tensor parallel logics, necessary communication operators are inserted for each sub-graph to ensure the correct running logic of these sub-graphs as a whole; wherein, the pipeline parallel means decomposing the inference task of a large model into multiple stages, and then letting these stages be executed on different computing units or chip units, and the stages are connected in sequence to form a pipeline; the tensor parallel means that during the inference of a large model, the calculation of a single tensor is distributed to multiple computing units or chip units for execution.

[0092] During the compilation stage, running an inference once according to the running logics of the pipeline parallel and tensor parallel can complete the compilation of the entire large model; wherein, the chip unit executes the compilation method of the single chip unit described above.

[0093] During application, taking advantage of the characteristic that a large model usually consists of multiple repeated computational structural blocks (i.e., block blocks), the pp_num parameter is set by evenly dividing according to the number of repeated blocks in the large model; and the tp_num is set according to the number of chips or dies in a multi-machine, multi-card, multi-die system.

[0094] The communication operators include communication operators such as AllReduce, Send, Receive, etc.; during the compilation process, the executable file generated by compiling each sub-graph and the binary file of the relevant weight parameters are saved to a directory named with the device number {server_id}{chip_id}{die_id}, so that during actual inference, each device only needs to load the corresponding file according to the device number.

[0095] The compilation method of this large model can map the same set of pp_num and tp_num parameters to device systems with different numbers of chips and dies, which significantly improves the versatility and flexibility of the chip system based on the near-memory computing architecture in deployment.

[0096] Refer to Figure 3 , which shows how to deploy a large model with 4 computational structural blocks (block number) to a multi-card, multi-die system with chip_num = 2, die_num = 2 and chip_num1, die_num = 4 when the parameters are pp_num = 2 and tp_num = 2.

[0097] Furthermore, due to the significant differences between the original data layouts of traditional large model operators and the data layouts of chip systems based on the near-memory computing architecture, the chips based on the near-memory computing architecture cannot directly use the original layouts. To avoid affecting the execution efficiency due to the conversion of data shapes during the execution of operators, a solution is proposed. During the compilation phase, the conversion of weight parameters from the original layout to the chip layout is completed, and at the same time, the input and output shapes of each operator are unified to ensure the continuity of the inference of the entire network's operators. Also, to improve the execution efficiency of operators as much as possible, operators are automatically fused during the compilation phase. Multiple operators with data dependencies are fused into one operator for execution, which can improve the execution efficiency and avoid redundant storage and loading of intermediate data. To support the dynamic inference of the entire network and enable inference in any situation with a single compilation, the dimensions that need to be dynamically supported in the operator shape are represented by variable expressions during the compilation process. The specific implementation is as follows:

[0098] Preprocessing and Conversion of Weight Parameters

[0099] During the compilation phase, by analyzing the original layout of the weight parameters of each operator and combining the characteristics of the chip's near-memory computing architecture, an efficient layout conversion algorithm is designed. According to the chip's memory access mode, the computing degree of the computing unit, and the data alignment requirements, such as 16-byte alignment, 32-byte alignment, etc., the weight parameters in the original layout are rearranged into a layout that adapts to the chip architecture. For example, for the gemm operator, the weight can be converted from the traditional N x K to the block layout of tile_num x N / / tile_num x tile_num x K / / tile_num, which is suitable for the near-memory computing architecture chip accelerator, thereby improving the data access efficiency during the execution of the operator.

[0100] Unification of Operator Input and Output Shapes

[0101] To avoid additional operator operations or data copy conversions due to mismatched input and output shapes during the inference of the entire network, the input and output shapes of each operator are unified during the compilation phase. This method mainly determines the execution shape of each operator by statically analyzing the dependency relationships of the operators in the network. When necessary, the shape of the weight parameters can be changed or relevant data processing operators such as Pad or Depad operators can be inserted to achieve the unification of the input and output shapes of the entire network's operators. For example, the original input shape of the gemm operator M x N is converted to the layout of the near-memory computing architecture chip tile_num x N / / tile_num / / 16 x M x 16, where M represents the length of the inference input token and 16 represents 16-byte alignment of the data.

[0102] Fusion of Multiple Operators

[0103] Dependency relationships and data flow directions. As much as possible, fuse these adjacent operators into one operator, which can avoid redundant storage and loading of intermediate data, improve the inference efficiency of the entire network, and reduce chip storage.

[0104] Operator dynamic shape support

[0105] To support inference in any situation with a single compilation, during the compilation process, represent the dimensions that need dynamic support in the operator shape with variable expressions. This variable expression can be a single variable or a calculation expression composed of multiple variables. By implementing dynamic compilation in this way, the flexibility and generality of large model compilation are greatly improved.

[0106] It should be noted that, for example, for a matrix of M x N, split the N dimension, M x N -> M x tile_num x N / / tile_num / / 16 x 16 -> tile_num x N / / tile_num / / 16 x M x 16;

[0107] Then N / / tile_num / / 16 and 16 are the other dimensions mentioned above.

[0108] Refer to Figure 4 , which represents the compilation process of a sub-graph of a large model based on a near-memory computing architecture chip. When deploying a multi-card multi-die system, its compilation process consists of the compilation of the above-mentioned multiple sub-graphs.

[0109] In the above solution, by splitting the calculation of the operators of the large model on a single chip unit, distributing the calculation of the entire operator to multiple tiles for parallel processing, and polling and transferring the input tensors of the operators between adjacent tiles, the goal of only transferring data between adjacent cores is finally achieved, significantly improving the operator calculation efficiency, thereby accelerating the inference process of the large model and realizing the mapping process at the tile level;

[0110] At the same time, for models with a large parameter scale, when the hardware resources of a single chip unit cannot fully carry the model, by mapping the model to a multi-machine multi-card multi-die system for calculation, the large model is mapped to a near-memory computing architecture chip system with any number of chips and chip units, thereby significantly improving the generality and adaptability of the near-memory computing architecture chip system, as well as the efficient mapping and deployment of large models.

[0111] Based on the same inventive concept, the embodiments of the present invention also provide a large model compilation device based on a near-memory computing architecture. Refer to Figure 5, applied to a smart device, the smart device includes a deployed single-chip unit, where the chip unit is represented by a die, each chip unit contains multiple tiles, the tile represents the smallest allocation unit, and the smallest allocation unit contains a certain number of logic units, computing units, and storage units; the compilation device includes:

[0112] A splitting module, used in the compilation stage of the operator, to split the dimension N of the operator weight parameters split from the large model according to the number of tiles within the chip unit; where the first dimension of the shape of the weight parameters corresponding to the operator must be tile_num, and the remaining dimensions are the parameter shapes allocated to each tile respectively;

[0113] A constraint module, used to simultaneously constrain the shapes of the input and output tensors of the operator, requiring that the first dimension of the input and output tensors must be tile_num, so as to ensure the continuity of the operations between operators;

[0114] A polling module, used to poll and transfer the input tensor of the operator between adjacent tiles, so as to ensure that the input tensor data of each tile can appear on other tiles and realize the complete calculation process of the operator;

[0115] A processing module, used for:

[0116] After completing the parallel calculation on all tiles within the chip unit, obtain the output results of the operator on each tile;

[0117] Then merge the output results of all tiles along the tile_num dimension, so as to obtain the output data with tile_num as the first dimension; where tile_num represents the number of tiles in a chip unit;

[0118] Finally, convert the tile_num dimension and other dimensions to the original dimension N according to the data layout rule within the chip unit to generate the final output result.

[0119] In this embodiment, when splitting, if the weight parameters cannot be evenly distributed to each tile in dimension N, it is necessary to first pad the dimension N of the original parameters to ensure that they can be evenly distributed to each tile within the chip unit after splitting.

[0120] Furthermore, to improve the operation speed of the operator on the near-memory computing architecture chip and avoid frequent shape conversions of weight parameters during the execution of the operator as much as possible, the present invention proposes to complete the conversion of the original arrangement of weight parameters at the compilation stage and adjust it to a form more suitable for the near-memory computing architecture, thereby effectively improving the computing efficiency; enabling the generation of the final output result, specifically including:

[0121] Preprocessing and conversion of weight parameters, rearranging the weight parameters in the original layout into a layout suitable for the chip architecture;

[0122] Unification of the input and output shapes of the operator;

[0123] Fusion of multiple operators;

[0124] Support for dynamic shapes of operators.

[0125] It should be noted that for a more specific description of the working process of the compilation device embodiment, please refer to the foregoing method embodiment part and will not be elaborated here.

[0126] For a model with a large parameter scale, when the hardware resources of a single card and single die cannot fully carry the model and the model needs to be mapped to a multi-card and multi-die system for calculation, the embodiment of the present invention also provides a large model compilation device based on the near-memory computing architecture. Referring to Figure 6 , applied to an intelligent device, the intelligent device further includes a deployed multi-machine, multi-card and multi-die system. When applied to the multi-machine, multi-card and multi-die system, the device includes:

[0127] A decomposition module, configured to:

[0128] Evenly divide the total number of computational structural blocks of the entire large model into pp_num parts, decompose the computational graph of the large model into pp_num subgraphs, and these subgraphs are called pipeline parallel subgraphs; where pp_num represents the number of stages included when decomposing the large model inference task into multiple stages;

[0129] According to the tp_num parameter, further split the operators in each pipeline parallel subgraph into tp_num tp subgraphs, finally forming a total of pp_num * tp_num subgraphs; where tp_num represents the number of devices involved when distributing a single tensor to multiple devices for execution, and the device is a chip or a chip unit;

[0130] An allocation module, configured to allocate a chip_id and a die_id to each of the pp_num * tp_num subgraphs according to the number of chips and the number of chip units in the multi-card and multi-chip unit system to be deployed;

[0131] An insertion module, which is used to insert necessary communication operators for each sub - graph according to the chip_id and die_id of each sub - graph and the pipeline parallel and tensor parallel logics, so as to ensure the correct running logic of the overall sub - graphs; wherein, the pipeline parallel means decomposing the inference task of a large model into multiple stages, and then letting these stages be executed on different computing units or chip units, and the stages are connected in sequence to form a pipeline; the tensor parallel means that during the inference of a large model, the calculation of a single tensor is distributed to multiple computing units or chip units for execution;

[0132] A compilation module, which is used to complete the compilation of the entire large model by running an inference again according to the running logics of the pipeline parallel and tensor parallel during the compilation stage; wherein, the chip unit executes the compilation method of the single chip unit described above.

[0133] For the entire solution, by splitting the calculation of the operators of the large model on a single chip unit, distributing the calculation of the entire operator to multiple tiles for parallel execution, and polling and transferring the input tensors of the operator between adjacent tiles, the purpose of only transferring data between adjacent kernels is finally achieved, significantly improving the operator calculation efficiency, thereby accelerating the inference process of the large model and realizing the mapping process at the tile level;

[0134] Meanwhile, for models with a large parameter scale, when the hardware resources of a single chip unit cannot fully carry the model, by mapping the model to a multi - machine, multi - card, multi - die system for calculation, the large model is mapped to a near - memory computing architecture chip system with any number of chips and chip units, thereby significantly improving the versatility and adaptability of the near - memory computing architecture chip system, as well as the efficient mapping and deployment of the large model.

[0135] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A large model compilation method based on a near-memory computing architecture, characterized in that, Applied to intelligent devices, the intelligent devices include deployed single-chip units, where each chip unit contains multiple tiles, the chip unit is represented by a die, the tile represents the smallest allocation unit, and the smallest allocation unit contains a certain number of logic units, computing units, and storage units; the compilation method of the single-chip unit includes: In the compilation stage of the operator, split the dimension N of the operator weight parameters split from the large model according to the number of tiles in the chip unit; where the first dimension of the shape of the weight parameters corresponding to the operator must be tile_num, and the remaining dimensions correspond to the parameter shapes allocated to each tile; Simultaneously constrain the shapes of the input and output tensors of the operator, requiring that the first dimension of the input and output tensors must be tile_num, so as to ensure the continuity of the operations between operators; Poll and transfer the input tensors of the operator between adjacent tiles to ensure that the input tensor data of each tile can appear on other tiles, and implement the complete calculation process of the operator; After completing the parallel calculation on all tiles within the chip unit, obtain the output results of the operator on each tile; Then merge the output results of all tiles along the tile_num dimension to obtain the output data with tile_num as the first dimension; where tile_num represents the number of tiles in a chip unit; Finally, convert the tile_num dimension and other dimensions to the original dimension N through the data arrangement rules within the chip unit to generate the final output result.

2. The large model compilation method based on the near-memory computing architecture according to claim 1, wherein, When splitting, if the weight parameters cannot be evenly distributed to each tile in dimension N, it is necessary to first pad the dimension N of the original parameters to ensure that they can be evenly distributed to each tile within the chip unit after splitting.

3. The large model compilation method based on the near-memory computing architecture according to claim 2, characterized in that The generation of the final output result specifically includes: Preprocessing and conversion of weight parameters, rearranging the weight parameters in the original layout into a layout adapted to the chip architecture; Unification of the shapes of operator inputs and outputs; Fusion of multiple operators; Support for dynamic shapes of operators.

4. The large model compilation method based on the near-memory computing architecture according to claim 3, characterized in that The unification of the shapes of operator inputs and outputs specifically includes: By statically analyzing the dependency relationships of each operator in the network, determine the execution shape of each operator, and the shapes of the weight parameters can be changed or data processing operators can be inserted to achieve the unification of the shapes of operator inputs and outputs across the network.

5. The large model compilation method based on the near-memory computing architecture according to claim 3, characterized in that, The fusion of multiple operators specifically includes: Based on the dependency relationships and data flow directions, fuse adjacent operators into one operator to avoid redundant storage and loading of intermediate data, improve the inference efficiency, and reduce the chip storage.

6. The large model compilation method based on the near-memory computing architecture according to claim 3, wherein, The intelligent device also includes a deployed multi-machine multi-card multi-die system. When applied to the multi-machine multi-card multi-die system, the method includes: The total number of computational structure blocks of the entire large model is evenly divided into pp_num parts, that is, the computational graph of the large model is decomposed into pp_num subgraphs, and these subgraphs are called pipeline parallel subgraphs; where pp_num represents the number of stages included when the large model inference task is decomposed into multiple stages; According to the tp_num parameter, the operators in each pipeline parallel subgraph are further split into tp_num tp subgraphs, finally forming a total of pp_num * tp_num subgraphs; where tp_num represents the number of devices involved when distributing a single tensor to multiple devices for execution, and the device is a chip or a chip unit; According to the number of chips and the number of chip units in the multi-machine, multi-card, multi-die system to be deployed, a chip_id and a die_id are assigned to each of the pp_num * tp_num subgraphs; According to the chip_id and die_id of each subgraph and the pipeline parallel and tensor parallel logics, necessary communication operators are inserted into each subgraph to ensure the correct running logic of these subgraphs as a whole; where the pipeline parallel means decomposing the large model inference task into multiple stages, and then letting these stages execute on different computing units or chip units, and the stages are connected in sequence to form a pipeline; the tensor parallel means that during the inference of the large model, the calculation of a single tensor is distributed to multiple computing units or chip units for execution; In the compilation stage, running an inference once according to the running logics of the pipeline parallel and tensor parallel can complete the compilation of the entire large model; where the chip unit executes the compilation method of the single chip unit described above.

7. A large model compilation device based on a near-memory computing architecture, characterized in that, Applied to an intelligent device, the intelligent device includes a deployed single chip unit, where each chip unit includes multiple tiles, the chip unit is represented by a die, the tile represents the smallest allocation unit, and the smallest allocation unit includes a certain number of logic units, computing units and storage units; the compilation device includes: A splitting module, used in the compilation stage of the operator, to split the dimension N of the operator weight parameter split from the large model according to the number of tiles in the chip unit; where the first dimension of the shape of the weight parameter corresponding to the operator must be tile_num, and the remaining dimensions are the parameter shapes allocated to each tile; A constraint module, used to simultaneously constrain the shapes of the input and output tensors of the operator, and requires that the first dimension of the input and output tensors must be tile_num, so as to ensure the continuity of the operations between operators; A polling module, used to poll and transfer the input tensor of the operator between adjacent tiles, so as to ensure that the input tensor data of each tile can appear on other tiles and realize the complete calculation process of the operator; A processing module, used for: After completing the parallel computing on all tiles within the chip unit, obtain the output results of the operators on each tile; Then, merge the output results of all tiles along the dimension of tile_num, so as to obtain the output data with tile_num as the first dimension; where tile_num represents the number of tiles in a chip unit; Finally, convert the tile_num dimension and other dimensions to the original dimension N through the data layout rule within the chip unit to generate the final output result.

8. The large model compilation device based on the near-memory computing architecture according to claim 7, characterized in that, When performing splitting, if the weight parameters cannot be evenly distributed to each tile on dimension N, it is necessary to first pad dimension N of the original parameters to ensure that they can be evenly distributed to each tile within the chip unit after splitting.

9. A large model compilation device based on a near-memory computing architecture, characterized in that, Applied to an intelligent device, the intelligent device further includes a deployed multi-machine multi-card multi-die system. When applied to the multi-machine multi-card multi-die system, the device includes: A decomposition module, configured to: Evenly divide the total number of computational structural blocks of the entire large model into pp_num parts, decompose the computational graph of the large model into pp_num subgraphs, and these subgraphs are called pipeline parallel subgraphs; where pp_num represents the number of stages included when decomposing the large model inference task into multiple stages; According to the tp_num parameter, further split the operators in each pipeline parallel subgraph into tp_num tp subgraphs, finally forming a total of pp_num * tp_num subgraphs; where tp_num represents the number of devices involved when distributing a single tensor to multiple devices for execution, and the device is a chip or a chip unit; An allocation module, configured to allocate a chip_id and a die_id to each of the pp_num * tp_num subgraphs according to the number of chips and the number of chip units in the multi-machine multi-card multi-die system to be deployed; An insertion module, configured to insert necessary communication operators for each subgraph according to the chip_id and die_id of each subgraph and the pipeline parallel and tensor parallel logics to ensure the correct running logic of these subgraphs as a whole; where the pipeline parallel means decomposing the large model inference task into multiple stages, and then allowing these stages to be executed on different computing units or chip units, and the stages are connected in sequence to form a pipeline; the tensor parallel means that during the inference of the large model, the calculation of a single tensor is distributed to multiple computing units or chip units for execution; A compilation module, configured to, during the compilation stage, run an inference again according to the running logics of the pipeline parallel and tensor parallel to complete the compilation of the entire large model; where the chip unit executes the compilation method of the single chip unit described above.