Dynamic-length large model compiling method and device

By splitting the big model into an operator set, and using the address generator and instruction configuration register to adjust the operator configuration at runtime, the memory usage and performance loss problems of the dynamic shape compiler are solved, and efficient dynamic length big model compilation is achieved.

CN120335813APending Publication Date: 2025-07-18BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510329181.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When existing AI compilers deal with operators of dynamic shapes, they have problems with high memory footprint, performance loss and additional overhead, making it difficult to efficiently support dynamic-length large-scale model compilation.

Method used

Split the big model into a collection of several operators, and the operator configuration is globally adjusted according to the dynamic length when running through the address generator and instruction configuration registers, to realize the compilation of dynamic operators, and to use hardware resources for flexible and efficient compilation.

Benefits of technology

With extremely small additional storage and runtime overhead, the performance of dynamic models is no less than that of static models, and improves the utilization efficiency and application scenarios of hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335813A_ABST
    Figure CN120335813A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a dynamic-length large model compiling method and device. The method comprises the steps that a trained large model is split into a set of a plurality of operators in a preset mode; each operator comprises a plurality of micro operators; dynamic length is used as input for operators in the large model, compiling of dynamic operators is achieved, and the compiled operators are spliced according to a network structure to achieve compiling of the model; the dynamic length is the length corresponding to each operator; the runtime support comprises globally adjusting the configuration of a model operator according to the received dynamic length of the reasoning during the model reasoning, and then executing the reasoning; the operator configuration is obtained by configuring an address generator and an instruction configuration register in the AI chip; the method has the beneficial effects that the remainder micro-operator generated by each operator is configured according to the dynamic length during operation, so that flexible and efficient dynamic length large model compiling can be realized by only needing a small amount of hardware support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular to a large model compilation method and device with dynamic length. Background Art

[0002] The full name of a large model is a large language model (LLM, Large Language Model). The "large" mainly refers to the large capacity of the model structure, many parameters in the structure, and a large amount of data used for pre-training the large model.

[0003] Operator: An operator is a general term for a certain mathematical operation. With the development of deep learning technology, the scenarios of using dynamic shape for inference have increased sharply, and there are more and more operators with dynamic shapes. Due to the uncertainty of the dynamic shape semantics, the existing AI compiler optimization systems face many problems brought by the dynamic shape models. Most compilers have low support for the representation of dynamic shapes and are not good at handling dynamic shapes. In the existing technical solutions, the support for dynamic shapes is often achieved through staticization.

[0004] A commonly used solution is as follows: Prepare N models with maximum lengths of 64, 128, 512, 2048, etc. Select the smallest available model and pad the input to the corresponding length. This solution will bring additional memory occupation due to the need to compile multiple models. At the same time, the padding data will generate additional overhead during the inference process, reducing the model performance.

[0005] Another solution is: A set of pre-compiled "micro-operators". After knowing the specific shape information at runtime, through a specific search algorithm, select one or more combinations of "micro-operators" to form a complete operator. This solution requires additional search overhead at runtime, and there are limitations in operator scheduling and operator fusion, making the performance of the compiled model lag behind that of static compilation. At the same time, due to the dynamic nature of the memory address brought by the dynamic shape, this solution is difficult to handle the situation where dynamic operators need to interact with memory. Summary of the Invention

[0006] Aiming at the technical defects of the existing technology, the purpose of the embodiments of the present invention is to provide a large model compilation method and device with dynamic length.

[0007] To achieve the above purpose, in the first aspect, the embodiments of the present invention provide a large model compilation method with dynamic length, and the method includes:

[0008] Split the trained large model into a set of several operators in a preset mode; wherein, each operator includes a plurality of micro-operators;

[0009] Use the operators in the large model with dynamic length as input to implement the compilation of dynamic operators, and splice the compiled operators according to the network structure to achieve the compilation of the model; wherein, the dynamic length is the length corresponding to each operator.

[0010] Runtime support, including during model inference, globally adjusting the model operator configuration according to the dynamic length received for this inference, and then performing the inference; wherein, the operator configuration is obtained by configuring the address generator and instruction configuration register in the AI chip.

[0011] As a specific implementation manner of the present application, the large model is split into several operators, and dynamic operator compilation is performed one by one.

[0012] As a specific implementation manner of the present application, the operator configuration specifically includes:

[0013] Configure the base addresses of the input and output of each operator in the model through the address generator to record the instruction memory access address.

[0014] Configure the remainder micro-operators generated by each operator through the instruction configuration register to support dynamic micro-operators.

[0015] As a specific implementation manner of the present application, the address generator uses one or more groups of registers and counters to record the instruction memory access address.

[0016] As a specific implementation manner of the present application, the execution process of the address generator includes:

[0017] Configure the starting address and increment rule of the address generator;

[0018] Bind the address generator during instruction generation;

[0019] When executing the first instruction, use the starting address of the address generator and clear the counter;

[0020] After the instruction execution is completed, increment the counter by one;

[0021] Execute subsequent instructions, and use the counter and address increment rule to generate the instruction memory access address.

[0022] As a specific implementation manner of the present application, the instruction configuration register is specifically used for:

[0023] After the execution times are completed, configure the remainder micro-operator instructions; wherein, the execution times are determined by the AI chip according to the input size.

[0024] In a second aspect, an embodiment of the present invention further provides a large model compilation device with dynamic length, and the device includes:

[0025] A splitting module, configured to split a trained large model into a set of several operators in a preset mode; wherein each operator includes a plurality of micro-operators.

[0026] A compiling module, configured to compile the operators in the large model with a dynamic length as the input to implement the compilation of dynamic operators, and splice the compiled operators according to the network structure to implement the compilation of the model; wherein the dynamic length is the length corresponding to each operator.

[0027] A processing module, configured to support during runtime, including during model inference, globally adjusting the model operator configuration according to the received dynamic length for this inference, and then performing the inference; wherein the operator configuration is obtained by configuring the address generator and instruction configuration register in the AI chip.

[0028] The technical solution provided by the embodiments of the present invention, by combining the network structure characteristics of the large model, abstracts several operators from the large model network, and completes the compilation of dynamic operators by configuring the remainder micro-operators generated by each operator according to the dynamic length during runtime, achieving extremely small additional storage and runtime overhead, and at the same time, the compiled dynamic model has performance no less than that of the static model. Description of the Drawings

[0029] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art.

[0030] Figure 1 is a flowchart of a method for compiling a large model with a dynamic length provided by the embodiments of the present invention;

[0031] Figure 2 is the execution flow of an address generator provided by the embodiments of the present invention;

[0032] Figure 3 is the operation flow of a dynamic operator provided by the embodiments of the present invention;

[0033] Figure 4 is a structural block diagram of a device for compiling a large model with a dynamic length provided by the embodiments of the present invention. Detailed Embodiments

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be understood that, as used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0036] Throughout the specification, references to "one embodiment", "an embodiment", "one example" or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the present invention. Thus, the phrases "in one embodiment", "in an embodiment", "one example" or "an example" appearing throughout the specification do not necessarily all refer to the same embodiment or example. Additionally, the particular features, structures, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples.

[0037] Tensor: In the field of artificial intelligence, a tensor is a multi-dimensional array.

[0038] Dynamic Dimension / Dynamic Length: Since each tensor represents a physical meaning, such as the pixels of an image, a convolutional kernel, etc. In large models, due to the variable length of the input text, large models usually have tensors with dynamic lengths.

[0039] It should be noted that, unless otherwise specified, the technical terms in this embodiment have the ordinary meanings understood by those skilled in the art.

[0040] Please refer to Figure 1 , a method for compiling a large model with dynamic length provided by an embodiment of the present invention, the method includes:

[0041] S101, splitting the trained large model into a set of several operators in a preset mode; wherein, each operator includes multiple micro-operators; each "micro-operator" processes a part of the input data, and for ease of description, each "micro-operator" is called a kernel;

[0042] S102, taking the operators in the large model with the dynamic length as the input, implementing the compilation of dynamic operators, and splicing the compiled operators according to the network structure to achieve the compilation of the model; wherein, the dynamic length is the length corresponding to each operator;

[0043] S103, runtime support, including during model inference, globally adjusting the model operator configuration according to the dynamic length received for this inference, and then performing the inference; wherein, the operator configuration is obtained by configuring the address generator and instruction configuration register in the AI chip.

[0044] In practical applications, the hardware's support for compiling dynamic shapes is often very limited and cannot widely support dynamic deep learning models, which seriously restricts the application of deep learning technology. Therefore, it is of great value to jointly solve the problems encountered by dynamic deep learning models during deployment through software and hardware. A good dynamic shape compilation project can improve the utilization efficiency of hardware resources and enable AI processors to obtain higher performance and a wider range of usage scenarios.

[0045] During implementation, the operator configuration specifically includes:

[0046] Configure the base addresses of the inputs and outputs of each operator in the model through the address generator to record the instruction memory access addresses; the address generator uses one or more groups of registers and counters to record the instruction memory access addresses;

[0047] Configure the remainder micro-operators generated by each operator through the instruction configuration register to support dynamic micro-operators.

[0048] In this embodiment, the specific implementation is mainly divided into two parts: hardware and software:

[0049] Hardware part:

[0050] 1. Address generator

[0051] Add one or more groups of registers and counters to record the instruction memory access addresses and times. There is no need to specify the memory address to be accessed currently in the instruction. Just bind the instruction to the address incrementer. According to the execution times of the current instruction and the address increment rule pre-configured in the register, the automatic derivation of the instruction memory access address is realized. The usage process of the address generator is shown in Figure 2 .

[0052] Specifically, configure the starting address and increment rule of the address generator;

[0053] Bind the address generator during instruction generation;

[0054] When executing the first instruction, use the starting address of the address generator and clear the counter;

[0055] After the instruction execution is completed, increment the counter by one;

[0056] When executing subsequent instructions, use the counter and the address increment rule to generate the instruction memory access address.

[0057] 2. Instruction configuration register

[0058] Before the instruction execution, modify the instruction content of the remainder kernel by writing to the instruction register, and then execute the remainder kernel.

[0059] The instruction configuration register is specifically used for:

[0060] After the execution times are completed, configure the remainder micro-operator instruction; wherein, the execution times are determined by the AI chip according to the input size.

[0061] Software part:

[0062] 1. Model splitting

[0063] Currently, the operators applied by mainstream large models on the market are relatively unified and single. Therefore, we only need to adopt a fixed mode to split the large model into a few operators and perform dynamic operator compilation one by one.

[0064] According to the current mainstream Decoder_only model structure, we can split the large model into a few operators such as Gemm, Attention, Normalization, and position encoding.

[0065] 2. Dynamic operator compilation

[0066] Based on the above concept of "micro-operator", each "micro-operator" processes a part of the input data. Here, we call each "micro-operator" a kernel. Due to the dynamic length of the input, the number of kernels to be generated and the data processing granularity of each kernel cannot be predicted during compilation when the size is unknown.

[0067] It should be noted that each kernel here needs to complete three parts of work: reading a part of the input from memory, performing corresponding calculations on the part of the input, and writing the calculation results back to memory. If the data processed by each kernel is of the same size, only the memory read and write addresses change between each kernel. At the same time, the offset of each kernel's memory access address relative to the start address of the input and output is also fixed. Therefore, only by implementing the corresponding address auto-increment function in hardware can the reuse of kernels be achieved.

[0068] If the data size processed by each kernel is fixed, the supported dynamic length can only be an integer multiple of the length processed by each kernel, and there is still a "remainder";

[0069] Therefore, it is also necessary to support a "remainder kernel", that is, a dynamic kernel, which dynamically configures the length of the kernel processed according to the remainder of the dynamic length at runtime. However, since there is only one dynamic kernel for each operator, the overhead of this operation is negligible. The premise is that the hardware needs to support the implementation of the dynamic kernel by rewriting registers and other means. See the running process of the dynamic operator in Figure 3 .

[0070] In another embodiment, on the basis of the above technical solution, when performing dynamic operator compilation, a combination of runtime compilation technology and static compilation can also be adopted. Only the "remainder kernel" is dynamically compiled at runtime, and the remaining kernels can be determined in advance by static compilation.

[0071] 3. Runtime Support

[0072] Since all the "static" content in the dynamic operator is retained in the operator compilation stage, in the runtime stage, compared with static runtime, only two additional tasks need to be completed:

[0073] One is to configure the base addresses of the inputs and outputs of each operator in the model (the base address here is the starting address described above);

[0074] The other is to configure the remainder kernel generated by each operator.

[0075] Both of the above aspects can be calculated using simple formulas or algorithms based on the input length, and basically no additional overhead is incurred.

[0076] For ease of understanding, the GEMM operator is used as an example for illustration.

[0077] Scenario setting:

[0078] Suppose it is necessary to calculate the matrix multiplication C = A × B, where:

[0079] The size of the input matrix A is 1040×1024 (dynamic number of rows is not aligned);

[0080] The size of the input matrix B is 1024×1024;

[0081] The fixed processing block size is 32×1024.

[0082] Configure the register to determine the number of micro-operator executions: floor(1040 / 32) = 32 complete row blocks; it can also be understood as static Kernel preprocessing;

[0083] Corresponding to dynamic Kernel processing (remainder part);

[0084] Remainder calculation:

[0085] Row remainder: 1040 % 32 = 16 (the last 16 rows);

[0086] Dynamic configuration:

[0087] Set the processing size of the dynamic kernel to 16×1024 through the hardware register;

[0088] Execution process:

[0089] Only 1 dynamic kernel needs to be started;

[0090] Automatically capture the last 16 lines of A for multiplication and addition;

[0091] After the static kernel is executed, the pointer is automatically offset by 32 × sizeof(float), eliminating the pointer calculation overhead in traditional loops;

[0092] Dynamically adjust the physical layout of the buffer according to the remainder (such as switching from 32 × 1024 to 16 × 1024);

[0093] Dynamic part acceleration:

[0094] The remainder kernel reuses the instruction pipeline of the static kernel

[0095] Implement mode switching through register reconfiguration (instead of recompilation), making the dynamic operation time < the average time of the static kernel.

[0096] In the above solution, by combining the network structure characteristics of the large model, several operators are abstracted from the large model network. By configuring the remainder micro-operators generated by each operator according to the dynamic length at runtime, the compilation of the dynamic operator is completed, achieving only minimal additional storage and runtime overhead. At the same time, the compiled dynamic model has performance no less than that of the static model.

[0097] Based on the same inventive concept, the embodiment of the present invention also provides a large model compilation device with dynamic length. Refer to Figure 4 , the device includes:

[0098] A splitting module, configured to split the trained large model into a set of several operators in a preset mode; wherein, each operator includes multiple micro-operators;

[0099] A compilation module, configured to use the dynamic length as input to compile the operators in the large model, realize the compilation of the dynamic operator, and splice the compiled operators according to the network structure to realize the compilation of the model; wherein, the dynamic length is the length corresponding to each operator;

[0100] A processing module, configured to support runtime, including during model inference, globally adjusting the model operator configuration according to the dynamic length received in this inference, and then performing inference; wherein, the operator configuration is obtained by configuring the address generator and instruction configuration register in the AI chip.

[0101] Further, the operator configuration specifically includes:

[0102] Configure the base addresses of the input and output of each operator in the model through the address generator to record the instruction memory access address;

[0103] Configure the remainder micro-operators generated by each operator through the instruction configuration register to support dynamic micro-operators.

[0104] The execution process of the address generator includes:

[0105] Configure the starting address of the address generator and the self-increment rule;

[0106] Bind the address generator during instruction generation;

[0107] When executing the first instruction, use the starting address of the address generator and clear the counter;

[0108] After the instruction execution is completed, increment the counter by one;

[0109] When executing subsequent instructions, generate the instruction memory access address using the counter and the address self-increment rule.

[0110] In this embodiment, the instruction configuration register is specifically used for:

[0111] After the execution times are completed, configure the remainder micro-operator instruction; wherein, the execution times are determined by the AI chip according to the input size.

[0112] It should be noted that for a more specific description of the working process of the device embodiment, please refer to the foregoing method embodiment section and will not be elaborated here.

[0113] For the entire solution, through a detailed analysis of the operator dynamics in the large model, reasonably divide the work content of the compilation stage and the runtime stage, and only a small amount of hardware support is required to achieve flexible and efficient compilation of large models with dynamic lengths.

[0114] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for compiling a large model with dynamic length, characterized in that, The method includes: Splitting the trained large model into a set of several operators in a preset mode; wherein, each operator includes multiple micro-operators; Taking the operators in the large model as input with a dynamic length, realizing the compilation of dynamic operators, and splicing the compiled operators according to the network structure to realize the compilation of the model; wherein, the dynamic length is the length corresponding to each operator; Runtime support, including during model inference, globally adjusting the model operator configuration according to the dynamic length received for this inference, and then performing the inference; wherein, the operator configuration is obtained by configuring the address generator and instruction configuration register in the AI chip.

2. The dynamic-length large model compilation method according to claim 1, wherein Splitting the large model into several operators and performing dynamic operator compilation one by one.

3. A method for compiling a large model with a dynamic length according to claim 1 or 2, characterized in that The operator configuration specifically includes: Configuring the base addresses of the input and output of each operator in the model through the address generator to record the instruction memory access address; Configuring the remainder micro-operators generated by each operator through the instruction configuration register to support dynamic micro-operators.

4. The dynamic-length large model compilation method according to claim 3, wherein The address generator uses one or more groups of registers and counters to record the instruction memory access address.

5. The dynamic-length large model compilation method according to claim 4, wherein, The execution process of the address generator includes: Configuring the starting address and increment rule of the address generator; Binding the address generator during instruction generation; When executing the first instruction, using the starting address of the address generator and clearing the counter; After the instruction execution is completed, incrementing the counter by one; Executing subsequent instructions and generating the instruction memory access address using the counter and the address increment rule.

6. The dynamic-length large model compilation method according to claim 4, wherein, The instruction configuration register is specifically used for: Configuring the remainder micro-operator instruction after the execution count is completed; wherein, the execution count is determined by the AI chip according to the input size.

7. A large model compilation device with a dynamic length, characterized in that The device includes: A splitting module, used to split the trained large model into a set of several operators in a preset mode; wherein, each operator includes multiple micro-operators; A compilation module, used to take the operators in the large model as input with a dynamic length, realize the compilation of dynamic operators, and splice the compiled operators according to the network structure to realize the compilation of the model; wherein, the dynamic length is the length corresponding to each operator; A processing module, used for runtime support, including during model inference, globally adjusting the model operator configuration according to the dynamic length received for this inference, and then performing the inference; wherein, the operator configuration is obtained by configuring the address generator and instruction configuration register in the AI chip.

8. The dynamic-length large model compilation device according to claim 7, wherein, The operator configuration specifically includes: Configuring the base addresses of the input and output of each operator in the model through the address generator to record the instruction memory access address; Configuring the remainder micro-operators generated by each operator through the instruction configuration register to support dynamic micro-operators.

9. The dynamic-length large model compilation device according to claim 8, wherein The instruction configuration register is specifically used for: Configuring the remainder micro-operator instruction after the execution count is completed; wherein, the execution count is determined by the AI chip according to the input size.

Citation Information

Cited By

  • Compiling method and device based on dynamic shape scene, storage medium and electronic device

    CN120723249A