A large model cross-platform system based on micro-operators
Through a cross-platform system based on micro-operators, the problem of large models inference and fine-tuning on different hardware platforms is solved, operator development and debugging are simplified, cross-platform deployment efficiency is improved, and efficient adaptation of hardware platforms is achieved.
Patent Information
- Application Number
- CN202510238510.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing technology cannot quickly realize the inference and fine-tuning of large models on multiple hardware platforms, the operator development, debugging and deployment are difficult, and the compilation scheme takes a long time and poor compatibility.
Design a cross-platform system based on microoperators, including hardware execution layer, hardware access layer, precompilation layer, microoperator layer, backend layer and framework layer, simplify the operator development process, and compatible with different hardware through the unified hardware access layer, the precompilation layer improves compilation efficiency, the microoperator layer defines common operators, the backend layer recognizes data types, and the framework layer provides functional modules to realize large-model inference and fine-tuning.
It significantly reduces the difficulty of developing large-model operators, improves cross-platform deployment efficiency, simplifies the development and debugging process, and improves hardware platform adaptability and compilation efficiency.
Smart Images

Figure CN119718484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital data processing, and in particular to a large model cross-platform system based on micro-operators. Background Art
[0002] With the rapid development of large model technology, the computing power required for training and inferring large models is increasing day by day, which poses higher requirements for efficient and easy-to-use artificial intelligence computing frameworks. Current mainstream artificial intelligence computing frameworks, such as PyTorch, TensorFlow, and Keras, provide tools for building large models. However, these frameworks often rely too much on specific hardware platforms, resulting in many difficulties in their applications on emerging or specific hardware architectures. Emerging or specific hardware architectures include domestic domain specific architectures (DSAs), Field Programmable Gate Arrays (FPGAs), General-purpose computing on graphics processing units (GPGPUs), etc. Currently, the main solutions include two types: one is to manually write a Deep Neural Network (DNN) operator library that conforms to the definition of the mainstream artificial intelligence computing framework for a specific hardware platform; the other is to automatically generate operator code using Multi-Level Intermediate Representation (MLIR) through Just-In-Time (JIT) compilation. However, in the era of large models, both of these solutions face severe challenges.
[0003] For the development of traditional operator libraries, developers need to deeply understand the underlying hardware architecture and algorithm optimization, and the development process is complex and time-consuming. In addition, with the continuous iteration of hardware and algorithms, the operator library needs to be continuously updated and maintained to ensure compatibility with new hardware and frameworks, resulting in high maintenance costs. In addition, traditional operator libraries also need to be adapted to different versions of front-end deep learning frameworks, which involves complex interface adaptation and compatibility issues, further increasing the development difficulty. In addition, debugging end-to-end models requires dealing with the superposition of multi-level complexities, making it extremely difficult to optimize operator performance and debug code. In contrast, although the compilation scheme can automatically generate operator codes for different platforms, there are still problems. For example, the compilation takes a long time, cache management is difficult, and efficiency is easily reduced due to repeated compilation. In addition, the debugging of the compilation scheme is more difficult, and the process of problem location and repair is very complex. Especially for dynamic inputs, the compilation scheme has poor support, and repeated compilation is required for variable-length inputs, making it difficult to meet the high-efficiency inference and fine-tuning requirements of large models. In summary, the prior art cannot provide a system that can quickly implement large model inference and fine-tuning on multiple hardware platforms. This makes the development, debugging, and deployment of large model operators face high difficulties and challenges. Summary of the Invention
[0004] The present invention provides a cross-platform system for large models based on Micro Kernels, aiming to solve the problem that the prior art cannot quickly perform large model inference and fine-tuning across hardware platforms, and thereby reduce the difficulty of operator development, debugging, and deployment. The system architecture proposed by the present invention is simple and efficient, which can significantly reduce the complexity of large model operator development under different hardware platforms and improve the efficiency of development and deployment.
[0005] The present invention provides a large model cross-platform system based on micro-operators, including: a hardware execution layer, a hardware access layer, a pre-compilation layer, a micro-operator layer, a backend layer, a framework layer, and a model layer; wherein, the hardware execution layer is used to execute memory access instructions and / or various micro-operators compiled by the pre-compilation layer; the hardware access layer is used to perform high-level abstraction on various interfaces of each hardware device accessing the large model cross-platform system and form a consistent call interface for the upper layer; the pre-compilation layer is used to call a platform compiler to compile micro-operator code into intermediate instructions or executable instructions of the hardware device; the micro-operator layer is used to call various micro-operators compiled by the pre-compilation layer and provide a programming interface based on micro-operators to the upper layer, so that the upper layer can combine and call multiple micro-operators to implement various operators required by mainstream large models; the backend layer is connected to the framework layer upwards and the micro-operator layer downwards, and is used to expose the function interfaces of the micro-operator layer to the outside. At the same time, the backend layer can identify tensor data types and implement different data types of the same micro-operator compiled by the pre-compilation layer through the micro-operator layer; the framework layer includes various function modules required to implement large model inference and fine-tuning, including a tensor module, a storage module, a device module, a weight loading module, a custom operator module, a neural network module, a weight quantization module, a word vector encoding module, a backpropagation module, and an output sampling module. Each function module realizes the corresponding large model inference function or large model fine-tuning function by calling the function interfaces in the backend layer; the model layer is used to provide a model layer specification for the user to build a large model instance based on the framework layer; wherein, providing a model layer specification for the user to build a large model instance based on the framework layer includes: using the tensor module, neural network module, and custom operator module provided by the framework layer to build a large model instance based on micro-operators according to the mainstream large model structure; using the weight loading module provided by the framework layer to load the large model weights and calling the weight quantization module provided by the framework layer according to the model requirements to perform quantization and inverse quantization of the large model weights; calling the function modules in the framework layer to perform large model inference and fine-tuning.
[0006] The present invention defines the hardware execution layer specification, and any hardware device that conforms to this specification can be connected to the system of the present invention for large model inference and fine-tuning. Hardware devices include but are not limited to CPUs, GPUs, GPGPUs, DSAs, FPGAs, etc.
[0007] Optionally, the hardware execution layer shall have the ability of parallel computing and provide interfaces to support starting a specified number of threads to execute specified memory access instructions and / or various micro-operators compiled by the pre-compilation layer, and support event, asynchronous and streaming processing at the same time. The hardware execution layer supports half-precision (FP16, BF16) or integer type (INT8) for inference and supports higher precision (FP32, TF32, FP64) for fine-tuning. The hardware execution layer supports starting pre-compiled operator code and has a caching function, and supports parameter passing between the host and the device. The hardware execution layer (hardware device) shall have a large-capacity internal high-speed storage (such as high-bandwidth memory HBM) and the ability of high-speed interconnection between the host and the device or between devices, and provide interfaces for memory access operations at the same time. Memory access instructions refer to instructions for data loading operations and data storage operations. Various micro-operators compiled are various different types of operators for computing or operating. Each type of micro-operator is an operator for performing a specified calculation or a specified operation.
[0008] The present invention uses a unified hardware access layer to be compatible with different hardware devices. The hardware access layer performs a higher-level abstraction on various interfaces in the hardware abstraction layers of different hardware manufacturers. Various interfaces are interfaces for performing various operations, including device query interfaces, computing core management / module interfaces, computing core interfaces, event interfaces, streaming interfaces, memory management interfaces, etc. A higher-level abstraction is performed on various interfaces to form call interfaces for calling various hardware functions.
[0009] The present invention defines the pre-compilation layer function. When the large model cross-platform system is compiled into a binary program, the pre-compilation layer calls the platform compiler to compile the pre-written micro-operator code into intermediate instructions or executable instructions for the hardware device. The pre-compilation layer is also used to provide a caching function for pre-compilation instructions. When compiling the binary program of the host-side inference and fine-tuning system, the pre-compilation layer calls the platform compiler to compile the micro-operator code into intermediate instructions or executable instructions for the hardware device, thereby obtaining the micro-operators. Specific error information is provided when the compilation fails. The micro-operator code of each micro-operator is pre-written code that can obtain the corresponding micro-operator after compilation. The intermediate instructions or executable instructions for the hardware device obtained after compiling the micro-operator code of various micro-operators are the various micro-operators. Intermediate instructions refer to the intermediate code generated during the compilation process. Executable instructions for the hardware device refer to instructions that can be executed and take effect on the hardware device connected to the large model cross-platform system. The platform compiler is a pre-set compiler used to compile the micro-operator code into intermediate instructions or executable instructions for the hardware device. The micro-operator code of various micro-operators can be written by technicians and compiled through the pre-compilation layer to generate various micro-operators. The micro-operators compiled by calling the platform compiler need to be compatible with the current hardware platform, have the ability of horizontal transplantation, and have downward compatibility, that is, the new hardware is compatible with the old instructions, or the old code can be recompiled to use the new hardware instructions. The pre-compilation layer provides a caching function for pre-compiled micro-operators, and there is no need to repeat the compilation for the unmodified micro-operator source code. If the pre-compiled micro-operator library has been cached locally and the micro-operator source code has not been modified, this pre-compilation process can be directly skipped. The pre-compilation layer can also provide a parallel compilation function. During the compilation of the host program, multiple micro-operator codes can be pre-compiled simultaneously to improve the compilation efficiency.
[0010] The present invention also discloses a micro-operator layer for defining and implementing various operators required by mainstream large models. The micro-operator layer includes device memory access management micro-operators, tensor calculation micro-operators, and various operators required by mainstream large models implemented by combining tensor calculation micro-operators. Device memory access management micro-operators include device memory allocation micro-operators, device memory release micro-operators, host-to-device communication micro-operators, device-to-host communication micro-operators, device-to-device communication micro-operators, device synchronization micro-operators, and device query micro-operators. Tensor calculation micro-operators include single operand operators, double operand operators, type conversion operators, index operators, general copy operators, conditional selection operators, aggregation operators, and matrix multiplication operators. Tensor calculation micro-operators are combined to implement various operators required by mainstream large models, and various operators required by mainstream large models include rotary position encoding operators, activation layer operators, normalization operators, standardization operators, and attention layer operators.
[0011] Optionally, the micro-operator layer includes device memory access management micro-operators. For example, device memory allocation micro-operator (alloc), device memory release micro-operator (free), host-to-device communication micro-operator (h2d), device-to-host communication micro-operator (d2h), device-to-device communication micro-operator (d2d), device synchronization micro-operator (sync), and device query micro-operator, etc. The micro-operator layer also includes tensor calculation micro-operators. For example, unary operator, binary operator, type conversion operator (cast), indexing operator, general copy operator, conditional selection operator (select), reduction operator (reduce), and matrix multiplication operator (matmul). Among them, the unary operator and the binary operator are general micro-operator names. The unary operator can further include multiple actual unary operators. For example, sine operator (sin), cosine operator (cos), exponential operator (exp), hyperbolic tangent operator (tanh), and rectified linear unit operator (relu). The binary operator can further include multiple actual binary operators. For example, addition operator (add), subtraction operator (sub), multiplication operator (mul), division operator (div), and equality operator (eq), etc. The tensor calculation micro-operators are combined to implement various operators required by mainstream large models. Various operators required by mainstream large models (fused kernels), including Rotary Position Embedding (RoPE), activation layer operator (activation), normalization operator (softmax), normalization operator (layernorm / rmsnorm), and attention layer operators (multiheaded attention, flash attention, paged attention), etc.
[0012] Optionally, the micro-operator layer exposes a micro-operator-based programming interface to the host code. The process of the host code calling micro-operators to perform calculations is the same as a normal function call. The host code combines and calls multiple micro-operators to implement various operators required by mainstream large models. The micro-operator-based programming interface is an interface for combining and calling various micro-operators. For the combined call of micro-operators, stream computing can be adopted, that is, a group of micro-operators execute on the same data stream / computation stream to improve the throughput of hardware execution. For the implementation of various micro-operators, generic programming can be used, that is, the same micro-operator supports calculations of different data types. At the same time, the suffix of the interface name exposed by the micro-operator layer is added with the data type for distinction. For example, the micro-operator cast_f32_f16 converts single-precision floating-point FP32 data to half-precision floating-point FP16 data, and the micro-operator softmax_bf16 normalizes 16-bit floating-point format BF16 data.
[0013] The present invention also discloses a backend layer, which is used to connect to the framework layer upward and the micro-operator layer downward, and at the same time exposes the function interfaces for calling various micro-operator layers outward. Corresponding to the tensor calculation micro-operator, the backend layer exposes the unary operation operator interface (unary_impl), binary operation operator interface (binary_impl), type conversion operator interface (binary_impl), indexing operator interface (indexing_impl), general copy operator interface (general_copy_impl), conditional selection operator interface (select_impl), aggregation operator interface (reduce_impl), and matrix multiplication operator interface (matmul_impl). The backend layer can identify the data type of the tensor data to be processed. For the functional implementation of all micro-operators, the backend layer parses the interface call parameters of the framework layer and identifies the data type of the parameters, and loads the corresponding micro-operators compiled by the pre-compilation layer in the form of kernel launch according to the data type, starts the execution through the hardware execution layer, and encapsulates the calculation results and returns them. The encapsulated calculation results are returned to the backend layer. Corresponding to the device memory access management micro-operator, the backend layer exposes the device memory allocation micro-operator interface (alloc_impl), device memory release micro-operator interface (free_impl), host-to-device communication micro-operator interface, device-to-host communication micro-operator interface, device-to-device communication micro-operator interface, device synchronization micro-operator interface (sync_impl), and device query micro-operator interface.
[0014] Optionally, for different hardware platforms, the corresponding function interfaces (i.e., the backend layer) and the micro-operator code for the corresponding platform can be implemented, so that the framework layer can be accessed through the present invention to realize the inference and fine-tuning of large models.
[0015] Optionally, the backend layer can convert non - contiguous memory blocks / non - contiguous tensors into contiguous tensors by calling a general - purpose copy micro - operator, so as to fully utilize the Single Instruction Multiple Data (SIMD) performance of the hardware and significantly reduce the difficulty of micro - operator development. That is, the backend layer uses the general - purpose copy micro - operator to ensure that other micro - operators only process contiguous tensors, without the need to develop computational kernel code for non - contiguous tensors. To convert all non - contiguous tensor calculations into contiguous tensor calculations, the backend layer also needs to implement the tracking of the tensor layout change process, which is used to record the original and target sizes (shapes) and strides when a contiguous tensor is converted into a non - contiguous tensor. Among them, the size and stride information are the key information for calling the general - purpose copy micro - operator. The backend layer calls different implementations of micro - operators by identifying the tensor data type. For example, add_fp32, add_fp16, and add_bf16 called by the backend layer are the hardware - platform implementations of the addition micro - operator for single - precision (FP32), half - precision (FP16), and half - precision (BF16) respectively. Different data - type implementations of the same micro - operator can refer to different operators generated based on the same micro - operator for performing specified operations or calculations on data of different data types. Each operator performs specified operations or calculations on data of one data type.
[0016] The present invention also discloses a framework layer. The framework layer includes various functional modules required to implement large - model inference and fine - tuning. Specifically, the framework layer includes a tensor module, a storage module, a device module, a weight - loading module, a custom op module, a neural network (layer) module, a weight quantization module (quantization), a tokenizer module, a backpropagation module, and a sampling module, etc.
[0017] Optionally, the tensor module is an abstract representation of tensors in a neural network. The tensor module contains tensor data and various common tensor operations. The various tensor operations are implemented by calling the functional interfaces of the corresponding micro-operators in the backend layer. The tensor module not only holds the corresponding tensor data but also contains common operations applied to tensors. For example, reshape, transpose, squeeze, stack, and various micro-operator operations such as unary, binary, broadcast, matmul, to_contiguous, index_select, reduce, etc. Through the tensor module, the data type, shape, layout, whether it is non-contiguous storage, the hardware type where the tensor data is located, etc. of the tensor can be queried. In addition, the tensor module also overloads various operators (such as +, -, *, / , =, ==, etc.). Therefore, tensors can be conveniently created and initialized through the tensor module, and various operations can be performed on tensors (or between tensors and tensors).
[0018] Optionally, the storage module is a lower-level representation of tensors and tensor operations, used to identify the device category and data type of tensors. The operations of identifying the device category and data type of tensors are implemented by calling the functional interfaces of the corresponding micro-operators in the backend layer. For example, the first tensor is a tensor with a data type of single-precision FP32 stored on a GPU device. Applying a rectified linear unit (ReLU) operation to the first tensor will call the relu_f32 micro-operator in the CUDA backend layer. It should be noted that only tensors in the same device and with the same data type can perform binary or multi-way operations. The second tensor is a tensor with a data type of half-precision FP16 stored on the CPU. It is not feasible to directly add the first tensor and the second tensor. The general copy operator provided by the backend layer needs to be used to move the second tensor from the CPU to the GPU, and then the type conversion operator is called to convert the data type to single-precision floating-point number FP32, and finally the addition operation can be performed with the first tensor. An exception is that for some special operators such as the matrix multiplication operator, the data types of its binary operands can be different. For example, when the matmul_f32_f16 micro-operator performs matrix multiplication, its left operand is single-precision FP32, and the right operand (usually the weight) is half-precision FP16 data. In this case, the operator needs to support data type conversion internally or the hardware supports binary operations of different data types, but even so, the operands must be in the same device.
[0019] Optionally, the device module is a further abstraction of the device functions provided by the backend layer. The device module is used to provide function interfaces required for memory access operations and micro-operator execution operations on different hardware devices. Memory access operations include operations required for tensor allocation, recycling, and inter-device migration on different devices. For example, by calling the micro-operator alloc_uninit of the device module (while specifying that the device type is GPU, the data type is single-precision FP32, and the shape is M×K×N), an uninitialized single-precision FP32 type area with a size of 4×M×K×N is allocated on the GPU to represent the data held by the first tensor. By calling the data migration (to_device) function interface of the first tensor to specify the target device to which the first tensor needs to be transferred, and further calling the function interface copy_to of the device module to implement the migration of the first tensor between different devices. By calling the relu interface of the first tensor to apply an activation function to the first tensor, and further calling the function interface launch of the device module to determine the device and data type of the tensor, and calling the corresponding micro-operator of the backend layer, such as the micro-operator relu_f32 (note that loading and calling a pre-compiled micro-operator such as relu_f32 requires the runtime support of the hardware platform). At the same time, the calculation execution mode can be set through the device module, that is, synchronous execution or asynchronous execution. For synchronous execution (launch), the execution result will be immediately returned after calling the micro-operator relu_f32. For asynchronous execution, a compute stream (stream) for the compute operation needs to be specified, and the upper layer calls the synchronous interface of the device module at an appropriate time to wait for the results of the micro-operator or multiple micro-operators to be returned.
[0020] Optionally, the weight loading module is used to load large model weights from a file. Among them, the weight loading process is triggered by the model layer. The weight loading module sets the target data type of the large model weights when creating the loader, and performs data type conversion on the large model weights during the weight loading process, converting the data type of the large model weights to the target data type. Data type conversion will be dynamically performed during the weight loading process. For example, for a weight file in single-precision FP32, if the set data type is half-precision FP16, then during the weight loading process, for each weight tensor, its function interface to_dtype operation will be called, which will further call the backend layer micro-operator cast_f32_f16 to convert the single-precision FP32 tensor to a half-precision FP16 type tensor. In addition, it should be noted that the weight loader usually loads layer by layer according to the neural network layers, and the loading order and the names of the weights of each layer are generally specified by the upper layer (model layer).
[0021] Optionally, the custom operator module is used to provide function interfaces for extending micro-operators. Among them, the function interfaces for extending micro-operators include extension interfaces for single-operand operators, double-operand operators, and multi-operand operators. Non-standard micro-operators or aggregation operators can be implemented through the function interfaces for extending micro-operators. Non-standard micro-operators can be new micro-operators that need to be added. Aggregation operators can be operators with new functions obtained by aggregating different micro-operators or fusion operators written directly using platform programming tools. When implementing non-standard micro-operators or when aggregation operators need to be implemented, corresponding custom operators can be implemented at different backend layers. Only by specifying the type and parameters of the custom operator can the custom operator interface of the tensor be called to perform custom operator operations. For example, the custom operator RMSNorm is an extended single-operand operator, and its internal implementation is a combination of a set of micro-operators (composed of square root operators, broadcast operators, division operators, multiplication operators, and addition operators, etc.) or a single micro-operator (written by the corresponding hardware platform code, such as the micro-operator CUDA kernel).
[0022] Optionally, the neural network module defines common large model operators. The implementation forms of common large model operators are written using micro-operators or combinations of micro-operators, or are written by calling aggregation operators through the custom operator module. For example, the implementation forms of common large model operators such as linear layer (Linear) operators, activation layer operators, normalization operators, rotary position encoding operators, and attention layer operators are usually divided into two types, namely, directly written using micro-operators or combinations of micro-operators. The linear layer operator can be composed of matrix multiplication operators, broadcast operators, and addition operators, etc. The normalization operator can be composed of square root operators, broadcast operators, division operators, multiplication operators, and addition operators, etc. The rotary position encoding operator and the attention layer operator can be directly written by the hardware platform code. The neural network module is also responsible for loading the corresponding weights. For example, the linear layer operator needs to load the weights required by the matrix multiplication operator.
[0023] Optionally, the weight quantization module is an extension of the tensor module. In addition to providing the standard operations defined by the tensor module, it also provides quantization and dequantization function interfaces to support low-precision large model inference and fine-tuning. The standard tensor module supports operations on common data types such as single-precision FP32, single-precision TF32, half-precision FP16 type, and half-precision BF16. However, for the inference scenario, usually lower precision such as 8-bit signed integer INT8, 4-bit signed integer INT4, and 6-bit floating-point format FP6 is used to improve the operation speed. The weight quantization module provides, in addition to the standard operations defined by the tensor module, additional quantization (quantize) and dequantization (dequantize) function interfaces. The quantization and dequantization function interfaces can perform lossy compression (quantization) on the held tensor data and perform corresponding quantization operations (such as the micro-operator qmatmul_f16_fp6, which performs matrix multiplication on tensors of half-precision FP16 type and tensors of 6-bit precision floating-point FP6 type). For the output result, reverse quantization is performed to restore the precision. It should be noted that the quantization function requires additional micro-operator support. For example, micro-operators such as quantize_f16_fp6, qmatmul_f16_fp6, and dequantize_fp6_f16. The quantization, quantized matrix multiplication, and dequantization functions all require writing corresponding hardware platform code. The quantization function can be accessed at the framework layer in the form of a custom operator. Low-precision large model inference and fine-tuning refer to large model inference and fine-tuning based on the quantized model weight data.
[0024] Optionally, the word vector encoding module is used to implement the word vector encoding operation required for large model inference or large model fine-tuning by calling the micro-operators in the micro-operator layer. The word vector encoding module is used to encode (encode) the input text sequence (sequence) of the model into word vectors (embedding vectors). Word vectors are also called tokens. Among them, there are various implementation forms of the word vector encoding module, and usually algorithms BPE and WordPiece are adopted, which will not be elaborated here.
[0025] Optionally, the backpropagation module is used to implement fine-tuning of the large model. The backpropagation module depends on the tensor module to record the operation types and parameters during the tensor operation process, and calculates the backward gradient through the backward calculation interface of the tensor module; in the backpropagation module, the backward gradient calculation operators corresponding to various tensor operations are defined, and the backward gradient calculation operators corresponding to various tensor operations are implemented by combining various micro-operators or are written and implemented by calling the aggregation operator through the custom operator module. The implementation of the backpropagation module requires the support of the tensor module (recording the operation types and parameters during the tensor operation process for backward gradient calculation), and provides an upward functional interface backward. Among them, the functional interface backward is used to calculate the gradient. The process is to traverse backward the tensor operations applied to it according to the difference between the output tensor and the expected result (or called loss), and calculate the gradient sequentially from back to front according to the loss. In the backpropagation module, the backpropagation operators corresponding to various tensor operations are defined to calculate the gradient, and their implementation forms are similar to the custom operators, and are all implemented by combining various basic micro-operators (such as addition operator, subtraction operator, multiplication operator, exponential operator, square root operator, type conversion operator, negative number operator, etc.) or implemented by the hardware platform code. After the gradient calculation is completed, the upper layer fine-tunes the model weights according to the gradient and the corresponding optimization scheme (for example, the optimization scheme using hyperparameters such as the optimizer and the learning rate).
[0026] Optionally, the output sampling module is used to convert the word vectors generated by the large model into a character sequence. The output sampling module is mainly used to convert the word vectors (tokens) generated by model inference into corresponding character representations. Common sampling algorithms include the top-k sampling algorithm, the argmax sampling algorithm, and the multinomial sampling algorithm. Similar to the word vector encoding module, the output sampling module can be implemented using industry standards, so it will not be elaborated here.
[0027] The present invention also discloses a model layer specification for providing a large model instance construction to the user based on the framework layer. The model layer specification for providing a large model instance construction to the user based on the framework layer includes: using the tensor module, neural network module, and custom operator module provided by the framework layer to build a micro-operator-based large model instance according to the mainstream large model structure; using the weight loading module provided by the framework layer to load the large model weights and calling the weight quantization module provided by the framework layer for quantization and dequantization of the large model weights according to the model requirements; calling the functional module in the framework layer for inference and fine-tuning of the large model. The micro-operator-based large model instance refers to various large models written according to the mainstream large model structure according to the model writing specification of the present invention.
[0028] An embodiment of the present invention provides a large model cross-platform system based on micro-operators, which can be used for quickly performing inference and fine-tuning of large models under different hardware platforms, and is applicable to quickly developing, debugging, inferring, and fine-tuning large models under different hardware platforms. It simplifies the development difficulty of large model operators, has a simple system structure, and is conducive to model development, debugging, and cross-platform deployment. It defines the hardware access method, the implementation form and call method of micro-operators, and each functional module required for realizing large model inference and fine-tuning based on micro-operators. Its overall system structure design is simple, which can effectively reduce the operator development difficulty and improve the development and deployment efficiency of large models.
[0029] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0031] Figure 1 It is a schematic structural diagram of a large model cross-platform system based on micro-operators provided by an embodiment of the present invention.
[0032] Figure 2 It is a schematic diagram of the example directory structure of the pre-compilation layer provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "target", "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising", "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] The present invention proposes a large model inference and fine-tuning technical solution based on micro-operators. Compared with the traditional operator library solution and the JIT just-in-time compilation solution, the micro-operators defined in the present invention are simple to develop, have high flexibility and scalability. At the same time, its pre-compilation ability solves the problems of long code generation time, low compatibility and poor dynamic supportability of the just-in-time compilation solution, thus effectively reducing the adaptation difficulty of mainstream large models on different hardware platforms.
[0036] Based on the micro-operator solution, the present invention proposes a complete end-to-end inference and fine-tuning solution, including the form of hardware access, the definition and call method of micro-operators, the adaptation methods for different backends, and the calculation framework and model building solution based on micro-operators, etc. Its overall system design structure is simple, with low development difficulty and high compatibility, solving the problems of high operator development and maintenance costs and poor cross-platform compatibility in mainstream deep learning systems.
[0037] The experimental results of the present invention on domestic DSA hardware platforms show that the solution of the embodiments of the present invention can significantly reduce the difficulty of large model operator development. The large model inference system constructed based on this solution can easily be compatible with and adapt to a variety of mainstream large models and multi-modal scenarios, and at the same time, its inference performance is comparable to that of the traditional operator library solution (by professional operator development teams, extreme optimization operators, mainstream deep learning frameworks). Therefore, the present invention has significant technical advantages and can bring high commercial value.
[0038] Figure 1Schematic diagram of a large model cross-platform system based on micro-operators provided by an embodiment of the present invention. This embodiment is applicable to the situation of reasoning and fine-tuning of large models for enterprise use. Mainstream large models include large language models (LLMs), large vision models (VLMs), vision-language models (VLMs), and large multimodal models (LMMs). The large models used by enterprises are large language models for machine translation or intelligent question answering, large vision models for face recognition, and large multimodal models for scene perception, etc. As Figure 1 shown, the large model cross-platform system based on micro-operators specifically includes: a hardware execution layer 101, a hardware access layer 102, a pre-compilation layer 103, a micro-operator layer 104, a backend layer 105, a framework layer 106, and a model layer 107. Their structures and functions will be described below.
[0039] Among them, the hardware execution layer 101 is used to execute memory access instructions and / or various micro-operators compiled by the pre-compilation layer 103. The hardware access layer 102 is used to perform high-level abstraction on various interfaces of each hardware device accessing the large model cross-platform system, and form a consistent call interface for the upper layer. The pre-compilation layer 103 is used to call the platform compiler to compile the micro-operator code into intermediate instructions or executable instructions for the hardware device. The micro-operator layer 104 is used to call various micro-operators compiled by the pre-compilation layer 103 and provide a programming interface based on the micro-operators to the upper layer, so that the upper layer can combine and call multiple micro-operators to implement various operators required by mainstream large models. The backend layer 105 interfaces with the framework layer 106 upwards and the micro-operator layer 104 downwards, and is used to expose the functional interfaces of the micro-operator layer 104 to the outside. At the same time, the backend layer 105 can recognize the tensor data type, and different data types of the same micro-operator compiled by the pre-compilation layer 103 are implemented through the call of the micro-operator layer 104. The framework layer 106 includes various functional modules required for implementing large model inference and fine-tuning, including a tensor module, a storage module, a device module, a weight loading module, a custom operator module, a neural network module, a weight quantization module, a word vector encoding module, a backpropagation module, and an output sampling module. Each functional module implements the corresponding large model inference function or large model fine-tuning function by calling the functional interfaces in the backend layer 105. The model layer 107 is used to provide a model layer specification for the user to build a large model instance based on the framework layer 106; among them, providing a model layer specification for the user to build a large model instance based on the framework layer 106 includes: using the tensor module, neural network module, and custom operator module provided by the framework layer 106 to build a micro-operator-based large model instance according to the mainstream large model structure; using the weight loading module provided by the framework layer 106 to load the large model weights and calling the weight quantization module provided by the framework layer 106 for quantization and de-quantization of the large model weights according to the model requirements; calling the functional modules in the framework layer 106 for large model inference and fine-tuning.
[0040] Taking the access of a domestic DSA hardware platform (Tencent's GCU platform) to the large model cross-platform system based on micro-operators provided by the present invention as an example, various aspects of the present invention will be explained in detail.
[0041] The pre-compilation layer is used to call the platform compiler to compile the micro-operator code into intermediate instructions or executable instructions for the hardware device. The function of the pre-compilation layer is to pre-compile the written micro-operator code into intermediate code or binary code (also known as kernel code) that can be loaded and executed by the hardware platform before the program execution. The corresponding compiler needs to be provided by the hardware platform where it is located. The implementation of the micro-operator is written with reference to the programming interface defined by the hardware platform where it is located. For example, for the TPU DSA hardware platform, the operator is written using the extended C++ programming language, and the compiler topscc (TPU platform compiler) is provided to compile the C++ operator source code into assembly code (executable code for the DSA hardware platform).
[0042] The implementation of the pre-compilation layer can be written in the Rust programming language, and the Cargo manager of the Rust programming language is used to create a new project. Among them, the written operator code file is placed in the src directory of the project or a separate directory (such as Figure 2 the kernels directory shown), and the compilation of the operator code is processed in the project compilation file build.rs. The project compilation file build.rs is called during the project compilation process to handle the pre-compilation related logic. Figure 2 It is a schematic diagram of the example directory structure of the pre-compilation layer provided by the embodiment of the present invention. The example directory structure of the pre-compilation layer is an exemplary directory structure for storing the operator code files in the pre-compilation layer.
[0043] In the project compilation file build.rs, the modification status of each file in the operator directory can be detected and compared with the generated operator binary file. If the operator binary code is not generated, the operator file is directly compiled. Otherwise, it is judged whether the modification status of the operator file is later than the generation time of the corresponding binary file. If it is later than the binary file, it means that the operator file has been modified since the last pre-compilation and needs to be recompiled. Otherwise, the pre-compilation of the operator file is skipped. Before the pre-compilation of the operator code, that is, before calling the platform compiler, relevant environment variables and the header files on which the operator code depends also need to be set, and the platform identifier is passed to the compiler according to the characteristics of the hardware platform to correctly compile the operator binary file that conforms to the current hardware platform. To accelerate the pre-compilation, a thread pool can be used to start the pre-compilation work of multiple operator codes simultaneously and wait for the completion of all compilation tasks. In addition to the on-demand compilation of the operator code, the pre-compilation layer needs to prompt the corresponding error message when the compilation fails, usually by returning the compiler error message.
[0044] In addition to containing intensive computing operators such as matrix multiplication operators and two-dimensional convolution operators (conv2d), the micro-operator layer usually also contains single-operand operators, double-operand operators, affine transformation operators (affine), type conversion operators, general copy operators, indexing operators, aggregation operators, ternary operation operators (ternary), fill operators (fill), etc. Operators required for large model inference and fine-tuning, such as rotary position encoding operators, attention operators, tensor masked fill operators (masked_fill), etc., can all be composed of the above-mentioned micro-operators. Users can also use the custom operators mentioned in the present invention for extension to achieve the development of fused operators.
[0045] The micro-operator indexing (index_select) in the micro-operator layer can be written on the DSA hardware platform. As shown below, the operator code of micro-operator indexing is declared as __device__ code, which uses templates to support type generics. Micro-operator indexing selects corresponding data from the input data according to the given condition sequence and copies it to the output data.
[0046] template <typename ID_TYPE, typename T>
[0047] __device__ void index_select_kernel(size_t num_ids, ID_TYPE *ids,
[0048] T *in, T *out, size_t l_size, size_t dim_size, size_t r_size) {
[0049] int thread_id = GetThreadIdx();
[0050] __local__ ID_TYPE ids_buffer[MAX_IDS_SIZE];
[0051] tops_dte_ctx_t ctx;
[0052] dte_scope s(ctx);
[0053] memcpy(ctx, mdspan(Private, ids_buffer, num_ids), mdspan(Global, ids,num_ids));
[0054] int THREAD_STEP = 1;
[0055] int thread_step = 1;
[0056] / / determine thread steps
[0057] get_thread_steps(THREAD_STEP, thread_step, num_ids);
[0058] for (int i = 0; i < thread_step; i++) {
[0059] int idx = thread_id * THREAD_STEP + i;
[0060] for (int j = 0; j < l_size; ++j) {
[0061] int _idx = ids_buffer[idx];
[0062] mdspan hbm_in(Global, in + (j * dim_size + _idx) * r_size, r_size);
[0063] mdspan hbm_out(Global, out + (idx + j * num_ids) * r_size, r_size);
[0064] memcpy(ctx, hbm_out, hbm_in); / / copy selected buffer
[0065] }
[0066] }
[0067] }
[0068] To support multiple data types, the micro-operator layer can be extended using C++ models and generic programming methods. Ten different types of micro-operators index_select can be defined as follows. For example, the micro-operator is_u32_bf16 represents the index_select micro-operator with the given data type of uint32 (u32) and the input type of bfloat16 (bf16); the micro-operator is_u8_f16 represents the index_select micro-operator with the given data type of uint8 (u8) and the input type of float16 (f16); the micro-operator is_u32_f32 represents the index_select micro-operator with the given data type of uint32 (u32) and the input type of float32 (f32).
[0069] It should be noted that the external interfaces of the micro-operator layer are all defined by the implementation interfaces of the micro-operators. For example, the definition of the micro-operator index_select is as follows:
[0070] #define IS_OP(TYPE, ID_TYPE, FN_NAME) \
[0071] extern "C" __global__ void FN_NAME( size_t id_num, \
[0072] ID_TYPE* ids, TYPE *in, TYPE *out, \
[0073] size_t l_size, size_t dim_size, size_t r_size) \
[0074] { \
[0075] index_select_kernel<ID_TYPE, TYPE> \
[0076] (id_num, ids, in, out, l_size, dim_size, r_size); \
[0077] } \
[0078] IS_OP(__bf16, uint32_t, is_u32_bf16)
[0079] IS_OP(__fp16, uint32_t, is_u32_f16)
[0080] IS_OP(float, uint32_t, is_u32_f32)
[0081] IS_OP(uint8_t, uint32_t, is_u32_u8)
[0082] IS_OP(uint32_t, uint32_t, is_u32_u32)
[0083] IS_OP(__fp16, uint8_t, is_u8_f16)
[0084] IS_OP(__bf16, uint8_t, is_u8_bf16)
[0085] IS_OP(float, uint8_t, is_u8_f32)
[0086] IS_OP(uint8_t, uint8_t, is_u8_u8)
[0087] IS_OP(uint32_t, uint8_t, is_u8_u32)
[0088] Among them, the input and output (in, out) are of the same data type. Given that the input condition ids is of integer data type and other parameters are of size_t type. In a specific instance, when the input and output types are bf16 and the given input condition type is u32 type, the external interface of this micro-operator is:
[0089] extern "C" __global__ void is_u32_bf16 (size_t id_num,
[0090] uint32_t * ids, __bf16 *in, __bf16 *out, size_t l_size,
[0091] size_t dim_size, size_t r_size)
[0092] Correspondingly, the backend layer provides the interface of the micro-operator index_select. If it is determined that the types of the input and output data and the input condition match, the micro-operator is_u32_bf16 will be called in the form of a kernel launch.
[0093] Multiple micro-operators can be combined arbitrarily to form the operators required for large model inference and fine-tuning. For example, the commonly used normalization (layernorm / rmsnorm) operators in large models are implemented using combinations of micro-operators as follows:
[0094] impl Forward for LayerNorm {
[0095] fn forward(&self, x: &Tensor) -> Tensor {
[0096] let in_dtype = x.dtype();
[0097] let scale = 1 / x.dim(-1) as f64;
[0098] let x = x.to_dtype(Type::F32);
[0099] let x = if self.rms_norm{ x} else {
[0100] let mean = x.reduce_sum(-1) * scale;
[0101] x.sub(&mean)
[0102] };
[0103] let norm_x = x.sqr().reduce_sum(-1) * scale;
[0104] let x = x.div(&(norm_x + self.eps).sqrt());
[0105] let x = x.to_dtype(in_dtype).mul(&self.weight);
[0106] return x.add_if(&self.bias);
[0107] }
[0108] }
[0109] Among them, to ensure accuracy, the input data type is first converted to a high-precision representation (FP32), and then the combination of single-operand operators and double-operand operators (such as subtraction operator, square root operator, division operator, square root operator, addition operator, multiplication operator, etc.) is used for calculation according to the operator category (whether it is rmsnorm). The micro-operator Rmsnorm is a simplified normalization operator that removes the operation of subtracting the mean value. It should be noted that the type conversion to_dtype operation will call the type conversion operator. For example, if the input type is FP16 and the output type is FP32, the type conversion operator cast_f16_f32 will be called. Similarly, other micro-operator calls, such as subtraction operator, addition operator, and division operator, will also call the relevant micro-operator implementations of the type.
[0110] Different hardware platforms need to implement their specific functions according to the interfaces defined by the backend layer. The backend layer is called through the interfaces of the backend layer, and the pre-compiled computing kernels of the hardware platform are loaded and called according to the characteristics of the hardware platform. Usually, this kind of call is carried out in the form of kernel launch. For example, the micro-operator index_select is called in the form of kernel launch as follows:
[0111] impl IndexSelect {
[0112] fn call<T: DType>(&self, src: &GcuSlice <t>, src_l: &Layout,
[0113] dim: usize, dev: &GcuDevice) -> GcuSlice <t>{
[0114] let ids_l = &self.id_layout;
[0115] let ids_shape = ids_l.shape();
[0116] let ids_el = ids_shape.elem_count();
[0117] let src = match src_l.offsets() {
[0118] Some((o1, o2)) => src.slice(o1..o2),
[0119] _ => Err("index-select data must be contiguous!"),
[0120] };
[0121] let l_size: usize = src_l.dims()[..dim].iter().product();
[0122] let r_size: usize = src_l.dims()[dim + 1..].iter().product();
[0123] let dim_size = src_l.dims()[dim];
[0124] let out = dev.alloc:: <t>(ids_el * l_size * ri_size);
[0125] match &self.slice {
[0126] GcuStorageSlice::U32(slice) => {
[0127] let ptr = slice.slice(ids_l.offset()..);
[0128] let func = dev.load_func(&name:: <t>("is_u32"));
[0129] let params = (ids_el, ptr.device_ptr(),
[0130] src.device_ptr(), out.device_ptr(),
[0131] l_size, dim_size, r_size);
[0132] func.launch(&dev.launch_cfg, params);
[0133] }
[0134] GcuStorageSlice::U8(slice) => {
[0135] let ptr = slice.slice(ids_l.offset()..);
[0136] let func = dev.get_or_load_func(&name:: <t>("is_u8"));
[0137] let params = (ids_el, ptr.device_ptr(),
[0138] src.device_ptr(), out.device_ptr(),
[0139] l_size, dim_size, r_size);
[0140] func.launch(&dev.launch_cfg, params);
[0141] }
[0142] _ => Err("index_select ids should be u8 or u32"),
[0143] };
[0144] out
[0145] }
[0146] }
[0147] Among them, first, the call parameters are organized according to the interface requirements of the micro-operator index_select, including the input and output data types. Then, the output buffer is allocated using the device interface passed in by the backend layer (dev.alloc). Finally, the compute kernel is loaded and the computation is launched (dev.load_func, dev.launch(&func)).
[0148] To simplify the development difficulty of the operator, before the backend layer calls the micro-operator layer to perform the computation, it will convert all non-contiguous storage tensors into contiguous storage tensors. This conversion process is jointly implemented by a general copy operator and the backend layer. Specifically:
[0149] The micro-operator layer implements a general copy operator, which can convert the input non-contiguous storage tensor into a contiguous tensor output according to the input and output target sizes (shape) and strides (stride) information. The specific conversion process is as follows: for each output element, its position (index) in the input data is calculated according to the target size (shape) and stride (stride) information, and the data is copied to the corresponding output position according to the position. The core index calculation logic is as follows:
[0150] template <int RANK>
[0151] __device__ VType get_strided_index(VType &indexes, VType results,
[0152] VType dst_shape[], VType dst_strides[]) {
[0153] VType vec_rem[RANK];
[0154] vec_rem[0] = indexes;
[0155] for (int i = 0; i < RANK; i++) {
[0156] unsigned int dim_idx = RANK - 1 - i;
[0157] results = vadd(vmul(vrem(vec_rem[i], dst_shape[dim_idx]),
[0158] dst_strides[dim_idx]), results);
[0159] vec_rem[i + 1] = vdiv(vec_rem[i], dst_shape[dim_idx]);
[0160] }
[0161] return results;
[0162] }
[0163] As shown above, DSA vectorized instructions such as vadd, vmul, vdiv, etc. can be used to accelerate the calculation. The application of these instructions enables the position calculation to be carried out in batches. For example, each time 128 output positions are given (the length of indexes is 128, such as 0, 1, 2, 3... 127), the function get_strided_index returns 128 calculation results, that is, the input element positions where these 128 elements are located. Then, the gather (vgather) instruction is used to batch obtain data from the input and write it into the corresponding output positions at one time to complete the conversion of the non-continuous tensor to the continuously stored tensor.
[0164] Before calling each operator to perform calculations, the backend layer first uses a general copy operator to convert all tensors with non - contiguous storage into tensors with contiguous storage, and then uses the tensors with contiguous storage as input data to call the corresponding operator for implementation. For example, as shown in the following code, first, it is judged whether the layout of the input data (self.slice) of the unary operation is non - contiguous. For the input data with non - contiguous storage, temporary data is first allocated, and then the general copy operator is called to convert the input data into data with contiguous storage and place it in the temporary data dst. At the same time, the memory layout is modified to be contiguous. Finally, the specific implementation of the unary operation is called with the temporary data dst as the input.
[0165] fn unary_impl<T: UnaryOp>(&self, layout: &Layout) -> Self {
[0166] let device = self.device().clone();
[0167] if layout.is_contiguous() {
[0168] let slice = T::F.map(&self.slice, &device, layout);
[0169] Self { slice, device}
[0170] } else {
[0171] let mut dst= device.alloc(layout.shape(), self.dtype());
[0172] self.general_copy(&mut dst, 0, layout);
[0173] let slice = T::F.map(&dst.slice, &device, &layout.contiguous());
[0174] Self { slice, device}
[0175] }
[0176] }
[0177] As described above, the framework layer interfaces with the backend layer downward and provides a functional interface for writing model instances upward. Except that both word vector encoding and output sampling use industry standards, embodiments of other modules in the backend layer are introduced separately below.
[0178] 1) Tensor module: This module holds tensor data and provides various computational interfaces for operating on tensors. To represent different types of tensors, its structure is designed with a storage member (managed by the storage module) to store tensor data, a device member to indicate the device where the tensor is stored, a dtype member to indicate the tensor's data type, a layout member to indicate the tensor's memory layout, and a backprop member to record the forward operations applied to the tensor (used for backpropagation and fine-tuning). To support multithreaded access, the storage member of a tensor is protected by a read-write lock, allowing only one thread to write at a time but multiple threads to read simultaneously. Tensors can be converted between different data types or devices. For example, calling the to_dtype operation on the storage member invokes the to_dtype operation, which further calls the cast operator in the backend layer to convert the tensor's storage type to the target type. All other properties of the tensor, except the data type, remain unchanged. Calling the to_device operation matches the target device with the current device of the tensor. If they are on the same device, only the tensor is copied. If they are on different devices, the host-device copy of the tensor storage member is triggered. In cases where device-to-device (D2D) copying cannot be completed, CPU memory is used as a relay. As mentioned earlier, the tensor module provides encapsulation for various interface calls, such as common reshape, transpose, squeeze, stack, and various micro-operator operations such as unary, binary, broadcast, matmul, to_contiguous, index_select, reduce, etc. These operations will be called to different backend layers based on the storage type. For example, the GCU storage member will call the GCU backend layer to trigger the kernel launch on the Suiyuan DSA device. The layout attributes of the tensor, such as shape and stride, will change after operations such as reshape, broadcast, and reduce. For non-inplace operations, a copy after the operation is usually returned, and the original tensor operation remains unchanged. To implement the to_contiguous function, which converts non-contiguously stored tensors into contiguous ones, the tensor module also records the layout changes. This allows the generalcopy operator to be passed the original and target shape and stride information when calling to_contiguous . Furthermore, in addition to querying the data type, shape, layout, and device type, the tensor module also allows querying whether a tensor is non-contiguous (by determining its shape and stride information).The tensor module overloads various operators (such as +, -, *, / , =, ==, etc.), and its implementation is similar to that of ordinary binary and unary operations, which will not be elaborated here one by one.
[0179] 2) Storage module: It is a lower-level representation of tensors and tensor operations. It directly interfaces with the backend layer and calls different backend layers and different (data type-specific) micro-operators by identifying the device type and data type of the tensors. The storage module can be defined as an enumeration type in the Rust programming language as follows. Here, CpuStorage, CudaStorage, and GcuStorage represent the specific implementations for the CPU, GPU, and GCU platforms respectively. Public interfaces can be defined in storage, and different platforms implement the corresponding interfaces to complete the call to the backend layer of that platform. Similarly, to distinguish the device type and data type where storage is located, storage also has device type and data type (dtype) attributes. Storage supports data copying between different devices to complete the to_device function, and its to_dtype function is responsible for calling the cast operations of different backends to complete the conversion of data types. It should be noted that tensor calculations need to be performed on the same device because each tensor calculation is completed in the form of kernel launch by its specific storage (specific backend), and different backends are isolated from each other. If the binary or multi-operand (tensors) are on different devices, all tensors can be pulled to the same storage, such as GcuStorage, through the to_device operation, and then the calculation interface of the tensor is called (that is, by calling the corresponding interface of GcuStorage to call the gcu backend and start the gcu kernel launch). It should be noted that except for to_device, the storage type returned by tensor calculations is always the same as the source input.
[0180] pub enum Storage {
[0181] Cpu(CpuStorage),
[0182] Cuda(CudaStorage),
[0183] Gcu(GcuStorage),
[0184] }
[0185] 3) Device Module: It is a further abstraction of the device functions provided by the backend layer, and provides the functional interfaces required for device memory access and micro-operator execution (kernel launch). Generally speaking, the device module is a simple encapsulation of the runtime system (such as cudart) provided by the hardware platform. The basic runtime system functions should comply with the hardware execution layer specifications of the present invention, and generally should provide functions such as creating a device, creating a data stream / computation stream (stream), loading kernel code and returning the function address (kernel function), executing kernel tasks (kernel launch), synchronization (sync), device memory allocation (alloc), device memory recycling (free), and communication between host and device memory (memcpy, d2d: device to device, d2h: device to host, h2d: host to device). It should be noted that the runtime system needs to support both synchronous and asynchronous operations, that is, when the kernel task or memory access operation passes in data stream information or computation stream information, all operations should be executed asynchronously on the data stream or computation stream (return immediately after calling the kernel task or memory operation), and when no data stream information or computation stream information is passed in, all operations are executed synchronously. When executing asynchronously, the upper layer controls the synchronization timing of the data stream information or computation stream information. Generally for stream computing, the runtime system and the corresponding hardware platform need to provide a cache function, that is, during the execution of a group of operators, the intermediate results do not need to be frequently written back to the storage device (such as HBM), but are temporarily stored through a cache such as L2 cache to accelerate stream computing.
[0186] 4) Weight Loading Module: It is mainly used to load the pre-trained network model weights from model weight files, such as Pytorch pth files, Huggingface Safetensors files, or Numpy npy files. The model layer triggers the loading of the weight file when building a large model. Usually, the loading process is carried out layer by layer according to the network model, and a unique weight name, such as weight name layer1.weight1, is provided by the model layer during loading. Different loading methods can be adopted according to the type of the weight file. Taking the safetensors format of the standard model weights in the open-source community and the Huggingface platform as an example, its loading process can adopt the form of memory mapping, that is, mapping the model weight file to the physical host memory and converting the mapped model weight file into a memory object (such as an instance of safe tensor) in a deserialized manner. The deserialization of the instance safetensor first reads the metadata information according to the mapped memory address, and then reads the corresponding tensor data according to the metadata information (such as tensor shape, tensor data type). When the tensor is loaded into the memory, it is decided whether to convert the data type of the weight tensor according to the target data type. It should be noted that the memory mapping method for loading the weight file does not require loading all the weights into the host memory at one time. When each layer of the model loads the weight, it triggers the deserialization, data type conversion (if necessary), and copying from the host to the device memory (if the weight is specified to be stored in the device memory) of the corresponding weight file of this layer. When all the model layers are loaded, the weight file of the model is loaded into the specified device memory, and at this time the model can be executed on this device.
[0187] 5) Custom Operator Module: It provides a functional interface for extending micro-operators. That is, when the user considers accessing a new fused operator or an operator not supported by the system of the present invention, the operator library can be extended through custom operators (custom op). The steps to extend the operator library are as follows: First, the user writes the corresponding operator code using the programming interface provided by the hardware platform and accesses the pre-compilation layer to generate the corresponding operator binary file; Second, define a new operator type in the system of the present invention according to the custom operator specification; Finally, implement the forward calculation function of the custom operator, that is, load and call the newly written operator in the form of starting a computing kernel in the corresponding backend layer. Taking the custom operator of binary operation as an example below, name() in the CustomOp interface specifies the operator name of this custom operator, and forward and backward are the forward calculation interface and the backward calculation interface of the custom operator respectively.
[0188] pub trait CustomOp {
[0189] fn name(&self) -> String;
[0190] fn forward(&self, _: &GcuStorage, _: &Layout, _: &GcuStorage,
[0191] _: &Layout) -> (GcuStorage, Shape) {
[0192] todo!()
[0193] }
[0194] fn backward(&self, _arg1: &Tensor, _arg2: &Tensor, _res: &Tensor,
[0195] _grad_res: &Tensor) -> (Option <tensor>, Option <tensor>) {
[0196] todo!()
[0197] }
[0198] }
[0199] In addition, in order to integrate the custom operator into the tensor module of the system of the present invention, it is necessary to complete the call to the custom operator in the implementation of the tensor, and it is further called by the storage module to the backend layer and the micro-operator layer for implementation.
[0200] impl Tensor {
[0201] pub fn custom_forward<T: CustomOp>(&self,
[0202] rhs: &Tensor, c: &T) -> Tensor {
[0203] let (storage, shape) = self.storage()
[0204] .forward(self.layout(), &rhs.storage(), rhs.layout(), c);
[0205] from_storage(storage, shape, false)
[0206] }
[0207] pub fn custom_backward<T: CustomOp>(&self,
[0208] arg1: &Tensor, arg2: &Tensor,
[0209] res: &Tensor, grad: &Tensor, c: &T) -> Tensor {
[0210] let (storage, shape) = self.storage()
[0211] .backward(arg1, arg2, res, grad, c);
[0212] from_storage(storage, shape, false)
[0213] }
[0214] }
[0215] In a specific example, a new custom operator KVConcat is shown as follows. It mainly completes the concatenation operation of keys and values during the forward calculation of the large model. Among them, the dimension of the concatenation operation is specified by the concatenation dimension (concat_dim).
[0216] pub struct KVConcat {
[0217] pub concat_dim: i32,
[0218] }
[0219] pub fn kvconcat(lhs: &Tensor, rhs: &Tensor, dim: i32) -> Tensor {
[0220] let op = KVConcat { concat_dim: dim};
[0221] lhs.custom_forward(&rhs, op)
[0222] }
[0223] According to KVConcat, the kvconcat function can be further defined. It accepts two tensor inputs and a concatenation dimension (concat_dim). This function finally calls the functional interface custom_forward of the tensor module, that is, the custom_forward functional interface of the above tensor module. After the new operator completes the function and interface definition, the call to the new operator's compute kernel (kvconcat kernel launch) can be implemented in the corresponding backend layer, as shown below:
[0224] impl CustomOp for KVConcat {
[0225] fn name(&self) -> String { "kvconcat".to_string()}
[0226] fn forward(&self, l: &GcuStorage, l_l: &Layout,
[0227] fn(&mut GcuStorage, &Layout) -> (GcuStorage, Shape) {
[0228] let dev = &l.device;
[0229] let cfg = &dev.launch_cfg;
[0230] let elem_count =
[0231] l_l.shape().elem_count() + r_l.shape().elem_count();
[0232] let dims = l_l.shape().dims().len();
[0233] let ds = dev
[0234] .htod_copy([l_l.shape().dims(), r_l.shape().dims()]
[0235] .concat());
[0236] let slice = match (&l.slice, &r.slice) {
[0237] (GcuStorageSlice::BF16(l_), GcuStorageSlice::BF16(r_)) => {
[0238] let out = dev.alloc:: <bf16>(elem_count);
[0239] let func = dev.load_func("kvconcat_bf16");
[0240] let params = ( l_.device_ptr(), r_.device_ptr(),
[0241] out.device_ptr(), ds.device_ptr(),
[0242] dims, self.concat_dim);
[0243] func.launch(cfg, params);
[0244] GcuStorageSlice::BF16(out)
[0245] }
[0246] (GcuStorageSlice::F16(l_), GcuStorageSlice::F16(r_)) => {
[0247] let out = dev.alloc:: <f16>(elem_count);
[0248] let func = dev.load_func("kvconcat_f16");
[0249] let params = (l_.device_ptr(), r_.device_ptr(),
[0250] out.device_ptr(), ds.device_ptr(),
[0251] dims, self.concat_dim);
[0252] func.launch(cfg, params);
[0253] GcuStorageSlice::F16(out)
[0254] }
[0255] _ => InternalError("dtype mismatch in kvconcat op"),
[0256] };
[0257] let dst_shape = concat_shape(l_l, r_l, self.concat_dim);
[0258] (GcuStorage { slice, device: dev.clone()}, dst_shape.into())
[0259] }
[0260] }
[0261] As shown above, in the implementation of the forward calculation function, the kernel call parameters are first organized, and then the kernel function instance to be loaded is determined according to the data type of the tensor. For example, for an FP16 input tensor, the kernel kvconcat_f16 is loaded. Finally, the call is completed through the kernel launch function of the device module and the calculation result is returned. It should be noted that the input parameters (parameter type, order, quantity) in the kernel need to be consistent with the written kernel interface. The main implementation logic of the kernel Kvconcat is to extract the corresponding slices of the two input tensors according to the specified concatenation dimension and copy them to the specified area of the output tensor through a device-to-device (d2d) copy operation.
[0262] 6) Neural network module: Defines common large model operators. For example, operators such as linear layer (Linear) operator, activation layer operator, normalization operator, rotary position encoding operator, attention operator, etc. The operator implementation forms of the neural network module are divided into two types. One is to directly use micro-operators or combinations of micro-operators for writing, such as the normalization operator or the linear layer operator. The implementation of the linear layer operator (shown below) calls the matrix multiplication operator, transpose operator, and addition operator. The implementation of other neural network layers is similar, that is, the Forward (forward calculation) interface needs to be implemented. For inplace operations, the InPlaceForward forward calculation interface needs to be implemented, and its forward calculation function directly modifies the input tensor without returning a modified copy. It should be noted that the operators or neural network layers defined by the neural network module can use the operator functions provided in the tensor module, that is, the corresponding backend call is determined by the device where the tensor is located without explicitly specifying the backend. At the same time, it can also directly load and execute the computing kernels of the specified hardware platform in the form of computing kernel startup, similar to the implementation of custom operators.
[0263] pub struct Linear {
[0264] weight: Tensor,
[0265] bias: Option <tensor>,
[0266] }
[0267] impl Forward for Linear {
[0268] pub fn forward(&self, x: &Tensor) -> Tensor {
[0269] let x = x.matmul(&self.weight.transpose(1, 2));
[0270] match &self.bias {
[0271] Some(b) => x.add(&b),
[0272] None => x,
[0273] }
[0274] }
[0275] }
[0276] 7) Weight quantization module: It is an extension of the tensor module. The standard tensor module supports common data types such as FP16, BF16, FP32, etc. Lower precision data types such as INT8, INT4, FP6, etc. are often used to accelerate large model inference. To support the loading of low-precision model weights (or called quantized weights) and low-precision inference of large models, the weight quantization module needs to implement the following functions: First, the weight loading module needs to support loading low-precision weights. The common file formats for low-precision weights of large models are GGML / GGUF and GPTQ. The GGML / GGUF file format is further divided into Q4 (Q4_0, Q4_1, and Q4_K file formats), Q5 (Q5_0, Q5_1, and Q5_K file formats), Q8 (Q8_0, Q8_1, and Q8_K file formats), etc. according to the quantization bit width. Therefore, the weight loading module needs to implement the dequantization of quantized weights. Second, the compute kernel (usually the matrix multiplication operator, that is, quantizing the weights of the matrix multiplication operator or the fully linear layer operator) needs to support low-precision matrix operations or support real-time dequantization of low-precision weights (dequantizing the quantized weights to a higher precision and then performing high-precision matrix operations). Taking the file format Q8_0 in the common file format GGML / GGUF as an example, its Rust language definition of the data format is:
[0277] #[repr(C)]
[0278] pub struct Q8_0 {
[0279] public scale: f16,
[0280] public chunk: [i8; 32],
[0281] }
[0282] That is, a Q8_0 data contains a quantization scale factor (scale) member of FP16 and 32 INT8 elements. The quantization process of Q8_0 is to quantize and scale a given high-precision tensor (requiring 32-element alignment) in groups of 32 elements (chunk) to the INT8 representation range (-127 to +127), where the scale factor is the scale member in Q8_0 (if scaled down by 8 times, the scale is 8.0). In this way, only 32 * 1 byte + 2 bytes = 34 bytes are needed to represent 32 high-precision values (32 * 4 bytes = 128 bytes, taking FP32 as an example), greatly saving the storage of weights and the occupancy of device memory during the inference process.
[0283] Similarly, during the quantization weight loading process, weight dequantization can be performed, or real-time dequantization operations can be carried out during the inference process. The dequantization process is as follows: for a given low-precision tensor (such as the Q8_0 format), every 34 bytes are parsed into a Q8_0 data, where the first two bytes represent the quantization scale, and the last 32 bytes are a group (chunk) of INT8 quantization elements. First, convert each quantization element to a high-precision representation, and then multiply by the scale factor scale to restore the quantization element to a high-precision value (dequantization). The dequantized data can then be used for normal matrix operations.
[0284] During the inference process, there is usually a high requirement for the real-time dequantization speed. To accelerate the dequantization speed, a dequantization operator (micro-operator, compute kernel) can be written in the hardware platform programming language to make full use of the parallel execution ability of the hardware platform and improve the inference speed. The following is a pseudo-code implementation example of the Q8_0 format dequantization micro-operator implemented on the Tencent Cloud GCU platform.
[0285] #define CHUNK_SIZE 32
[0286] typedef struct {
[0287] __fp16 scale;
[0288] int8_t chunk[CHUNK_SIZE];
[0289] } Q8_0;
[0290] template <typename T>
[0291] __device__ void dequantize_q8_0(const void *vx,
[0292] const int idx, T* out) {
[0293] const Q8_0* x = (const Q8_0*) vx;
[0294] const T scale = T(x[idx].scale);
[0295] auto chunk = vcast<T, uint8_t>(x[idx].chunk);
[0296] auto deq = vmul_scalar(chunk, scale);
[0297] vstore(deq, out + idx * CHUNK_SIZE);
[0298] }
[0299] #define DEQUANTIZE_OP(DST_TYPE, FN_NAME) \
[0300] extern "C" __global__ FN_NAME( \
[0301] const void* vx, DST_TYPE* out, \
[0302] const size_t elements) { \
[0303] size_t chunks = elements / CHUNK_SIZE; \
[0304] for (unsigned int i = \
[0305] blockIdx.x * blockDim.x + threadIdx.x; \
[0306] i < chunks; i += blockDim.x * gridDim.x) { \
[0307] dequantize_q8_0<DST_TYPE>(vx, i, out); \
[0308] } \
[0309] } \
[0310] DEQUANTIZE_OP(float, dequantize_q8_0_f32)
[0311] DEQUANTIZE_OP(__fp16, dequantize_q8_0_f16)
[0312] DEQUANTIZE_OP(__bf16, dequantize_q8_0_bf16)
[0313] Among them, first, the data structure Q8_0 is defined, and then the dequantization function dequantize_q8_0 is defined, which is mainly used to dequantize the Q8_0 data of the specified chunk index and store it in the specified position of the output data. The dequantization operators dequantize_q8_0_f32, dequantize_q8_0_fp16, and dequantize_q8_0_bf16 are three dequantization operators exposed to the upper layer to handle inputs of different data types. The dequantization function dequantize_q8_0 is called according to the number of started threads, and the input data vx (each 34 bytes is a Q8_0 data) is dequantized in turn. Here, generics and template functions are used to implement the dequantization of Q8_0 data into FP32, FP16, and BF16 data.
[0314] 8) Backpropagation Module: Mainly used for fine-tuning large models. The implementation of the backpropagation module requires the support of the tensor module, that is, recording the operation type and parameters during tensor operations, and recording the operations and corresponding parameters applied to each tensor within each tensor, so that the gradient can be calculated backward step by step according to the chain rule from the output layer loss during backpropagation of the gradient. The backpropagation module provides a backward calculation interface for tensors. By calling the backward calculation interface of tensors, the chain rule for differentiation is triggered. That is, first, a topological sort is performed on the operation sequence recorded by the tensor, and then the operation sequence is traversed in reverse order. According to the output loss, the backward calculation interface of each sequence element is called. After calculating the gradient, it is passed down successively until the gradient of the specified layer is calculated. Finally, the optimizer is used to adjust the weight parameters of the specified layer to achieve fine-tuning of the model. The specific implementation of the backpropagation module is similar to the standard backward calculation, and the implementation details will not be elaborated here. It should be noted that the backward calculation can also use a combination of micro-operators to reduce the development difficulty of backward operators. For example, for the common activation function SiLU in large models, its backward calculation can use the following combination of micro-operators, which uses single-operand operators and double-operand operators, such as the negative operator, exponential operator, reciprocal operator (recip), multiplication operator, addition operator, subtraction operator, etc.
[0315] fn silu_grad(&x: Tensor) -> Tensor {
[0316] let sigmoid_x = (x.neg().exp() + 1.).recip();
[0317] let grad = &sigmoid_x * (1 + x * (1 - &sigmoid_x));
[0318] return grad;
[0319] }
[0320] 9) Model Layer: As mentioned before, the model layer uses various functional modules provided by the framework layer to build a large model instance based on micro-operators. Among them, each network layer of the large model can be combined using the tensor module, neural network module (Layer module), and custom operator module of the framework layer. The loading of weights for each network layer can be performed using the weight loading and quantization module of the framework layer.
[0321] Taking the mainstream large model network structure LLaMa as an example, its model contains multiple Blocks, and each Block consists of an RmsNorm, an Attention, and an MLP layer. The writing example of its Block is as follows. Among them, the forward function is responsible for the forward calculation of the current block, and its writing method refers to the official example of LLaMa. The load function is responsible for loading the weights of the current block (i.e., the weights of RmsNorm, Attention, and MLP).
[0322] struct Block {
[0323] norm_1: RmsNorm,
[0324] norm_2: RmsNorm,
[0325] attention: Attention,
[0326] mlp: Mlp,
[0327] }
[0328] impl Block {
[0329] fn forward(&mut self, x: &Tensor, pos: usize,
[0330] layer_idx: usize,cache: &mut KVCache) -> Tensor {
[0331] let skip = x;
[0332] let x = self.norm_1.forward(x);
[0333] let x = (self.
[0334] attention.forward(&x, pos, layer_idx, cache) + skip);
[0335] let skip = &x;
[0336] self.mlp.forward(&self.norm_2.forward(&x)) + skip
[0337] }
[0338] fn load(ld: &Loader, w: &str, cfg: &Config) -> Block {
[0339] let ld = ld.sub(&w);
[0340] Block {
[0341] norm_1: RmsNorm::load(&ld, "input_layernorm", cfg),
[0342] norm_2: RmsNorm::load(&ld, "post_attention_layernorm", cfg),
[0343] attention: Attention::load(&ld, "self_attn", cfg),
[0344] mlp: Mlp::load(&ld, "mlp", cfg),
[0345] }
[0346] }
[0347] }
[0348] The RmsNorm module in Block is a simplified LayerNorm and can be written using a custom operator or an aggregated operator; the Attention and MLP modules in Block can also be done in a similar way. For example, the structure definition of the attention layer (Attention) is as follows, which contains head, kv_heads, and head_dim used to record the multi-head attention layer, and four Linear layers for projecting input and output data:
[0349] struct Attention {
[0350] heads: usize,
[0351] kv_heads: usize,
[0352] head_dim: usize,
[0353] q_linear: Linear,
[0354] k_linear: Linear,
[0355] v_linear: Linear,
[0356] o_linear: Linear,
[0357] }
[0358] The forward calculation of the attention layer operator is as follows. It includes: First, perform a linear layer projection of the input tensor into query vector (q), key vector (k), and value vector (v). Then, reshape the projected q, k, and v tensors to adapt to subsequent multi-headed attention calculations. After that, perform rotary position encoding (rope) and key-value cache (kv cache) caching, and calculate the attention values using scaled dot-product attention (a combination of two matmuls and softmax). Finally, reshape the output and perform a linear layer projection of the output tensor to obtain the final output of the attention layer operator.
[0359] impl Forward for Attention {
[0360] fn forward(&self, x: &Tensor, index: usize,
[0361] layer_idx: usize, cache: &mut Cache) -> Tensor {
[0362] let (batch, seq_len, hidden_size) = x.dims();
[0363] let q = self.q_linear.forward(x);
[0364] let k = self.k_linear.forward(x);
[0365] let v = self.v_linear.forward(x);
[0366] let q = q.reshape((batch, seq_len, self.heads, self.dim))
[0367] .transpose(1, 2);
[0368] let k = k.reshape((batch, seq_len, self.kv_heads, self.dim))
[0369] .transpose(1, 2);
[0370] let v = v.reshape((batch, seq_len, self.kv_heads, self.dim))
[0371] .transpose(1, 2);
[0372] let (q, mut k) = nn::rope(&q, &k, &cache, index);
[0373] if let Some((k_cache, v_cache)) = &cache.kvcache[layer_idx] {
[0374] k = nn::kvconcat(k_cache, &k, 2);
[0375] v = nn::kvconcat(v_cache, &v, 2);
[0376] }
[0377] cache.kvcache[layer_idx] = Some((k.clone(), v.clone()));
[0378] let k = nn::repeat(k);
[0379] let v = nn::repeat(v);
[0380] let attn = (q.matmul(&k.transpose()) / self.dim.sqrt());
[0381] let attn = match &cache.mask {
[0382] Some(mask) => {
[0383] nn::softmax(attn + mask)
[0384] }
[0385] _ => nn::softmax(attn)
[0386] };
[0387] let out = attn.matmul(&v).transpose(1, 2);
[0388] let out = out.reshape(&[batch, seq_len, hidden_size]);
[0389] self.o_linear.forward(&out)
[0390] }
[0391] }
[0392] The weight loading of the attention layer operator is similar to that of the block layer. It uses a weight loading module to load the corresponding weights for the input and output projection layers as follows:
[0393] impl Attention {
[0394] fn load(ld: Loader, w: &str, c: &Config) -> Attention {
[0395] let ld = ld.sub(&w);
[0396] let h_sz = c.hidden_size;
[0397] let i_sz = c.inner_size;
[0398] let q_sz = (h_sz / c.heads) * c.heads;
[0399] let kv_sz = (h_sz / c.heads) * c.kv_heads;
[0400] Attention {
[0401] heads: c.heads,
[0402] kv_heads: c.kv_heads,
[0403] head_dim: c.hidden_size / c.kv_heads,
[0404] q_linear: Linear::load(i_sz, q_sz, &ld, "q_proj"),
[0405] k_linear: Linear::load(i_sz, kv_sz, &ld, "k_proj"),
[0406] v_linear: Linear::load(i_sz, kv_sz, &ld, "v_proj"),
[0407] o_linear: Linear::load(q_sz, i_sz, &ld, "o_proj"),
[0408] }
[0409] }
[0410] }
[0411] After completing the writing of the attention layer operator, you can continue to write the multi-layer perceptron layer (MLP). As shown below, the load function is responsible for loading the weights fc_1 and fc_2 required for the input linear projection, as well as the output projection weight fc_o. The forward function first performs a linear projection of the input using fc_1, then performs a non-linear activation (silu activation) on the projected tensor, multiplies it with the result of the tensor obtained by the linear projection of fc_2, and finally performs an output projection to obtain the final output of the multi-layer perceptron layer.
[0412] struct Mlp {
[0413] fc_1: Linear,
[0414] fc_2: Linear,
[0415] fc_o: Linear,
[0416] }
[0417] impl Mlp {
[0418] fn forward(&self, x: &Tensor) -> Tensor {
[0419] let x1 = self.fc_1.forward(x);
[0420] let x = nn::silu(&x1) * self.fc_2.forward(x);
[0421] self.fc_o.forward(&x)
[0422] }
[0423] fn load(ld: Loader, w: &str, cfg: &Config) -> Mlp {
[0424] let ld = ld.sub(&w);
[0425] let h_sz = cfg.hidden_size;
[0426] let i_sz = cfg.inner_size;
[0427] Mlp {
[0428] fc_1: Linear::load(h_sz, i_sz, &ld, "gate_proj"),
[0429] fc_2: Linear::load(h_sz, i_sz, &ld, "up_proj"),
[0430] fc_o: Linear::load(i_sz, h_sz, &ld, "down_proj"),
[0431] }
[0432] }
[0433] }
[0434] After writing the standardized operator, attention layer operator, multi-layer perceptron layer, and other modules, the mainstream large model LLaMa can be written using the model layer specification shown in the present invention, as exemplified below. In the forward function of LLaMa, first, the input tensor is encoded using the word vector encoding module. Then, according to the number of model layers of LLaMa (different numbers of layers correspond to different-sized LLaMa models, such as 7B, 13B, 70B), the forward functions of each module are called in sequence for forward calculation. After that, normalization (RmsNorm normalization) is performed on the module output, and the slice operator is used to obtain the output results of the corresponding slices. Finally, the final prediction result of LLaMa is obtained using the output layer linear projection. The LLaMa output also needs to go through word vector encoding and output collection to obtain the text result, also known as a token. Correspondingly, the load function of LLaMa is responsible for loading the word vector encoding weights, the weights of each module layer of the model, the weights of the standardized operator, and the weights of the output projection layer.
[0435] pub struct LlamaModel {
[0436] emb: Embedding,
[0437] blocks: Vec <block>,
[0438] norm: RmsNorm,
[0439] head: Linear,
[0440] }
[0441] impl Forward for LlamaModel {
[0442] pub fn forward(&self, x: &Tensor, index: usize,
[0443] cache: &mut Cache) -> Tensor {
[0444] let (batch, seq_len) = x.dims();
[0445] let mut x = self.emb.forward(x);
[0446] for (idx, block) in self.blocks.iter().enumerate() {
[0447] x = block.forward(&x, index, idx, cache)?;
[0448] }
[0449] let x = self.norm.forward(&x);
[0450] let x = x.slice(seq_len - 1, 1);
[0451] self.head.forward(&x)
[0452] }
[0453] }
[0454] impl LlamaModel {
[0455] pub fn load(ld: Loader, w: &str, cfg: &Config) -> LlamaModel {
[0456] let ld = ld.sub(w);
[0457] let emb = Embedding::load(cfg.vocab_size, cfg.hidden_size,
[0458] &ld, "model.embed_tokens");
[0459] let blocks: Vec<_> = (0..cfg.num_layers)
[0460] .map(|i| Block::load(&ld, &format!("model.layers.{i}"), &cfg))
[0461] .collect();
[0462] let norm = RmsNorm::load(cfg.hidden_size, cfg.eps,
[0463] &ld, "model.norm");
[0464] let head = Linear::load(cfg.hidden_size, cfg.vocab_size,
[0465] &ld, "lm_head");
[0466] LlamaModel { emb, blocks, norm, head}
[0467] }
[0468] }
[0469] For the fine-tuning of large models, generally, a small number of linear layers (low-rank weights) need to be added to the pre-trained model. The reverse computation micro-operator or micro-operator combination of the corresponding linear layer can be written according to the backpropagation module in the framework layer, and the backward function is implemented for the large model (such as LlamaModel). Among them, in the backward function, the gradients of the newly added linear layer are calculated according to the input loss, and the gradient tensor is returned. The returned gradient tensor is called by the backpropagation module in the framework layer to update the weights of this linear layer by the optimizer to achieve the purpose of fine-tuning the model weights. The writing method of the backward function for model fine-tuning is similar to that of the forward computation function.
[0470] An embodiment of the present invention provides a large model cross-platform system based on micro-operators, which can be used for quickly performing inference and fine-tuning of large models under different hardware platforms, and is applicable to quickly developing, debugging, inferring, and fine-tuning large models under different hardware platforms. It simplifies the development difficulty of large model operators, has a simple system structure, and is conducive to model development, debugging, and cross-platform deployment. It defines the hardware access method, the implementation form and call method of micro-operators, and each functional module required for realizing large model inference and fine-tuning based on micro-operators. Its overall system structure design is simple, which can effectively reduce the operator development difficulty and improve the large model development and deployment efficiency.
[0471] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0472] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.< / block> < / tensor> < / tensor> < / tensor> < / t> < / t> < / t> < / t> < / t>
Claims
1. A large model cross-platform system based on micro-operators, characterized in that Including: Hardware execution layer, hardware access layer, pre-compilation layer, micro-operator layer, backend layer, framework layer, and model layer; Among them, the hardware execution layer is used to execute memory access instructions and / or various micro-operators compiled by the pre-compilation layer. A micro-operator is an intermediate instruction for calculation or operation or an executable instruction of a hardware device. Various micro-operators include micro-operators for device memory access management operations and micro-operators for tensor calculation; The hardware access layer is used to form call interfaces for calling various hardware functions of each hardware device for accessing and integrating into the large model cross-platform system; The pre-compilation layer is used to call the platform compiler to compile the micro-operator code into intermediate instructions or executable instructions of the hardware device; The micro-operator layer is used to call various micro-operators compiled by the pre-compilation layer and provide a programming interface based on micro-operators to the upper layer, so that the upper layer can combine and call multiple micro-operators to implement various operators required by mainstream large models; The backend layer interfaces with the framework layer upward and the micro-operator layer downward, is used to expose the function interfaces of the micro-operator layer outward, and at the same time the backend layer identifies the tensor data type and realizes it through different data types of the same micro-operator compiled by the pre-compilation layer called by the micro-operator layer; The framework layer contains various function modules required for realizing large model inference and fine-tuning, including a tensor module, a storage module, a device module, a weight loading module, a custom operator module, a neural network module, a weight quantization module, a word vector encoding module, a backpropagation module, and an output sampling module. Each function module realizes the corresponding large model inference function or large model fine-tuning function by calling the function interfaces in the backend layer; The model layer is used to provide a model layer specification for the user to build a large model instance based on the framework layer; among them, providing a model layer specification for the user to build a large model instance based on the framework layer includes: using the tensor module, neural network module, and custom operator module provided by the framework layer to build a micro-operator-based large model instance according to the mainstream large model structure; using the weight loading module provided by the framework layer to load the large model weights and calling the weight quantization module provided by the framework layer for quantization and de-quantization of the large model weights according to the model requirements; calling the function modules in the framework layer for large model inference and fine-tuning.
2. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that When the large model cross-platform system is compiled into a binary program, the pre-compilation layer calls the platform compiler to compile the pre-written micro-operator code into intermediate instructions or executable instructions of the hardware device; The pre-compilation layer is also used to provide a cache function for pre-compilation instructions.
3. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that The micro-operator layer contains device memory access management micro-operators, tensor calculation micro-operators, and various operators required by mainstream large models realized by combining tensor calculation micro-operators; The device memory access management micro-operators include device memory allocation micro-operators, device memory release micro-operators, host-to-device communication micro-operators, device-to-host communication micro-operators, device-to-device communication micro-operators, device synchronization micro-operators, and device query micro-operators; The tensor calculation micro-operator includes a single operand operator, a double operand operator, a type conversion operator, an index operator, a general copy operator, a conditional selection operator, an aggregation operator, and a matrix multiplication operator; The tensor calculation micro-operators are combined to implement various operators required by mainstream large models. The various operators required by mainstream large models include rotary position encoding operators, activation layer operators, normalization operators, standardization operators, and attention layer operators.
4. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that, The backend layer parses the interface call parameters of the framework layer and identifies the parameter types, loads the micro-operators compiled by the pre-compilation layer in the form of computing core startup, starts execution through the hardware execution layer, and encapsulates and returns the calculation results.
5. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that, The framework layer includes a tensor module; the tensor module is an abstract representation of tensors in a neural network. The tensor module contains tensor data and various common tensor operations, and various tensor operations are implemented by calling the functional interfaces of the corresponding micro-operators in the backend layer.
6. The cross-platform system of the large model based on micro-operators according to claim 1, wherein The framework layer includes a storage module; the storage module is a lower-level representation of tensors and tensor operations, used to identify the device category and data type of tensors. The operations of identifying the device category and data type of tensors are implemented by calling the functional interfaces of the corresponding micro-operators in the backend layer.
7. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that, The framework layer includes a device module; the device module is used to provide functional interfaces required for memory access operations and micro-operator execution operations on different hardware devices. The memory access operations include operations required for tensor allocation, recycling, and inter-device migration on different devices.
8. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that, The framework layer includes a weight loading module; the weight loading module is used to load the large model weights from a file. Among them, the weight loading process is triggered by the model layer. The weight loading module sets the target data type of the large model weights when creating a loader, and performs data type conversion on the large model weights during the weight loading process, converting the data type of the large model weights to the target data type.
9. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that The framework layer includes a custom operator module; the custom operator module is used to provide functional interfaces for extending micro-operators. Among them, the functional interfaces of the extended micro-operators include extension interfaces for single operand operators, double operand operators, and multi-operand operators. Non-standard micro-operators or aggregation operators are implemented through the functional interfaces of the extended micro-operators.
10. The large model cross-platform system based on micro-operators according to claim 1, characterized in that The framework layer includes a neural network module; the neural network module defines common large model operators. The implementation forms of common large model operators are written using micro-operators or combinations of micro-operators, or written by calling aggregation operators through the custom operator module.
11. The cross-platform system for large models based on micro-operators according to claim 1, characterized in that, The framework layer includes a weight quantization module; the weight quantization module is an extension of the tensor module. In addition to providing the standard operations defined by the tensor module, it also provides quantization and de-quantization functional interfaces to support low-precision large model inference and fine-tuning.
12. The large model cross-platform system based on micro-operators according to claim 1, wherein The framework layer includes a word vector encoding module; the word vector encoding module is used to implement the word vector encoding operations required for large model inference or large model fine-tuning by calling the micro-operators in the micro-operator layer.
13. The cross-platform system of the large model based on micro-operators according to claim 1, characterized in that, The framework layer includes a backpropagation module; the backpropagation module is used to implement fine-tuning of the large model. The backpropagation module depends on the tensor module to record the operation types and parameters during the tensor operation process, and calculates the backward gradient through the backward calculation interface of the tensor module. In the backpropagation module, backward gradient calculation operators corresponding to various tensor operations are defined. The backward gradient calculation operators corresponding to various tensor operations are implemented by combining various micro-operators or are written and implemented by calling aggregation operators through the custom operator module.
Citation Information
Patent Citations
Tensor calculation unit-oriented convolution operator optimization implementation method
CN115983356A
Operator fusion method, system, equipment and medium
CN119272234A