Large language model performance optimization method and device, computer equipment, readable storage medium and program product

By obtaining and compiling the language data parameters and processing parameters of the large language model, generating calculation diagrams and updating operators, the problem of large language model not adapting to the chip at the operator level is solved, and the operation efficiency and performance of the model are improved.

CN119990265APending Publication Date: 2025-05-13BEIJING QINGCHENG JIZHI TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510209806.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art lacks abstraction of adapting to chips at the operator level of large language models, resulting in low task execution efficiency and low model performance when performing language translation and other tasks after deployment by users.

Method used

By obtaining the language data parameters and processing parameters of the large language model, the objective function is compiled to generate a calculation graph, and multiple object operators are generated based on the calculation graph and the parameters of the hardware to be deployed, and the large language model is updated to optimize its performance.

Benefits of technology

Improve the operation efficiency and portability of large language models on hardware to be deployed, and enhance the performance and user experience of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990265A_ABST
    Figure CN119990265A_ABST
Patent Text Reader

Abstract

The invention relates to a large language model performance optimization method and device, computer equipment, a readable storage medium and a program product, and belongs to the technical field of deep learning model compilers. The method comprises the steps of obtaining data parameters and processing parameters of language data corresponding to a large language model; the processing parameters are determined according to parameters of to-be-deployed hardware corresponding to the large language model; in response to a programming operation aiming at a preset programming primitive, obtaining a target function representing arithmetic logic of the large language model; according to the data parameters and the processing parameters, compiling the target function to obtain a calculation graph corresponding to the target function; and according to the computational graph and the parameters of the to-be-deployed hardware, generating a plurality of target operators, and updating the large language model based on the plurality of target operators to obtain an optimized large language model. By adopting the method, the performance of the large language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of deep learning model compilers, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for optimizing the performance of a large language model. Background Art

[0002] In deep learning and large model reasoning, updating specific operators in the model to improve the model's reasoning performance is a common acceleration method. However, traditional update methods often require developers to have an in-depth understanding of the chip architecture and the corresponding programming interface, which increases the difficulty of development and the promotion of domestic chips. Therefore, related technologies will hide the details of the chip through compilation to provide developers with a simple development method. However, although the related technology provides an overall update method at the model level, it does not build an abstraction that is compatible with the chip at the operator level. After the user deploys a large language model, when performing language translation processing tasks, there will be a problem of low task execution efficiency. Therefore, the performance optimization method of the large language model provided in the related technology has the problem of low model performance. Summary of the invention

[0003] Based on this, it is necessary to provide a large language model performance optimization method, device, computer equipment, computer-readable storage medium and computer program product that can improve the performance of a large language model in response to the above technical problems.

[0004] In a first aspect, the present application provides a method for optimizing the performance of a large language model, comprising:

[0005] Acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model;

[0006] In response to a programming operation for a preset programming primitive, an objective function representing the operation logic of the large language model is obtained;

[0007] Compiling the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function;

[0008] According to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model.

[0009] In one embodiment, compiling the objective function according to the data parameter and the processing parameter to obtain a calculation graph corresponding to the objective function includes:

[0010] Compiling the target function according to the data parameters and the processing parameters to obtain intermediate layer parameters of the target function;

[0011] According to the intermediate layer parameters, a computation graph corresponding to the objective function is constructed.

[0012] In one embodiment, compiling the target function according to the data parameter and the processing parameter to obtain the intermediate layer parameters of the target function includes:

[0013] Determining a plurality of target programming primitives involved in the target function;

[0014] Compiling the target function according to the data parameters and the processing parameters to obtain input and output types and tensor shapes of each target programming primitive;

[0015] The input and output types of each of the target programming primitives and the tensor shape are used as the intermediate layer parameters.

[0016] In one embodiment, constructing a computation graph according to the intermediate layer parameters includes:

[0017] constructing a plurality of nodes and edges according to the tensor shape and the input and output types of each of the target programming primitives;

[0018] The computation graph is generated according to the plurality of nodes and the edges.

[0019] In one embodiment, generating multiple target operators according to the computation graph and the parameters of the hardware to be deployed, and updating the large language model based on the multiple target operators includes:

[0020] Performing graph optimization processing on the computation graph to obtain an optimized computation graph;

[0021] According to the optimized computation graph and the parameters of the hardware to be deployed, multiple target operators are generated to update the large language model.

[0022] In one of the embodiments, the programming primitive includes a calculation part and a memory access part, and the input and output of each of the programming primitives include data size, data quantity and data storage location.

[0023] In a second aspect, the present application also provides a large language model performance optimization device, comprising:

[0024] A parameter acquisition module, used to acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model;

[0025] A function programming module, for obtaining a target function representing the operation logic of the large language model in response to a programming operation on a preset programming primitive;

[0026] A function compiling module, used to compile the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function;

[0027] A model updating module is used to generate multiple target operators according to the calculation graph and the parameters of the hardware to be deployed, and update the large language model based on the multiple target operators to obtain an optimized large language model.

[0028] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0029] Acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model;

[0030] In response to a programming operation for a preset programming primitive, an objective function representing the operation logic of the large language model is obtained;

[0031] Compiling the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function;

[0032] According to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model.

[0033] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0034] Acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model;

[0035] In response to a programming operation for a preset programming primitive, an objective function representing the operation logic of the large language model is obtained;

[0036] Compiling the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function;

[0037] According to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model.

[0038] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0039] Acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model;

[0040] In response to a programming operation for a preset programming primitive, an objective function representing the operation logic of the large language model is obtained;

[0041] Compiling the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function;

[0042] According to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model.

[0043] The above-mentioned large language model performance optimization method, device, computer equipment, computer-readable storage medium and computer program product obtain data parameters and processing parameters of language data corresponding to the large language model; wherein the processing parameters are determined according to the parameters of the hardware to be deployed corresponding to the large language model, so that the large language model can be more closely matched with the hardware to be deployed when deployed; in response to programming operations for preset programming primitives, a target function representing the operation logic of the large language model is obtained for subsequent compilation operations; further, according to the data parameters and processing parameters, the target function is compiled to obtain a calculation graph corresponding to the target function, the calculation graph is traceable, which helps to analyze the performance of the large language model and perform troubleshooting, facilitates the model to run on the hardware to be deployed, and improves the portability of the model; and according to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model, and the target operator is a series of high-performance operators, thereby improving the performance of the large language model after update. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0045] Figure 1 A schematic diagram of a flow chart of a method for optimizing performance of a large language model in one embodiment;

[0046] Figure 2 A schematic diagram of a flow chart of the steps of constructing a computational graph in one embodiment;

[0047] Figure 3 A schematic diagram of a flow chart of a computation graph construction step in another embodiment;

[0048] Figure 4 It is a structural block diagram of a large language model performance optimization device in one embodiment;

[0049] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] As described in the background technology, the large language model performance optimization method in the prior art has the problem of low model performance after updating. The inventor has found that the reason for this problem is that in deep learning and large model reasoning, optimizing specific operators in the model to improve the reasoning performance of the model is a common acceleration method. However, traditional optimization methods often require developers to have an in-depth understanding of the chip architecture and the corresponding programming interface, which greatly increases the difficulty of development and the promotion of domestic chips. Therefore, by compiling, hiding chip details, and providing developers with a concise development method, the above problems can be effectively solved. In the related technology, triton is a high-performance programming language for operators, and tvm is an end-to-end deep learning model compilation system. For triton technology, based on the traditional compilation system, there is no way to abstract the calculation graph well, and there is no way to be compatible with the programming interface provided by the current artificial intelligence chip, resulting in unnecessary analysis and processing during the generation of operators, increasing compilation time, and the performance of the generated operators may not necessarily reach the best. For tvm: although an overall optimization method is provided at the model level, no suitable abstraction is constructed at the operator level, making it very difficult to improve the performance of specific operators. Similar technologies are often built in a static way, requiring the size of the data to be built and processed, memory layout on the chip, explicit type construction, etc., with a long development cycle and high development difficulty. Therefore, after the large language model updated by the above two technical means is deployed on the hardware, when executing language translation tasks, due to the mismatch between the model operator and the underlying logic of the hardware, the execution delay of tasks such as customer service robots is too high, the waiting time is too long, affecting the user experience, and there is a problem of low performance of the updated large language model.

[0052] Based on the above reasons, the present invention provides an update solution for a large language model, aiming to solve the problem of low performance of the updated large language model.

[0053] In one embodiment, Figure 1 As shown, a method for optimizing the performance of a large language model is provided. This embodiment uses the method to be applied to a deep learning model compilation system of a server as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0054] Step S102: obtaining data parameters and processing parameters of language data corresponding to the large language model.

[0055] The large language model may be a BERT model (Bidirectional Encoder Representations from Transformers), which is an open source machine learning framework designed for the field of natural language processing. The language data may be the data to be processed by the large language model, and the data to be translated that is input into the large language model. The data parameters may include parameters such as the type of language data and the size of the data, and the processing parameters may be a segmentation method for the language data, wherein the data segmentation method may refer to distributing the data stored in the same database to multiple databases (hosts) under certain specific conditions to achieve the effect of distributing the load of a single device. Data segmentation can be divided into two segmentation modes according to the type of segmentation rules. One is to segment data to different databases (hosts) according to different tables, which can be called vertical (vertical) segmentation of data; the other is to split the data in the same table to multiple databases (hosts) according to certain conditions based on the logical relationship of the data in the table, which is called horizontal (horizontal) segmentation of data.

[0056] The processing parameters are determined according to the parameters of the hardware to be deployed corresponding to the large language model, wherein the hardware to be deployed may be a chip that needs to run the large language model. It can be understood that the processing parameters are the segmentation method of the language data, and the segmentation rules of the atmosphere method are determined by factors such as the memory hierarchy of the hardware to be deployed and the access mode of the hardware, so that the large language model can fit the hardware to be deployed, thereby improving the operating efficiency.

[0057] Optionally, the system obtains data parameters and processing parameters of the language data to be processed corresponding to the large language model, which serve as a data basis for subsequent operation derivation of the large language model.

[0058] Step S104 , in response to the programming operation for the preset programming primitive, obtain the target function representing the operation logic of the large language model.

[0059] Among them, programming primitives can refer to the most basic building blocks or operations in a programming language, which are used to build more complex program logic. They are usually the underlying functions provided by the language, and programmers can directly use these primitives to perform basic calculations or control operations. Among them, programming operations can refer to specific tasks or instructions performed during the programming process.

[0060] The operation logic may be the mechanism by which a model processes input data and generates output in machine learning or deep learning.

[0061] Optionally, in response to programming operations for a plurality of preset programming primitives, the system constructs an objective function representing the operational logic of the large language model using the preset programming primitives according to instructions of the programming operations.

[0062] Step S106, compile the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function.

[0063] Among them, the compilation operation can be a process of converting a high-level programming language (such as C++, Java) into a low-level language (such as machine language or bytecode).

[0064] Among them, the computational graph can be a graphical structure used to represent computing operations. In machine learning, nodes represent operations (such as addition and multiplication) and edges represent data flows (tensor flows).

[0065] Optionally, the system compiles the objective function based on the data parameters and processing parameters of the language data input into the large language model, uses implicit deduction technology to derive the parameters involved in the operation process of the code of the large language model at runtime, and converts them into a computational graph form, thereby obtaining a computational graph corresponding to the objective function.

[0066] Step S108, generating multiple target operators according to the calculation graph and the parameters of the hardware to be deployed, and updating the large language model based on the multiple target operators to obtain an optimized large language model.

[0067] Among them, the operator is a node in the computational graph, representing a specific mathematical operation.

[0068] Optionally, the system generates multiple target operators based on the computation graph and the parameters of the hardware to be deployed, updates the code of the large language model based on the multiple target operators, and obtains an updated large language model. It should be noted that the parameters of the hardware to be deployed are obtained by calling the application program interface of the hardware to be deployed.

[0069] In the above-mentioned large language model performance optimization method, data parameters and processing parameters of language data corresponding to the large language model are obtained; wherein the processing parameters are determined according to the parameters of the hardware to be deployed corresponding to the large language model, so that the large language model can be more closely matched with the hardware to be deployed when deployed; in response to programming operations for preset programming primitives, a target function representing the operation logic of the large language model is obtained for subsequent compilation operations; further, according to the data parameters and processing parameters, the target function is compiled to obtain a calculation graph corresponding to the target function, and the calculation graph is traceable, which is helpful to analyze the performance of the large language model and perform troubleshooting, so as to facilitate the model to run on the hardware to be deployed, thereby improving the portability of the model; and according to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model, and the target operator is a series of high-performance operators, thereby improving the performance of the large language model after update.

[0070] In an exemplary embodiment, Figure 2 As shown, step S106 compiles the target function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the target function, including the following steps S202 and S204. Among them:

[0071] Step S202, compile the target function according to the data parameters and the processing parameters to obtain the intermediate layer parameters of the target function.

[0072] The intermediate layer parameters may be data parameters generated by the intermediate layer operation during the operation of the large language model, including the type of output data and the size of the data.

[0073] Optionally, the system compiles the target function according to the data parameters and processing parameters of the language data input into the large language model to derive the intermediate layer parameters of the target function. For example, the operation logic represented by the target function is multiplication, and the data parameters of the language data include that the data type is a floating point number and the size of the data is two. Then, by compiling the multiplication of two floating point numbers, it can be deduced that the output data must be floating point data, and the output type obtained is also a floating point number.

[0074] Step S204, constructing a computation graph corresponding to the objective function according to the intermediate layer parameters.

[0075] Optionally, the system determines the data flow, data type, and operation type of the target function during the compilation process based on the derived intermediate layer parameters, thereby constructing a computational graph corresponding to the target function.

[0076] In this embodiment, by compiling the objective function and deriving the intermediate layer parameters, the system can optimize specific data flows and operation types, improve the calculation speed, and thus further improve the performance of the model. Through the visualized calculation graph, researchers and developers can more intuitively understand the working principle of the model, which helps to analyze the performance of large language models and perform troubleshooting, making it easier for the model to run on the hardware to be deployed, thereby improving the portability of the model and eliminating the need for developers to understand the specific architecture of the hardware to be deployed.

[0077] In an exemplary embodiment, Figure 3 As shown, step S202 compiles the target function according to the data parameters and the processing parameters to obtain the intermediate layer parameters of the target function, including the following steps S302 to S306. Among them:

[0078] Step S302, determining multiple target programming primitives involved in the target function.

[0079] The programming primitive includes a calculation part and a memory access part, and the input and output of each programming primitive include data size, data quantity and data storage location.

[0080] Optionally, the system determines preset programming primitives involved in programming the target function, and uses multiple preset programming primitives involved in the target function as target programming primitives.

[0081] Step S304, compile the target function according to the data parameters and the processing parameters to obtain the input and output types and tensor shapes of each target programming primitive.

[0082] The input and output types may refer to the input data type and output data type of each target programming primitive.

[0083] Among them, a tensor can be a mathematical object, which can be regarded as a generalization of a scalar (0th-order tensor), a vector (1st-order tensor) and a matrix (2nd-order tensor). A tensor can represent data and relationships in multidimensional space and is often used in physics, engineering, and deep learning. Among them, the shape of a tensor can refer to the size of each dimension of the tensor.

[0084] Optionally, the system compiles the target function according to the data parameters and processing parameters to obtain the input and output types of each target programming primitive and the tensor shape. In deep learning, data structures and algorithms can be identified and implemented by tensors, which simplifies the model design, training and reasoning process.

[0085] Step S306, taking the input and output types and tensor shapes of each target programming primitive as intermediate layer parameters.

[0086] Optionally, the system uses the input and output types and tensor shapes of each target programming primitive as intermediate layer parameters and as the basis for subsequent computational graph transformations.

[0087] In this embodiment, by compiling the target function through the deep learning model compilation system, a more efficient calculation process can be generated to avoid unnecessary runtime overhead. During the compilation process, the system determines the input and output types and tensor shapes, which helps to find potential errors before execution. Using input and output types and tensor shapes as intermediate layer parameters can support more complex operations and calculation graphs, making it easier to integrate new algorithms.

[0088] In an exemplary embodiment, step S204 constructs a computation graph according to the intermediate layer parameters, including:

[0089] According to the tensor shape and the input and output types of each target programming primitive, multiple nodes and edges are constructed; and a computational graph is generated according to the multiple nodes and edges.

[0090] Among them, nodes represent operations or data in the computational graph, and edges connect nodes, indicating the data flow or dependency relationship between nodes.

[0091] Optionally, the system constructs multiple nodes and edges according to the tensor shape and the input and output types of each target programming primitive. Specifically, multiple nodes are constructed through the data corresponding to the input and output types, the tensor shape, and the operations involved in the target programming primitives. According to the operation logic of the target function and the input and output types of the target programming primitives, the association between the target programming primitives is determined to determine the data flow direction and data dependencies, thereby constructing multiple edges, integrating multiple nodes and edges, and generating a computational graph.

[0092] In this embodiment, constructing a computational graph enables computing resources to be managed at the graph level. For example, memory and computing resources can be allocated separately for nodes of different operations. Effective allocation of resources can improve overall execution efficiency. When processing large language models, computing time and memory usage can be significantly reduced, thereby improving the performance of large language models.

[0093] In an exemplary embodiment, step S108 generates multiple target operators according to the computation graph and the parameters of the hardware to be deployed, and updates the large language model based on the multiple target operators, including:

[0094] The computation graph is optimized to obtain an optimized computation graph. According to the optimized computation graph and the parameters of the hardware to be deployed, multiple target operators are generated to update the large language model.

[0095] Among them, graph optimization processing can be an optimization of the computational graph structure in deep learning and machine learning to improve the training and reasoning efficiency of the model.

[0096] Optionally, the system performs graph optimization processing on the computational graph, such as operation fusion, weight sharing, pruning, quantization, and dynamic computational graphs, to reduce the computational workload and storage requirements of the computational graph, and generates multiple high-performance target operators based on the optimized computational graph and the parameters of the hardware to be deployed, which are used to update the operator code of the large language model. Among them, operation fusion is to merge multiple operations into one operation to reduce the overhead of computing and memory transmission; weight sharing is to share some weights among multiple operations to reduce the storage requirements of the model; pruning is to remove redundant neurons or connections to reduce the computational workload and storage requirements; quantization is to convert floating-point weights into low-precision representations to reduce the model size and accelerate reasoning; dynamic computational graphs are to dynamically adjust the computational graphs according to actual input data to save unnecessary calculations.

[0097] In this embodiment, the system optimizes the computational graph and generates high-performance operators to update the large language model, which can effectively improve the performance of the large language model, reduce the amount of computation and storage requirements, and enhance the deployment capability of the model on the hardware to be deployed. It not only improves the reasoning efficiency of the model and reduces energy consumption, but also expands the application scenarios of the model, enabling it to play a role in a wider range of practical applications.

[0098] In an exemplary embodiment, another performance optimization method for a large language model is provided, which is applied to a deep learning compilation system, and specifically includes:

[0099] Step 1: Provide a set of preset programming primitives to express the required calculation and memory access operators in constructing operators. Users need to implement a function based on the preset compilation primitives. The function entry provides the data type to be processed, the data size to be processed, and the data segmentation method. Then use the provided programming primitives to describe the relevant logic (the operation logic of the target function) to complete the writing of a function (target function).

[0100] Among them, the deep learning model compilation system is a system that uses a mixture of C++ and python programming languages. Since mainstream deep learning frameworks such as pytorch are based on python, a python interface is provided to facilitate integration into the deep learning framework for easy calling. At the same time, the development, compilation, and adaptation of high-performance operators to the underlying chip need to be completed using c++. Therefore, the present invention uses a mixture of python and c++ to achieve both ease of use and high performance.

[0101] The programming primitives are divided into the calculation part, the unary operator, the binary operator, and the memory access part, including sending from global memory to shared memory and local memory. The input and output of each programming primitive include the size of the data, the amount of data to be processed, and the storage location of the data. These basically do not need to be specified by the user, but are implicitly deduced by the compiler.

[0102] Step 2: Compile the function provided by the user, and use type inference technology and tensor shape inference technology to derive the specific type (input and output type) of each programming primitive (target programming primitive) in the implementation function and the shape of the tensor to be processed.

[0103] Among them, the input of the function includes the type of data. After the relevant type reaches a specific execution statement, the compiler knows the input type and input size of the statement. Through the calculation logic of the corresponding statement, for example, if two floating-point numbers are multiplied, the result must be a floating-point number, and the output type is also a floating-point number. To obtain the output type and data size of the statement. In this way, the input and output types and tensor sizes of each statement in the function are deduced, and the static calculation graph of this function is constructed.

[0104] Step 3: After obtaining the function with complete type annotations and tensor shapes, the compilation system converts it into a computational graph and generates a high-performance implementation for a specific chip for each element of the computational graph, thereby achieving high-performance operator code (target operator).

[0105] Among them, after obtaining the static calculation graph, the compiler can make full use of the type information and tensor information of each operator in the graph to generate high-performance code. One business scenario is that after the user deploys the BERT model, the delay is too high and the waiting time is too long when performing tasks such as language translation and customer service robots, which seriously affects the user experience. For example, when performing translation, waiting for a word for 10 milliseconds will cause the user to wait for a long time during translation, which seriously affects the use of the translation function. By applying the compiler mentioned in the present invention, the core operators in the model can be optimized, the performance of the actual deployment of the model can be improved, and the execution time of related tasks can be increased to a few tenths of a millisecond to ensure the availability of related tasks.

[0106] In this embodiment, the problem that the development difficulty of building specific artificial intelligence operators in the prior art is too high and that it is necessary to take into account the characteristics of the implementation logic and the underlying hardware, effectively implement high-performance operators, and enable developers to obtain high-performance operator implementations without understanding the underlying hardware design.

[0107] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0108] Based on the same inventive concept, the embodiment of the present application also provides a large language model performance optimization device for implementing the large language model performance optimization method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more large language model performance optimization device embodiments provided below can refer to the limitations of the large language model performance optimization method above, and will not be repeated here.

[0109] In an exemplary embodiment, Figure 4 As shown, a large language model performance optimization device 400 is provided, including: a parameter acquisition module 402, a function programming module 404, a function compilation module 406 and a model update module 408, wherein:

[0110] The parameter acquisition module 402 is used to acquire data parameters and processing parameters of the language data corresponding to the large language model; the processing parameters are determined according to the parameters of the hardware to be deployed corresponding to the large language model.

[0111] Functional programming module 404 is used to obtain a target function representing the operation logic of the large language model in response to programming operations for preset programming primitives.

[0112] The function compilation module 406 is used to compile the target function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the target function.

[0113] The model updating module 408 is used to generate multiple target operators according to the calculation graph and the parameters of the hardware to be deployed, and update the large language model based on the multiple target operators to obtain an optimized large language model.

[0114] Furthermore, in one embodiment, the function compilation module 406 is also used to compile the target function according to the data parameters and the processing parameters to obtain the intermediate layer parameters of the target function; and to construct a calculation graph corresponding to the target function according to the intermediate layer parameters.

[0115] Furthermore, in one embodiment, the function compilation module 406 is also used to determine multiple target programming primitives involved in the target function; compile the target function according to the data parameters and processing parameters to obtain the input and output types and tensor shapes of each target programming primitive; and use the input and output types and tensor shapes of each target programming primitive as intermediate layer parameters.

[0116] Furthermore, in one embodiment, the function compilation module 406 is also used to construct multiple nodes and edges according to the tensor shape and the input and output types of each target programming primitive; and generate a computational graph according to the multiple nodes and edges.

[0117] Furthermore, in one embodiment, the model update module 408 is also used to perform graph optimization processing on the computation graph to obtain an optimized computation graph; and generate multiple target operators based on the optimized computation graph and parameters of the hardware to be deployed to update the large language model.

[0118] Each module in the above-mentioned large language model performance optimization device 400 can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each of the above modules.

[0119] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data parameters, processing parameters, objective functions, calculation graphs, target operators and other data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a large language model performance optimization method is implemented.

[0120] Those skilled in the art will understand that Figure 5The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0121] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0123] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0125] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0126] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0127] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for optimizing the performance of a large language model, characterized in that: The method comprises: Acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model; In response to a programming operation for a preset programming primitive, an objective function representing the operation logic of the large language model is obtained; Compiling the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function; According to the calculation graph and the parameters of the hardware to be deployed, multiple target operators are generated, and the large language model is updated based on the multiple target operators to obtain an optimized large language model.

2. The method according to claim 1, characterized in that The compiling operation on the objective function according to the data parameter and the processing parameter to obtain a calculation graph corresponding to the objective function includes: Compiling the target function according to the data parameters and the processing parameters to obtain intermediate layer parameters of the target function; According to the intermediate layer parameters, a computation graph corresponding to the objective function is constructed.

3. The method according to claim 2, characterized in that The step of compiling the target function according to the data parameter and the processing parameter to obtain the intermediate layer parameters of the target function includes: Determining a plurality of target programming primitives involved in the target function; Compiling the target function according to the data parameters and the processing parameters to obtain the input and output types and tensor shapes of each target programming primitive; The input and output types of each of the target programming primitives and the tensor shape are used as the intermediate layer parameters.

4. The method according to claim 3, characterized in that: The step of constructing a calculation graph according to the intermediate layer parameters includes: constructing a plurality of nodes and edges according to the tensor shape and the input and output types of each of the target programming primitives; The computation graph is generated according to the plurality of nodes and the edges.

5. The method according to claim 1, characterized in that The step of generating a plurality of target operators according to the computation graph and the parameters of the hardware to be deployed, and updating the large language model based on the plurality of target operators includes: Performing graph optimization processing on the computation graph to obtain an optimized computation graph; According to the optimized computation graph and the parameters of the hardware to be deployed, multiple target operators are generated to update the large language model.

6. The method according to any one of claims 1 to 5, characterized in that: The programming primitive includes a calculation part and a memory access part, and the input and output of each programming primitive include data size, data quantity and data storage location.

7. A large language model performance optimization device, characterized in that: The device comprises: A parameter acquisition module, used to acquire data parameters and processing parameters of language data corresponding to the large language model; the processing parameters are determined according to parameters of the hardware to be deployed corresponding to the large language model; A function programming module, for obtaining a target function representing the operation logic of the large language model in response to a programming operation on a preset programming primitive; A function compiling module, used to compile the objective function according to the data parameters and the processing parameters to obtain a calculation graph corresponding to the objective function; A model updating module is used to generate multiple target operators according to the calculation graph and the parameters of the hardware to be deployed, and update the large language model based on the multiple target operators to obtain an optimized large language model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • LLM output result verification method and device, storage medium and equipment

    CN120542582A

  • Machine learning model optimization method and device, computer equipment, readable storage medium and program product

    CN121707015A