Large model reasoning acceleration method and system
By dividing the inference process of the large model into multiple stages and using different quantitative models for quantization processing, the problems of slow inference speed and high resource consumption of large models are solved, and the inference efficiency and optimal use of resources are achieved.
Patent Information
- Application Number
- CN202510159505.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
During the inference process, large models have problems such as slow inference speed, large computing resource consumption and high memory usage due to the inference process.
By dividing the inference process of the large model into multiple stages and using different quantitative models for each stage for quantization processing, the respective corresponding quantitative models are obtained to accelerate the inference process.
It effectively reduces the storage space and computing overhead of the large model, improves the inference efficiency, and meets the performance and resource efficiency requirements of the large model.
Smart Images

Figure CN120087476A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and particularly to a method and system for accelerating large model inference. Background Art
[0002] With the development of artificial intelligence technology, large models have been widely applied to inference in various fields.
[0003] However, large models usually contain billions or even hundreds of billions of parameters. During the process of inference based on large models, various problems such as slow inference speed may be faced.
[0004] It should be noted that the content of the above related technologies is only information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this specification, nor does it represent that it can become the prior art of this specification. Summary of the Invention
[0005] This specification provides a method and system for accelerating large model inference to avoid at least one of the above technical problems.
[0006] In a first aspect, this specification provides a method for accelerating large model inference. The inference process of the large model includes multiple stages, and different stages correspond to different quantization models. The quantization model is obtained by quantizing a pre-trained original large model. The method includes:
[0007] Obtain an inference request, where the inference request is used to request inference on the information to be inferred based on the original large model;
[0008] Infer the information to be inferred based on the quantization model corresponding to each stage, and obtain and output a target inference result corresponding to the information to be inferred;
[0009] Among them, the multiple stages include the Prefill stage and the Decode stage; the quantization model corresponding to the Prefill stage is the first quantization model, and the quantization model corresponding to the Decode stage is the second quantization model; the first quantization model and the second quantization model are obtained by quantizing the original large model using two different quantization schemes;
[0010] The two different quantization schemes include: the first quantization scheme and the second quantization scheme; the first quantization scheme is used to obtain the first quantization model, and the second quantization scheme is used to obtain the second quantization model; the first quantization scheme is determined based on the first inference requirement of the Prefill stage, and the second quantization scheme is determined based on the second inference requirement of the Decode stage.
[0011] In a second aspect, this specification provides a system for accelerating large model inference, including:
[0012] At least one storage medium storing at least one instruction set for accelerating large model inference;
[0013] At least one processor communicatively connected to the at least one storage medium, wherein when the at least one processor runs, it reads the at least one instruction set and executes the method described in the first aspect according to the instructions of the at least one instruction set.
[0014] In a third aspect, this specification provides a computer-readable non-transitory storage medium, wherein the computer-readable non-transitory storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to implement the method described in the first aspect.
[0015] As can be seen from the above technical solutions, the method and system for accelerating large model inference provided in this specification divide the inference process of the large model into multiple stages, and for each stage, the pre-trained original large model is quantized separately to obtain the quantization models corresponding to each stage. Correspondingly, when an inference request is obtained, the information to be inferred requested by the inference request is accelerated and inferred based on the quantization models corresponding to each stage. Exemplarily, the multiple stages include the Prefill stage and the Decode stage. The quantization model corresponding to the Prefill stage is the first quantization model, and the quantization model corresponding to the Decode stage is the second quantization model. The first quantization model and the second quantization model are obtained by quantizing the original large model using two different quantization schemes. For example, the two different quantization schemes include: the first quantization scheme and the second quantization scheme. The first quantization scheme is used to obtain the first quantization model, and the second quantization scheme is used to obtain the second quantization model. Among them, the first quantization scheme is determined based on the first inference requirement of the Prefill stage, and the second quantization scheme is determined based on the second inference requirement of the Decode stage. It can effectively reduce the deployment cost of the large model and improve the inference efficiency.
[0016] Other functions of the method and system for accelerating large model inference provided in this specification will be partially listed in the following description. The creative aspects of the method and system for accelerating large model inference provided in this specification can be fully explained through practice or the use of the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0018] Figure 1 Schematic diagram of the application scenario of the method for accelerating large model inference provided by the embodiments of this specification;
[0019] Figure 2 Schematic diagram of the structure of the system for accelerating large model inference provided by the embodiments of this specification;
[0020] Figure 3 Schematic diagram of the process of the method for accelerating large model inference provided by the embodiments of this specification;
[0021] Figure 4 Schematic diagram of the application scenario where the large model in the embodiments of this specification is a large language model;
[0022] Figure 5 Schematic diagram of the principle of quantization processing provided by the embodiments of this specification;
[0023] Figure 6 Schematic diagram of the principle of the method for accelerating large model inference provided by the embodiments of this specification. Detailed implementation manners
[0024] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are only examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0025] It should be understood that the terms "include" and "have" and any variations thereof in the embodiments of this specification are intended to cover but not be exclusive of inclusion. For example, a product or device including a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0026] The term "and / or" in the embodiments of this specification describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0027] In the embodiments of this specification, the term "a plurality of" refers to two or more, and other quantifiers are similar thereto.
[0028] The terms "first", "second", "third", etc. in this specification are used to distinguish similar or like objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be interchanged under appropriate circumstances, for example, it is possible to implement in an order other than those given in the illustrations or descriptions of the embodiments of this specification.
[0029] The term "unit / module" used in this specification refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to the element.
[0030] For the convenience of readers' understanding of this specification, the application scenarios of this specification are now introduced.
[0031] The technical solutions provided in this specification are applicable to scenarios that require inference based on large models. Inference can be understood as a process of using a trained large model to process input data and obtain and output results.
[0032] For example, in the case where the large model is a large language model (which can be abbreviated as LLM), inference can be understood as a process of inputting text to the large language model and obtaining and outputting corresponding responses or continued writing content.
[0033] A large language model can be understood as an artificial intelligence model based on deep learning technology with billions or more parameters, capable of understanding and generating human language.
[0034] Exemplarily, the technical solutions provided in this specification can be applied to scenarios such as natural language processing, computer vision, speech recognition and synthesis, recommendation systems, healthcare, the financial field, autonomous driving, etc.
[0035] Among them, the scenarios of natural language processing can include: text generation, machine translation, question answering systems, sentiment analysis, etc. The scenarios of computer vision can include: image recognition and classification, object detection, image generation, etc. The scenarios of speech recognition and synthesis can include: speech-to-text and text-to-speech, etc. The scenarios of healthcare can include disease diagnosis assistance and drug research and development, etc.
[0036] Taking the scenario of a question answering system as an example:
[0037] The user can raise questions to be answered to the question-answering system. After obtaining the questions to be answered raised by the user, the question-answering system can perform accelerated reasoning based on the method provided in this specification to quickly and accurately feedback the reasoning results corresponding to the questions to be answered to the user.
[0038] It is worth noting that the above examples are only used to illustrate the application scenarios to which the technical solutions of this specification can be applied, and cannot be understood as limiting the application scenarios. When examples are mentioned later in this specification, the scenario of the problem system will be used as an example for exemplary explanation.
[0039] Figure 1 The schematic diagram of the application scenario of the method for accelerating the reasoning of a large model in the embodiment of this specification is shown in FIG. Figure 1 The application scenario 100 shown. Figure 1 As shown, the application scenario 100 may include a target user 101 , a client 102 , a server 103 , and a network 104 .
[0040] The target user 101 may be a user who triggers the method for accelerating the reasoning of a large model provided in this specification. For example, the target user 101 may perform a target operation on the client 102 to achieve accelerated reasoning of a large model.
[0041] In some embodiments, the target user 101 may initiate an inference request on the client 102 to trigger the large model to perform accelerated inference on the information to be inferred based on the method provided in this specification. For example, the target user 101 may input the information to be inferred in the client 102 to initiate an inference request to the client 102, thereby triggering the large model to perform accelerated inference on the information to be inferred based on the method provided in this specification.
[0042] The client 102 may be an electronic device that provides interactive functions to the target user 101. For example, the client 102 may provide an interactive interface to the target user 101, and the target user 101 may perform interactive operations (such as target operations) in the interactive page. In some embodiments, the client 102 executes the large model reasoning acceleration method described in this specification in response to detecting the target operation triggered by the target user 101. At this time, the client 102 may store data or instructions for executing the large model reasoning acceleration method described in this specification, and may execute or be used to execute the data or instructions. In some embodiments, the client 102 may include a hardware device with a data information processing function and the necessary programs required to drive the hardware device to work, so as to execute the large model reasoning acceleration method described in this specification.
[0043] In some embodiments, the client 102 may include a mobile device, a tablet computer, a laptop computer, a built-in device of a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device, or the like, or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination. In some embodiments, the smart mobile device may include a smart phone, a personal digital assistant, a gaming device, a navigation device, etc., or any combination thereof. In some embodiments, the built-in device in a motor vehicle may include an in-vehicle computer, an in-vehicle TV, etc. In some embodiments, the client 102 may include a collection device for collecting a target operation. For example, the collection device in the client 102 may be a keyboard, and the target user 101 inputs information to be inferred to the client 102 based on the keyboard.
[0044] In some embodiments, one or more applications (APPs) may be installed on the client 102. The APP can provide the target user 101 with the ability to interact with the outside world through the network 104 and an interface. The APP includes but is not limited to: web browser APP programs, search APP programs, chat APP programs, shopping APP programs, video APP programs, financial management APP programs, instant messaging tools, email clients, social platform software, and so on. In some embodiments, a target APP may be installed on the client 102. The target APP can collect target operations for the client 102.
[0045] As Figure 1 shown, the client 102 can be communicatively connected to the server 103. Among them, the server 103 can be communicatively connected to one client 102 or multiple clients 102. In some embodiments, the client 102 can interact with the server 103 through the network 104 to receive or send messages, etc. For example, the client 102 can interact with the server 103 through the network 104 to send information to be inferred to the server 103.
[0046] The server 103 can be a server that provides various services. For example, the server 103 can be a cloud server or a local server. The server 103 can be communicatively connected to one client 102 and receive the data sent by the client 102, or can be communicatively connected to multiple clients 102 and receive the data sent by each client 102 respectively.
[0047] In some embodiments, the method for accelerating large model inference described in this specification can be executed on server 103. At this time, server 103 may store data or instructions for executing the method for accelerating large model inference described in this specification, and may execute or be used to execute the data or instructions. Server 103 may include a hardware device with data information processing capabilities and necessary programs for driving the hardware device to work.
[0048] Network 104 is a medium for providing a communication connection between client 102 and server 103. Network 104 can facilitate the exchange of information or data. As Figure 1 shown, client 102 and server 103 can be respectively connected to network 104, and transmit information or data to each other through network 104.
[0049] In some embodiments, network 104 can be any type of wired or wireless network, or a combination thereof. For example, network 104 can include a cable network, a wired network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network TM, a short-range wireless network (ZigBee TM), a near field communication (NFC) network, or a similar network.
[0050] In some embodiments, network 104 can include one or more network access points. For example, network 104 can include a wired or wireless network access point, such as a base station or an Internet exchange point, through which one or more components of client 102 and server 103 can be connected to network 104 to exchange data or information.
[0051] It is worth noting that Figure 1 the numbers of client 102, server 103, and network 104 in
[0052] That is to say, Figure 1 and forFigure 1 The above description is only used to exemplarily illustrate the application scenarios to which the method for accelerating large model inference in this specification may apply, and should not be construed as a limitation on the application scenarios.
[0053] Figure 2 FIG. shows a hardware structure diagram of a system 200 for accelerating large model inference provided according to an embodiment of this specification. The system 200 can execute the method for accelerating large model inference described in this specification. The method for accelerating large model inference is introduced in other parts of this specification. When the method for accelerating large model inference is executed on the client 102, the system 200 can be the client 102. When the method for accelerating large model inference is executed on the server 103, the system 200 can be the server 103. When the method for accelerating large model inference is partially executed on the client 102 and partially executed on the server 103, the system 200 can be a system including the client 102 and the server 103.
[0054] As Figure 2 shown, the system 200 may include at least one storage medium 203 and at least one processor 202. In some embodiments, the system 200 may further include a communication port 204 and an internal communication bus 201. The system 200 may further include I / O components 205.
[0055] The internal communication bus 201 can connect different system components. For example, the internal communication bus 201 can connect the storage medium 203, the processor 202, the communication port 204, and the I / O components 205.
[0056] The I / O components 205 support querying the input / output between the system 200 and other components.
[0057] The communication port 204 is used for data communication between the system 200 and the outside world. For example, the communication port 204 can be used for data communication between the system 200 and the network 104. The communication port 204 can be a wired communication port or a wireless communication port.
[0058] The storage medium 203 may include a data storage device. The data storage device can be a non-temporary storage medium or a temporary storage medium. For example, the data storage device can include one or more of a magnetic disk 2031, a read-only storage medium (ROM) 2032, or a random access storage medium (RAM) 2033. The storage medium 203 further includes at least one instruction set stored in the data storage device. The instruction set includes computer program code, and the computer program code can include programs, routines, objects, components, data structures, processes, modules, etc. for executing the method for accelerating large model inference provided in this specification.
[0059] At least one processor 202 can be communicatively connected to at least one storage medium 203. The at least one processor 202 is configured to execute the above-mentioned at least one instruction set. When the system 200 is running, the at least one processor 202 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the query method provided in this specification. The processor 202 can execute all steps included in the method for accelerating large model inference. The processor 202 can be in the form of one or more processors. In some embodiments, the processor 202 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.
[0060] For illustrative purposes only, only one processor 202 is shown in the system 200 in the drawings. However, it should be noted that the system 200 in this specification can also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification can be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 202 of the system 200 executes step A and step B, it should be understood that step A and step B can also be jointly or separately executed by two different processors 202 (for example, the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0061] Combined with the content of the above background technology, large models are widely used in inference in various fields. However, due to characteristics such as the large number of parameters of the large model itself, there are problems in many aspects such as the inference speed during the inference process of the large model. For example, in actual deployment and inference processes, there are also huge problems of computational resource consumption and memory occupation.
[0062] To avoid the problems existing in the prior art, the inventors of this specification thought of first performing quantization processing on the trained large model based on model quantization technology to obtain a quantized model, and then performing inference based on the quantized model.
[0063] Among them, quantization processing can be understood as a process of converting the numerical values (usually 32-bit floating-point numbers) in the trained large model into a lower-precision representation (such as 8-bit or 4-bit fixed-point numbers). That is, quantization processing can reduce the size and computational resource consumption of the trained large model.
[0064] The inventor further realized that if the quantization model is based on a whole trained large model. That is, the quantization model is obtained by quantizing the trained large model based on a single quantization scheme. However, the large model has corresponding requirements in terms of performance, resource efficiency, etc. Then, the above quantization model is very likely to fail to meet the requirements of the large model's performance and resource efficiency.
[0065] Therefore, to further avoid the above problems, this specification proposes a technically creative concept: further analysis of the inference process of the large model shows that the inference process of the large model includes multiple stages. Then, different stages can correspond to different quantization models. On this basis, for the information that needs to be inferred by the large model, inference can be performed based on the quantization models corresponding to each stage of the large model to obtain and output the final inference result. Moreover, different quantization models can be implemented based on different quantization schemes. Further, different quantization schemes can be implemented based on different inference requirements.
[0066] Based on the above analysis, the above technical solution can reduce the storage space and computational overhead of the large model by splitting the inference process of the large model into multiple stages, configuring different quantization models for different stages, and completing the inference according to each quantization model. Moreover, it can relatively meet the requirements of different aspects such as the performance and resource efficiency of the large model at the same time. In particular, different quantization models correspond to different quantization schemes, and different quantization schemes are implemented based on different inference requirements, so the accuracy, effectiveness, and reliability of the entire inference process of the large model can be ensured.
[0067] Exemplarily, based on the above technical concept, this specification provides a method for accelerating the inference of a large model. Among them, the inference process of the large model includes multiple stages, different stages correspond to different quantization models, and the quantization model is obtained by quantizing a pre-trained original large model.
[0068] Among them, different types of large models can have different inference goals. To achieve the inference goal, the inference process of the large model may be divided into different stages.
[0069] For example, the large model can be a large language model. The inference process of the large language model can include: Prefill (pre-filling) stage and Decode (decoding) stage.
[0070] In practical applications, the Prefill stage and the Decode stage are usually closely connected. The Prefill stage is the first stage of large language model inference. The Prefill stage mainly processes the context information input to the large language model, that is, this stage needs to perform parallel computations on a relatively long input sequence. The Decode stage is the second stage of large language model inference. After the context is processed in the Prefill stage, the Decode stage mainly generates tokens one by one to generate the complete text output. The Prefill stage and the Decode stage jointly determine the quality and coherence of the finally generated text.
[0071] Among them, a token is the basic processing unit of text, which can be a word, a subword, or a character, and is the smallest unit when the large language model processes text.
[0072] In this embodiment, each stage of the large model inference process has a corresponding quantization model. For example, the Prefill stage and the Decode stage each have their own corresponding quantization models.
[0073] That is to say, in this embodiment, the inference process of the large model is accelerated at the granularity of stages.
[0074] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the method for accelerating large model inference provided by the embodiments of this specification.
[0075] As Figure 3 shown, this method includes the following S301 and S303:
[0076] S301: Obtain an inference request, where the inference request is used to request inference on the information to be inferred based on the original large model.
[0077] Combined with the above analysis, it can be seen that the method provided in this specification can be applied to different application scenarios. For different application scenarios, the content of the information to be inferred may be different.
[0078] For example, for the scenario of a question-and-answer system, the information to be inferred can be the question initiated by the target user.
[0079] In addition, combined with the above examples, the entity that executes the method for accelerating large model inference in this specification can be a client or a server. This specification mainly takes the server as an example for exemplary elaboration.
[0080] Exemplarily, combined with Figure 4 it can be seen that this step can be understood as:
[0081] The client 102 can output an interactive interface. The target user 101 can initiate an inference request to the client 102 through the interactive interface based on their inference needs, such as inputting the information to be inferred to the client 102 through the interactive interface. For example, the information to be inferred can be a question input by the target user 101 through the interactive interface.
[0082] Correspondingly, the client 102 receives the information to be inferred input by the target user 101. The client 102 sends the information to be inferred to the server 103 through the network 104 ( Figure 4 not shown in the figure).
[0083] Correspondingly, the server 103 receives the information to be inferred sent by the client 102.
[0084] S302: Infer the information to be inferred based on the quantization models corresponding to each stage, and obtain and output the target inference result corresponding to the information to be inferred.
[0085] Combined with the above example and Figure 4 , this step can be understood as:
[0086] The server 103 infers the information to be inferred according to each quantization model to obtain the target inference result. The server 103 sends the target inference result to the client 102 through the network 104.
[0087] Correspondingly, the client 102 receives the target inference result sent by the server 103. The client 102 outputs the target inference result through the interactive interface. For example, in the scenario of a question-and-answer system, the target inference result is the answer corresponding to the question.
[0088] Based on the above analysis of S301 and S303, it can be seen that in this embodiment, for multiple stages of the inference process, there are respective corresponding quantization models. During the inference process, the server performs inference based on the quantization models corresponding to each stage. It can effectively reduce the deployment cost of the large model and improve the inference efficiency.
[0089] Combined with the above analysis, it can be seen that the inference process can include the above-mentioned multiple stages. In some embodiments, for different stages, the server can adopt different quantization schemes to quantize the original large model to obtain the quantization models corresponding to each stage.
[0090] Exemplarily, in some embodiments, the multiple stages may include a Prefill stage and a Decode stage. When the multiple stages include a Prefill stage and a Decode stage, the quantization model corresponding to the Prefill stage may be a first quantization model, and the quantization model corresponding to the Decode stage may be a second quantization model. Among them, the first quantization model and the second quantization model may be obtained by quantizing the original large model using two different quantization schemes.
[0091] In this embodiment, in order to make the quantization models corresponding to the Prefill stage and the Decode stage different, the server may perform quantization processing using different quantization schemes. That is, the server uses two different quantization schemes to quantize the original large model to respectively obtain a first quantization model corresponding to the Prefill stage and a second quantization model corresponding to the Decode stage.
[0092] In this embodiment, by using different quantization schemes to quantize the original large model, the server can effectively and reliably obtain the quantization models corresponding to the different quantization schemes.
[0093] In some embodiments, for a target stage among the multiple stages, the target quantization model corresponding to the target stage is determined based on the target inference requirement of the target stage. Wherein, the target stage is any stage among the multiple stages.
[0094] Exemplarily, the inference requirements or purposes of the large model are different in different stages. For example, in combination with the above example, the inference requirement of the Prefill stage may be to provide an initial text basis, while the inference requirement of the Decode stage is to gradually expand on the basis of the Prefill stage to generate a complete text output.
[0095] Therefore, in this embodiment, the server may first determine the different inference requirements of different stages; then, for each stage, determine the quantization scheme for this stage according to the inference requirement of this stage; finally, quantize the original large model according to the quantization scheme for this stage to obtain the quantization model for this stage.
[0096] In some embodiments, in combination with the above analysis, when the multiple stages include a Prefill stage and a Decode stage, there are two quantization schemes. The two different quantization schemes may include: a first quantization scheme and a second quantization scheme. The first quantization scheme may be used to obtain the first quantization model, and the second quantization scheme may be used to obtain the second quantization model. The first quantization scheme may be determined based on the first inference requirement of the Prefill stage, and the second quantization scheme may be determined based on the second inference requirement of the Decode stage.
[0097] Exemplarily, the server can first determine the inference requirements in the Prefill stage. For ease of distinction, this inference requirement can be referred to as the first inference requirement. Then, the server can determine the quantization scheme applicable to the first inference requirement. For ease of distinction, this quantization scheme can be referred to as the first quantization scheme.
[0098] Correspondingly, as Figure 5 shown, the server can perform quantization processing on the original large model based on the first quantization scheme to obtain the first quantized model.
[0099] Similarly, the server can first determine the inference requirements in the Decode stage, that is, the second inference requirement. Then, the server can determine the quantization scheme applicable to the second inference requirement, that is, the second quantization scheme.
[0100] Correspondingly, as Figure 5 shown, the server can perform quantization processing on the original large model based on the second quantization scheme to obtain the second quantized model.
[0101] In this embodiment, the server determines the quantization scheme through the inference requirements, so as to obtain the corresponding quantized model on this basis, and then performs accelerated inference based on the quantized model. It fully considers the inference characteristics and requirements in different stages of the inference process, improves the effectiveness and reliability of the accelerated inference. In addition, it avoids performance loss in the inference process and improves the resource utilization efficiency. Moreover, a relative balance between performance and efficiency can be achieved.
[0102] Combined with the above analysis, it can be seen that the inference requirements in the Prefill stage (i.e., the first inference requirement) are mainly to process a large amount of context information. The inference requirements in the Decode stage (i.e., the first inference requirement) are mainly to generate tokens.
[0103] Therefore, in some embodiments, the first inference requirement may include a precision requirement; the second inference requirement includes a resource occupancy requirement.
[0104] In this embodiment, the server obtains the first quantized model through the precision requirement and obtains the second quantization through the resource occupancy requirement. It can avoid reducing the inference precision due to complex calculations in the Prefill stage during the inference acceleration process. In addition, quantization processing can be performed through the resource occupancy requirement to reduce the memory occupancy, which is helpful for inference acceleration.
[0105] In some embodiments, the first quantization scheme may include the W8A8 quantization scheme; the second quantization scheme may include the W4A16 quantization scheme.
[0106] Exemplarily, the original large model includes model parameters. For example, the model parameters can include weights and activation values. Weights are used for neural network calculations. W8 represents storing weights with 8-bit precision, and W4 represents storing weights with 4-bit precision. Activation values are intermediate results during the calculation process of the original large model. A8 represents storing activation values with 8-bit precision, and A16 represents storing activation values with 16-bit precision.
[0107] Correspondingly, the W8A8 quantization scheme can be understood as a technical solution that converts both the weights and activation values of the original large model into 8-bit fixed-point numbers. The W4A16 quantization scheme can be understood as a technical solution that quantizes the weights of the original large model into 4-bit fixed-point numbers and quantizes the activation values into 16-bit fixed-point numbers.
[0108] Among them, fixed-point numbers can be understood as a numerical representation format, which has a simpler hardware implementation and lower storage overhead compared to floating-point numbers.
[0109] Continuing to refer to Figure 5 It can be known that in the case of obtaining the original large model, the server can perform quantization processing on the original large model based on the W8A8 quantization scheme (which can be simply referred to as W8A8 quantization processing) to obtain a first quantization model. The first quantization model can also be called the W8A8 quantization model.
[0110] Similarly, the server can perform quantization processing on the original large model based on the W4A16 quantization scheme (which can be simply referred to as W4A16 quantization processing) to obtain a second quantization model. The second quantization model can also be called the W4A16 quantization model.
[0111] Relatively speaking, the W8A8 quantization scheme has relatively high quantization accuracy and will cause relatively small loss to the performance of the original large model. Therefore, by performing inference using the first quantization model of the W8A8 quantization scheme in the Prefill stage, the accuracy of context processing can be ensured. In addition, large precision losses can be avoided. Moreover, by reasonably utilizing 16-bit internal values, the inference quality can be ensured.
[0112] The W4A16 quantization scheme adopts a hybrid quantization strategy of 4-bit weights and 16-bit activation values. It can significantly reduce the weight storage space and can maintain a relatively high precision of activation values. Therefore, by performing inference using the second quantization model of the W4A16 quantization scheme in the Decode stage, the memory occupancy can be significantly reduced. In addition, by using lower-bit weights, the inference efficiency can be improved. Moreover, resource waste caused by using 8-bit quantization throughout the entire inference process can be avoided.
[0113] Correspondingly, in the case where the multiple stages include a Prefill stage and a Decode stage, and the quantization model corresponding to the Prefill stage is the first quantization model, and the quantization model corresponding to the Decode stage is the second quantization model, S303 may include the following steps 1 and 2:
[0114] Step 1: Perform inference on the information to be inferred in the Prefill stage based on the first quantization model to obtain an intermediate inference result.
[0115] Exemplarily, still taking the above Q&A scenario as an example, combined with Figure 6 it can be known that the information to be inferred may be the input text of the question initiated by the target user.
[0116] In the Prefill stage, the server may process the context information of the input text based on the first quantization model, that is, the W8A8 quantization model.
[0117] In some embodiments, the first quantization model includes first quantization parameters, and the first quantization parameters include first weights and / or activation values; Step 1 may include: performing inference on the information to be inferred in the Prefill stage based on the first quantization parameters to obtain an intermediate inference result.
[0118] Exemplarily, the first quantization model may be the W8A8 quantization model, and the first quantization parameters may be the W8A8 quantization parameters. And the W8A8 quantization parameters may include 8-bit weights and activation values.
[0119] Correspondingly, in the Prefill stage, the server may perform inference based on the W8A8 quantization parameters (8-bit weights and activation values) in the W8A8 quantization model to obtain the corresponding intermediate inference result.
[0120] For example, when the server obtains the W8A8 quantization model, it may store the model parameters of the W8A8 quantization model (such as 8-bit weights and activation values). For example, the server stores the model parameters of the W8A8 quantization model in the memory. Continuing to refer to Figure 6 , after receiving the input text, the server may load the model parameters of the W8A8 quantization model to process the context information based on the model parameters of the W8A8 quantization model. Thus, the inference in the Prefill stage is completed to obtain an intermediate inference result (that is, the result corresponding to the processed context information).
[0121] Step 2: Perform inference on the intermediate inference result in the Decode stage based on the second quantization model to obtain and output the target inference result corresponding to the information to be inferred.
[0122] Continuing with the above example, in the Decode stage, the server can process the intermediate inference result (i.e., the result corresponding to the processed context information) based on the second quantization model, namely the W4A16 quantization model.
[0123] In some embodiments, the second quantization model includes second quantization parameters, and the second quantization parameters include second weights and / or activation values; step 2 may include: performing inference in the Decode stage on the intermediate inference result based on the second quantization parameters to obtain and output a target inference result corresponding to the information to be inferred.
[0124] Exemplarily, the second quantization model may be the W4A16 quantization model, and the second quantization parameters may be the W4A16 quantization parameters. And the W4A16 quantization parameters may include 4-bit weights and 16-bit activation values.
[0125] Correspondingly, in the Decode stage, the server can perform inference based on the W4A16 quantization parameters (4-bit weights and 16-bit activation values) in the W4A16 quantization model to obtain and output the final target inference result.
[0126] For example, when the server obtains the W4A16 quantization model, it can store the model parameters of the W4A16 quantization model (such as 4-bit weights and 16-bit activation values). For example, the server stores the model parameters of the W4A16 quantization model in the memory. Continuing to refer to Figure 6 , after the server receives the intermediate inference result (i.e., the result corresponding to the processed context information), it can load the model parameters of the W4A16 quantization model to generate tokens one by one based on the model parameters of the W4A16 quantization model. Thus, the inference in the Decode stage is completed, and the generated result (such as the answer corresponding to the question) is obtained and output.
[0127] Combining the above analysis of step 1 and step 2, it can be seen that in this embodiment, the server can first load the model parameters of the W8A8 quantization model to perform inference in the Prefill stage, and then switch to loading the parameters of the W4A16 quantization model to perform inference in the Decode stage. This not only ensures the accuracy of the inference but also optimizes the resource occupancy of the inference. In addition, it can achieve a fast and smooth automatic switch from the inference in the Prefill stage to the inference in the Decode stage without affecting the overall inference performance.
[0128] As can be seen from the above analysis, in some embodiments, the large model includes a large language model; the information to be inferred includes text information to be inferred (such as input text, specifically the input text corresponding to a question); the intermediate inference result includes the processing result of the context information of the text information to be inferred; the target inference result (such as the generated result, specifically the generated result corresponding to the question, that is, the answer corresponding to the question) is determined based on the tokens generated from the processing result of the context information.
[0129] When the method of this specification is applied to other scenarios, the large model can be other types of models. However, the principle of accelerating the inference process can be referred to the above examples, and will not be listed one by one here.
[0130] It should be noted that the above examples are only used to exemplarily illustrate the possible implementation manners of the method for accelerating large model inference in this specification, and should not be construed as a limitation on the implementation manners of the method for accelerating large model inference in this specification. Exemplarily, based on the above technical concepts, some of the above technical features can be combined to obtain new embodiments; new technical features can be added based on the above examples to obtain new embodiments; some technical features can be reduced based on the above examples to obtain new embodiments; some of the above technical features can be replaced with other technical features; the order of some of the above technical features can be adjusted to obtain new embodiments, etc., and will not be listed one by one here.
[0131] Based on the above technical concept, this specification also provides a computer-readable non-transitory storage medium, in which at least one instruction set is stored. When the at least one instruction set is executed by a processor, the steps of the method for accelerating large model inference described in this specification are implemented.
[0132] In some possible embodiments, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on system 200, the program code is used to cause system 200 to execute the steps of the method for accelerating large model inference described in this specification. The program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) including program code and can run on system 200. However, the program product of this specification is not limited to this. In this specification, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system. The program product can be any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection with one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on system 200, partially on system 200, executed as an independent software package, partially on system 200 and partially on a remote system 200, or entirely on a remote system 200.
[0133] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0134] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to embrace various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0135] In addition, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0136] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may well mark out some of the devices as separate embodiments when reading this specification. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0137] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, references, etc. (excluding any historical prosecution files associated therewith), are hereby incorporated by reference for all purposes relevant hereto, e.g., in the specification and claims of this application. However, in the event of any inconsistency or conflict between the description, definition, and / or terminology of such materials and those used in this application, the description, definition, and / or terminology used in this application shall prevail.
[0138] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Accordingly, the embodiments disclosed in this specification are presented by way of example and not limitation. Those skilled in the art may implement the application in this specification by taking alternative configurations based on the embodiments in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A method for accelerating large model reasoning, wherein: The reasoning process of the large model includes multiple stages, and different stages correspond to different quantization models. The quantization model is obtained by quantizing the pre-trained original large model. The method includes: Obtaining an inference request, wherein the inference request is used to request to perform inference on the inference information based on the original large model; Reasoning the information to be inferred based on the quantization models corresponding to each of the stages, and obtaining and outputting a target reasoning result corresponding to the information to be inferred; The multiple stages include a Prefill stage and a Decode stage; the quantization model corresponding to the Prefill stage is a first quantization model, and the quantization model corresponding to the Decode stage is a second quantization model; the first quantization model and the second quantization model are obtained by quantizing the original large model using two different quantization schemes; The two different quantization schemes include: a first quantization scheme and a second quantization scheme; the first quantization scheme is used to obtain the first quantization model, and the second quantization scheme is used to obtain the second quantization scheme; the first quantization scheme is determined based on the first reasoning requirement of the Prefill stage, and the second quantization scheme is determined based on the second reasoning requirement of the Decode stage.
2. The method according to claim 1, wherein: The first reasoning requirement includes an accuracy requirement; the second reasoning requirement includes a resource occupancy requirement.
3. The method according to claim 1, wherein: The first quantization scheme includes a W8A8 quantization scheme; the second quantization scheme includes a W4A16 quantization scheme.
4. The method according to any one of claims 1 to 3, wherein: The reasoning the information to be inferred based on the quantization models corresponding to each of the stages to obtain and output a target reasoning result corresponding to the information to be inferred includes: Performing reasoning at the Prefill stage on the information to be inferred based on the first quantization model to obtain an intermediate reasoning result; and The reasoning of the Decode stage is performed on the intermediate reasoning result based on the second quantization model to obtain and output a target reasoning result corresponding to the information to be reasoned.
5. The method according to claim 4, wherein: The first quantization model includes a first quantization parameter, and the first quantization parameter includes a first weight and / or activation value; The reasoning of the information to be inferred in the Prefill stage based on the first quantization model to obtain an intermediate reasoning result includes: The information to be inferred is inferred in the Prefill stage based on the first quantization parameter to obtain an intermediate inference result.
6. The method according to claim 5, wherein: The second quantization model includes a second quantization parameter, and the second quantization parameter includes a second weight and / or an activation value; the reasoning of the Decode stage on the intermediate reasoning result based on the second quantization model to obtain and output a target reasoning result corresponding to the information to be reasoned includes: The inference of the Decode stage is performed on the intermediate inference result based on the second quantization parameter to obtain and output the target inference result.
7. The method according to claim 4, wherein: The large model includes a large-scale language model; the information to be inferred includes text information to be inferred; the intermediate inference result includes a processing result of context information of the text information to be inferred; and the target inference result is determined by a word element generated according to the processing result of the context information.
8. A system for accelerating large model reasoning, comprising: At least one storage medium storing at least one instruction set for large model reasoning acceleration; At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, the at least one instruction set is read, and the method as described in any one of claims 1 to 7 is executed according to the instructions of the at least one instruction set.