Post-quantification method and system for large language model end-side deployment
By adopting post-quantization methods in the end-side deployment of large language models, including channel smoothing, channel grouping and rearrangement and integer polynomial computing, the energy consumption and computing resource limitation problems of large language models during the end-side deployment of terminal devices is solved, and efficient compression and accuracy improvement are achieved.
Patent Information
- Application Number
- CN202411831403.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-13
AI Technical Summary
Large language models face high energy consumption and limited computing resources when deploying end-side equipment, resulting in large memory usage and large latency during inference.
The post-quantization method is adopted to deploy to the end side of the large language model. By determining the channel value range, quantization is performed based on the preset channel smoothing algorithm and the channel grouping and rearrangement algorithm, quantization parameters are calculated, and replacement index calculation is calculated using integer polynomials in the nonlinear layer inference process.
It realizes efficient compression of large language models and is suitable for end-side device deployment, saving computing and communication overhead, and significantly improving the overall accuracy of the quantization process.
Smart Images

Figure CN119990228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software and hardware collaborative optimization, and in particular to a post-quantization method and system for terminal-side deployment of a large language model. Background Art
[0002] The emergence of Transformer-based Large Language Models (LLM) marks the beginning of a new era, and it has achieved great success in understanding and generating language. However, large language models face problems such as large memory usage, high latency, and high power consumption during reasoning. In order to apply LLM more widely to actual production and life, the end-side deployment of large language models has also become a key development area. Since end-side devices usually face strict energy consumption and computing resource constraints, how to efficiently run large language models in these restricted environments has become an important challenge that needs to be solved urgently. Summary of the invention
[0003] The present invention provides a post-quantization method and system for the terminal-side deployment of a large language model, which is used to solve the defects of energy consumption and computing resource limitations in the deployment of large language models on terminal-side devices in the prior art. The present invention can achieve efficient compression of large language models and is suitable for deployment on terminal-side devices, saving computing and communication overhead, and significantly improving the overall accuracy of the quantization process.
[0004] The present invention provides a post-quantization method for terminal-side deployment of a large language model, comprising: determining a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating a quantization parameter of the smoothed channel; performing channel quantization according to the quantization parameter of the smoothed channel to obtain a quantized matrix; and the quantized matrix is used for deployment on the terminal side to run the quantized large language model on the terminal side.
[0005] According to a post-quantization method for terminal-side deployment of a large language model provided by the present invention, the method further includes: performing channel grouping rearrangement based on a preset channel grouping rearrangement algorithm according to the channel value range, and calculating the quantization parameters of each grouped channel; the quantization parameters of the grouped channels are used to determine the quantized matrix.
[0006] According to a post-quantization method for terminal-side deployment of a large language model provided by the present invention, the method further includes: in the nonlinear layer inference process of the large language model, replacing the exponential calculation with integer polynomial calculation to perform full quantization inference of the large language model.
[0007] According to a post-quantization method for terminal-side deployment of a large language model provided by the present invention, the channel value range includes an activation value range and a weight value range; the channel smoothing is performed based on a preset channel smoothing algorithm according to the channel value range, including: calculating a channel smoothing factor according to the activation value range and the weight value range; the channel smoothing factor is used to characterize the data distribution in the channel so as to transfer part of the difficulty of quantizing the activation value to the weight value; and channel smoothing is performed according to the channel smoothing factor.
[0008] According to a post-quantization method for terminal-side deployment of a large language model provided by the present invention, channel grouping rearrangement is performed based on a preset channel grouping rearrangement algorithm according to the channel value range, including: according to the activation value range of each activation channel, based on a clustering algorithm, the activation channels with the activation value range in a preset similar range are divided into a group to obtain multiple groups of activation channels, so as to calculate quantization parameters for each group of activation channels.
[0009] The present invention also provides a post-quantization system for terminal-side deployment of a large language model, including: a determination module, used to determine a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; a smoothing module, used to perform channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculate the quantization parameters of the smoothed channel; a quantization module, used to perform channel quantization according to the quantization parameters of the smoothed channel to obtain a quantized matrix; the quantized matrix is used to be deployed on the terminal side to run the quantized large language model on the terminal side.
[0010] According to a post-quantization system for end-side deployment of a large language model provided by the present invention, it also includes: a grouping rearrangement module, which is used to perform channel grouping rearrangement according to the channel value range and based on a preset channel grouping rearrangement algorithm, and calculate the quantization parameter of each grouping channel; the quantization parameter of the grouping channel is used to determine the quantized matrix.
[0011] According to a post-quantization system for terminal-side deployment of a large language model provided by the present invention, the system also includes: a replacement module, which is used to replace exponential calculations with integer polynomial calculations in the nonlinear layer inference process of the large language model to perform full quantization inference of the large language model.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements any of the post-quantization methods for end-side deployment of a large language model as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the post-quantization method for terminal-side deployment of a large language model as described in any one of the above is implemented.
[0014] The present invention provides a post-quantization method and system for terminal-side deployment of a large language model, the method comprising: determining a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating the quantization parameters of the smoothed channel; performing channel quantization according to the quantization parameters of the smoothed channel to obtain a quantized matrix; the quantized matrix is used for deployment on a terminal-side device to achieve low-cost and efficient operation of a large language model. The present invention can achieve efficient compression of a large language model and is suitable for deployment on a terminal-side device, saving computing and communication overhead, and significantly improving the overall accuracy of the quantization process. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 This is one of the flow charts of a post-quantization method for terminal-side deployment of a large language model provided by the present invention.
[0017] Figure 2 This is the second flow chart of a post-quantization method for terminal-side deployment of a large language model provided by the present invention.
[0018] Figure 3 This is one of the structural schematic diagrams of a post-quantization system for terminal-side deployment of a large language model provided by the present invention.
[0019] Figure 4 This is the second structural diagram of a post-quantization system for terminal-side deployment of a large language model provided by the present invention.
[0020] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] With the rapid growth of the parameter scale of large language models, their deployment on end-side devices faces major challenges. In order to meet these challenges, the present invention designs a post-quantization PTQ (Post Training Quantization) method for the end-side deployment of large language models for the deployment of large language models on end-side devices. The present invention introduces an enhanced channel smoothing technology based on channel value ranges, and combines channel reordering to alleviate quantization errors related to channel value range differences and activation outliers, reduce the need for quantization and dequantization steps in the inference process, and make it more suitable for end-side deployment. Activation outliers refer to channels whose maximum or minimum values are much larger than other channels, which results in that setting the quantization range to cover a large value range may cause channels with small values to have large quantization errors, while setting it to cover a small value range may cause significant truncation of outliers, resulting in significant quantization errors.
[0023] Please refer to Figure 1 , Figure 1 One of the flow charts of a post-quantization method for terminal-side deployment of a large language model provided by the present invention.
[0024] The present invention provides a post-quantization method for terminal-side deployment of a large language model, including: 101: Determine a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model.
[0025] 102: performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating a quantization parameter of the smoothed channel.
[0026] As a preferred embodiment, the channel value range includes an activation value range and a weight value range; according to the channel value range, channel smoothing is performed based on a preset channel smoothing algorithm, including: calculating a channel smoothing factor according to the activation value range and the weight value range; the channel smoothing factor is used to characterize the data distribution in the channel so as to transfer the activation value to the weight value; and channel smoothing is performed according to the channel smoothing factor.
[0027] In order to solve the problem of activation outliers, in this embodiment, channel smoothing can effectively alleviate the difficulty of activation quantization and reduce the quantization error caused by activation outliers. Specifically, the range of values of each channel of activation and weight is used to calculate the channel smoothing factor to better capture the data distribution in the channel, thereby transferring the quantization difficulty from the activation value to the weight. By scaling the activation channel with a large value by the channel smoothing factor 1 / s and scaling the corresponding weight in the weight channel by s, the adverse effect of the large range of activation values on the quantization accuracy is reduced, which can bring greater performance improvement.
[0028] The expressions related to the channel smoothing algorithm are: in, P x To activate the value range, is the weight value range, X max is the maximum activation value, X min is the minimum activation value, W max is the maximum weight value, W min is the minimum weight value, s is the channel smoothing factor, is the activation hyperparameter, with a value range of [0, 1]. is a weight hyperparameter, and its value range is [0, 1].
[0029] 103: Perform channel quantization according to the quantization parameter of the smoothed channel to obtain a quantized matrix; the quantized matrix is used for deployment on the end side to run the quantized large language model on the end side.
[0030] PTQ is a low-bit quantization of the trained large language model. It does not require the design of the training process of the large language model and has a lower cost of use. PTQ includes the quantization of activations and the quantization of weights. Since weights are the main parameters of large language models, quantizing only weights can effectively reduce the parameter scale, thereby reducing the cost of memory access and storage.
[0031] Quantization refers to the process of converting a full-precision model to a fixed-point model.
[0032] Please refer to Figure 2 , Figure 2 The second flowchart of a post-quantization method for terminal-side deployment of a large language model provided by the present invention.
[0033] As a preferred embodiment, it also includes: 104: According to the channel value range, the channels are grouped and rearranged based on a preset channel grouping and rearrangement algorithm, and a quantization parameter of each grouped channel is calculated; the quantization parameter of the grouped channel is used to determine a quantized matrix.
[0034] As a preferred embodiment, channel grouping rearrangement is performed according to the channel value range and based on a preset channel grouping rearrangement algorithm, including: according to the activation value range of each activation channel, based on a clustering algorithm, the activation channels with activation value ranges in a preset similar range are divided into a group to obtain multiple groups of activation channels, so as to calculate the quantitative parameters of each group of activation channels.
[0035] In order to reduce the quantization error caused by the range difference between channels, in this embodiment, the calibration data set is first used to determine the maximum and minimum values of each activated channel, and the activated channels with activation value ranges in a preset similar range are divided into a group by a clustering method to obtain multiple groups of activated channels. Then, the quantization parameters (scaling factors and zero points) are calculated independently for each group to ensure that channels with similar ranges share the same quantization parameters, reducing the quantization error introduced by the range difference between channels. The overall accuracy of the quantization process can be significantly improved by the preset channel grouping rearrangement algorithm.
[0036] The expressions related to the channel grouping rearrangement algorithm are: set up ,in B is the number of calibration samples, C is the number of channels, N is the number of markers. First, the maximum and minimum values of each channel are calculated using the calibration data set, which are expressed as and .
[0037] Then, according to each channel The channels are grouped into g Clusters. S 1 , S 2 , ..., S g is the set of channel indices in each cluster, where and Finally, reorder the channels in X according to the index. Concatenate all the indices into a vector . Get the reordered activation vector .
[0038] The channel smoothing algorithm and the channel grouping rearrangement algorithm are orthogonal, and the combination of the two will bring better quantization accuracy.
[0039] As a preferred embodiment, it also includes: 105: In the nonlinear layer inference process of the large language model, exponential calculation is replaced by integer polynomial calculation to perform full quantization inference of the large language model.
[0040] In order to meet the requirements of the end-side devices for low power consumption and high efficiency, in this embodiment, in the Softmax function inference process of the large language model, the Softmax quantization method based on Log2 is used to replace the traditional exponential calculation with the bit shift operation to achieve efficient integer domain Softmax calculation, and the maximum activation value is subtracted by integer operation to prevent overflow problems and reduce computing costs, thereby significantly reducing computing and storage overheads while maintaining accuracy.
[0041] Specifically, considering the exponential function used in Softmax e x , this function has no boundaries and can take very large values. Using a polynomial to approximate a function with a large range will result in a very large error. Because first convert Softmax to: For each x i Subtract a maximum value x max , in this way the size of the exp function value is indirectly reduced.
[0042] is a non-positive number . Any non-positive number can be expressed as: in, z is an integer, , p is a value in the range Real numbers between .
[0043] In fact, a negative number is decomposed into -ln2, and the remainder is p Added representation.
[0044] Calculating the exp function, we can get: It can be seen that 2 -z It can be quickly implemented in the form of bit shifting.
[0045] exp can be calculated using polynomials, and Softmax can also be calculated using polynomials, which can then be used to quantify reasoning.
[0046] Of course, the nonlinear layer of the present invention may also be a quantization of LayerNorm, and the present invention does not make any special limitation here.
[0047] The method of the present invention first performs channel smoothing, then performs channel grouping rearrangement quantization, and finally performs nonlinear layer quantization. This order can minimize quantization errors and optimize hardware adaptability, ensuring that the accuracy of the model remains close to the original floating point accuracy under compressed 4-bit accuracy (W4A4), providing a solution for large language models that is efficient in model compression and easy to deploy on accelerators, saving computing and communication overhead.
[0048] The present invention has achieved remarkable results in the practical application of multiple large language models. Models such as OPT-125m, OPT-1.3b and OPT-6.7b were evaluated on WIKItext2, PTB and C4 datasets. The results show that compared with the unquantized full-precision model, the method of the present invention achieves an 8-fold model compression rate, while the model perplexity (PPL) performance at W4A4 accuracy is better than other existing methods. The present invention is expected to achieve significant energy consumption reduction at the same compression ratio, and efficient reasoning can be achieved on edge devices.
[0049] The method of the present invention can be deployed on various end-side devices through software. The end-side device involved in the present invention may refer to a device with a wireless connection function. The wireless connection function refers to the ability to connect to other terminal devices through wireless connection methods such as WiFi and Bluetooth. The terminal device may also have the function of communicating through a wired connection. The end-side device of the present invention may be a touch screen, a non-touch screen, or a screenless device. A touch screen device can control the terminal device by clicking and sliding on the display screen with a finger or a stylus; a non-touch screen device can be connected to a mouse, keyboard, touch panel, etc. to control the terminal device through an input device; a device without a screen, for example, may be a Bluetooth speaker without a screen. For example, the end-side device of the present invention may be a computer, a smart phone, a netbook, a tablet computer, a wearable electronic device (such as a smart bracelet, a smart watch, etc.), a television (TV), a virtual reality device, etc.
[0050] The method of the present invention can also be deployed on a server, which can be located in the cloud or locally, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., with a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. It can refer to a device with a wireless connection function, and the wireless connection function means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The server of the present invention can also have the function of communicating through a wired connection. For example, the server of the present invention can be located in the cloud and communicate with the terminal device. The server receives the large language model to be quantized sent by the terminal device, and quantizes the large language model using the quantization method deployed on the server, obtains the quantized large language model and returns it to the terminal device; or, the server can receive input data sent by the terminal device, process the input data using the inference method deployed on the server, obtain the processing result of the input data and return it to the terminal device.
[0051] The post-quantization system for terminal-side deployment of a large language model provided by the present invention is described below. The post-quantization system for terminal-side deployment of a large language model described below and the post-quantization method for terminal-side deployment of a large language model described above can refer to each other.
[0052] Please refer to Figure 3 , Figure 3 One of the structural schematic diagrams of a post-quantization system for terminal-side deployment of a large language model provided by the present invention.
[0053] The present invention also provides a post-quantization system for terminal-side deployment of a large language model, including: a determination module 1, used to determine a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; a smoothing module 2, used to perform channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculate the quantization parameters of the smoothed channel; a quantization module 3, used to perform channel quantization according to the quantization parameters of the smoothed channel to obtain a quantized matrix; the quantized matrix is used to be deployed on the terminal side to run the quantized large language model on the terminal side.
[0054] Please refer to Figure 4 , Figure 4 The second structural diagram of a post-quantization system for terminal-side deployment of a large language model provided by the present invention.
[0055] As a preferred embodiment, it also includes: a grouping rearrangement module 4, which is used to perform channel grouping rearrangement according to the channel value range and based on a preset channel grouping rearrangement algorithm, and calculate the quantization parameter of each grouping channel; the quantization parameter of the grouping channel is used to determine the quantized matrix.
[0056] As a preferred embodiment, it also includes: a replacement module 5, which is used to replace the exponential calculation with integer polynomial calculation in the nonlinear layer reasoning process of the large language model to perform full quantization reasoning of the large language model.
[0057] Figure 5 An example of a structural diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor (processor) 501, a communication interface (Communications Interface) 502, a memory (memory) 503 and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504. The processor 501 can call the logic instructions in the memory 503 to execute a post-quantization method for terminal-side deployment of a large language model, the method including: determining a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating the quantization parameter of the smoothed channel; performing channel quantization based on the quantization parameter of the smoothed channel to obtain a quantized matrix; the quantized matrix is used for deployment on the terminal side to run the quantized large language model on the terminal side; also including: performing channel grouping rearrangement based on a preset channel grouping rearrangement algorithm according to the channel value range, and calculating the quantization parameter of each grouped channel; the quantization parameter of the grouped channel is used to determine the quantized matrix; also including: in the nonlinear layer inference process of the large language model, replacing the exponential calculation with integer polynomial calculation to perform full quantization inference of the large language model.
[0058] In addition, the logic instructions in the above-mentioned memory 503 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0059] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the post-quantization method for the terminal side deployment of a large language model provided by the above methods, the method including: determining a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating the quantization parameter of the smoothed channel; performing channel quantization based on the quantization parameter of the smoothed channel to obtain a quantized matrix; the quantized matrix is used for deployment on the terminal side to run the quantized large language model on the terminal side; it also includes: performing channel grouping rearrangement based on a preset channel grouping rearrangement algorithm according to the channel value range, and calculating the quantization parameter of each grouped channel; the quantization parameter of the grouped channel is used to determine the quantized matrix; it also includes: in the nonlinear layer inference process of the large language model, replacing the exponential calculation with integer polynomial calculation to perform full quantization inference of the large language model.
[0060] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the post-quantization method for the terminal-side deployment of a large language model provided by the above methods, the method comprising: determining a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating the quantization parameter of the smoothed channel; performing channel quantization based on the quantization parameter of the smoothed channel to obtain a quantized matrix; the quantized matrix is used for deployment on the terminal side to run the quantized large language model on the terminal side; further comprising: performing channel grouping rearrangement based on a preset channel grouping rearrangement algorithm according to the channel value range, and calculating the quantization parameter of each grouped channel; the quantization parameter of the grouped channel is used to determine the quantized matrix; further comprising: replacing exponential calculation with integer polynomial calculation in the nonlinear layer inference process of the large language model to perform full quantization inference of the large language model.
[0061] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0062] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A post-quantization method for client-side deployment of a large language model, characterized in that: include: Determine a channel value range corresponding to an input channel of any current matrix to be quantized of the large language model; According to the channel value range, channel smoothing is performed based on a preset channel smoothing algorithm, and a quantization parameter of the smoothed channel is calculated; Channel quantization is performed according to the quantization parameter of the smoothed channel to obtain a quantized matrix; the quantized matrix is used for deployment on the end side to run the quantized large language model on the end side.
2. The post-quantization method for large language model terminal deployment according to claim 1, characterized in that: Also includes: According to the channel value range, the channels are grouped and rearranged based on a preset channel grouping and rearrangement algorithm, and a quantization parameter of each grouped channel is calculated; The quantization parameter of the grouped channels is used to determine the quantized matrix.
3. The post-quantization method for large language model terminal deployment according to claim 1, characterized in that: Also includes: In the nonlinear layer inference process of the large language model, exponential calculation is replaced by integer polynomial calculation to perform full quantization inference of the large language model.
4. The post-quantization method for large language model terminal deployment according to claim 1, characterized in that: The channel value range includes an activation value range and a weight value range; The performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range includes: Calculating a channel smoothing factor according to the activation value range and the weight value range; the channel smoothing factor is used to characterize the data distribution in the channel so as to transfer the difficulty of quantizing the activation value to the weight value; Channel smoothing is performed according to the channel smoothing factor.
5. The post-quantization method for terminal-side deployment of a large language model according to any one of claims 1 to 4, characterized in that: The performing channel grouping rearrangement according to the channel value range and based on a preset channel grouping rearrangement algorithm includes: According to the activation value range of each activation channel, the activation channels with activation value ranges in a preset similar range are divided into a group based on a clustering algorithm to obtain multiple groups of activation channels, so as to calculate the quantization parameters for each group of activation channels.
6. A post-quantization system for large language model client-side deployment, characterized in that: include: A determination module, used to determine a channel value range corresponding to an input channel of any current matrix to be quantized of a large language model; A smoothing module, used for performing channel smoothing based on a preset channel smoothing algorithm according to the channel value range, and calculating the quantization parameter of the smoothed channel; A quantization module is used to perform channel quantization according to the quantization parameter of the smoothed channel to obtain a quantized matrix; the quantized matrix is used to be deployed on the end side to run the quantized large language model on the end side.
7. The post-quantization system for large language model terminal deployment according to claim 6, characterized in that: Also includes: A grouping rearrangement module, used to perform channel grouping rearrangement according to the channel value range and based on a preset channel grouping rearrangement algorithm, and calculate the quantization parameter of each grouping channel; The quantization parameter of the grouped channels is used to determine the quantized matrix.
8. The post-quantization system for large language model terminal deployment according to claim 6 or 7, characterized in that: Also includes: A replacement module is used to replace exponential calculation with integer polynomial calculation in the nonlinear layer reasoning process of the large language model to perform full quantization reasoning of the large language model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the post-quantization method for terminal-side deployment of a large language model is implemented as described in any one of claims 1 to 5.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the post-quantization method for terminal-side deployment of a large language model is implemented as described in any one of claims 1 to 5.