Large model quantification method based on convergence response and related system

Through the method based on convergence response, the weight compression and quantification of the large model is solved, and the application challenges of the large model in resource-constrained scenarios are realized, and the lightweight deployment and efficient operation of the model are achieved.

CN119940573APending Publication Date: 2025-05-06CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510026263.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In scenarios where large models are subject to resource limitations, high computational complexity, long inference and high hardware resource requirements, making it difficult to achieve lightweight deployment.

Method used

Through a method based on convergence response, each layer and channel of the large model is compressed and quantified, and the quantization parameters with the smallest loss are obtained to realize the quantitative deployment of the model.

Benefits of technology

It significantly reduces the storage requirements and computing burden of the model, improves the adaptability and operation efficiency of the model on resource-constrained devices, and ensures that the key performance of the model remains at a high level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940573A_ABST
    Figure CN119940573A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and discloses a large model quantification method based on convergence response and a related system, which can significantly reduce the storage demand and calculation burden of a model by performing weight compression and quantification on each layer and channel of the large model. Therefore, the model which originally needs a large amount of computing resources can be converted into a smaller model which adapts to resource-constrained equipment (such as edge equipment or mobile equipment), and the dependence of the model on hardware resources is reduced. The compressed model generally contains fewer parameters and operations, and the reasoning process becomes more efficient, so that the response speed of the model is accelerated. For tasks applied in real-time scenes such as an end scene and an edge scene, the speed improvement is crucial. The parameter precision of the quantized model is generally reduced, so that the storage space required by the model is greatly reduced. In equipment with limited resources, the memory and the storage space can be effectively saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a large model quantization method based on convergent response and a related system. Background Art

[0002] In recent years, large models (LLMs) have made significant progress in the field of natural language processing (NLP). Since OpenAI released GPT-3, ultra-large-scale pre-trained models have gradually become the mainstream direction of artificial intelligence research. Large models rely on their large parameter scale to give them powerful learning capabilities. The application of the Transformer mechanism makes it possible to conduct unsupervised training based on massive data. With the support of big data, LLMs can not only learn rich knowledge, but also further explore the potential logic and patterns between data, making them have extremely strong generalization capabilities. However, large models have high computational complexity, long reasoning time, and high hardware resource requirements when used in reasoning applications. In scenarios with limited hardware resources such as distribution networks, safety supervision, and mobile, there are great challenges in deploying and applying large models. In order to reduce the computational complexity and hardware resource requirements of large models and realize lightweight deployment of large models, large model compression and optimization technology has become an important research direction at present. Specifically, the existing mainstream compression methods for large models include knowledge distillation, model pruning, and model quantization. Quantization can well balance the compression ratio and model performance, has strong generalization ability, and is widely used in practice. Algorithms such as GPTQ and AWQ can achieve a 60-70% reduction in video memory and time consumption with a slight decrease in accuracy (<5%).

[0003] Although algorithms such as GPTQ and AWQ have achieved remarkable results in compressing models and maintaining model performance, the requirements for computing and storage of quantization models in actual deployment are still very high. At the same time, according to relevant research such as LoRA and AWQ, there are many redundant parameters in the model, and the weights that significantly affect the model effect only account for a small part. Summary of the invention

[0004] The purpose of the present invention is to overcome the application limitations of the above-mentioned large model in resource-constrained scenarios such as terminals and edges, and to provide a large model quantization method based on convergent response and a related system.

[0005] In order to achieve the above object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a large model quantification method based on convergent response, comprising the following steps: Obtain the data set that needs to be calibrated in the large language model to form a calibration data set; The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model. The weight response convergence and weight importance of each layer and channel in the full model are determined based on the response characteristics. According to the weight response convergence and weight importance of each layer and channel in the full model, the weights of each layer and channel in the full model are compressed, and a quantized full model is obtained according to the compressed weights of each layer and channel in the full model. The quantized full model is iterated to obtain the quantization parameters with the minimum loss. Deploy the quantization model based on the quantization parameters with the smallest loss.

[0006] A further improvement of the present invention is that the specific method for obtaining the data set that needs to be calibrated in the large language model is as follows: Select the documents that need to be calibrated according to business needs; Parse the document and obtain the content of the document; De-duplicate and desensitize the content of the document, delete low-quality data, and obtain cleaned data; According to business needs, the data that needs to be calibrated is selected from the cleaned data to form a calibration data set.

[0007] A further improvement of the present invention is that the full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model. The specific method for determining the weight response convergence and weight importance of each layer and channel in the full model according to the response characteristics is as follows: The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model; According to the response characteristics of each layer and channel in the full model, the response characteristics of each layer and channel in the full model are spliced; According to the response characteristics after splicing, the convergence of each layer and channel is calculated; According to the convergence of each layer and channel, the average response amplitude of each layer and channel is inferred by weight, so as to obtain the weight response convergence and weight importance of each layer and channel in the full model.

[0008] A further improvement of the present invention is that, according to the response characteristics after splicing, the specific method for calculating the convergence of each layer and channel is as follows:

[0009]

[0010] in, For channels in the same layer and Channel The convergence of For Channel The response distribution of For Channel The response distribution of , so The threshold of is (0,1], For Channel The convergence of for The number of channels of the layer.

[0011] A further improvement of the present invention is that, according to the convergence of each layer and channel, a method of using weighted reasoning to calculate the average response amplitude of each layer and channel is as follows:

[0012] in, For Channel The average response amplitude, To calibrate the number of test set samples, For Channel In the sample The response amplitude on .

[0013] A further improvement of the present invention is that, according to the weight response convergence and weight importance of each layer and channel in the full model, the weights of each layer and channel in the full model are compressed, and a quantized full model is obtained according to the compressed weights of each layer and channel in the full model. The quantized full model is iterated, and the specific method of obtaining the quantization parameter with the minimum loss is as follows: Obtain the weight response convergence of each layer and channel in the full model, compare the weight response convergence with the preset threshold, select the weight that can be mapped according to the comparison result and the convergence multiple, compress the weight, and obtain the compressed full model; According to the weight importance of each layer and channel in the full model, the compressed full model is quantized to obtain a quantized full model; The quantized full-volume model is used to infer the loss parameters, and then the loss parameters are sent to the quantized full-volume model for iteration until the quantization parameters with the smallest loss are obtained while the compression ratio meets the requirements.

[0014] A further improvement of the present invention is that, according to the quantization parameter with the smallest loss, the specific method of deploying the quantization model is as follows: According to the weight importance of each layer and channel in the full model, the weight of each layer and channel in the full model is obtained; The weights of each layer and channel in the full model are divided by the quantization scaling factor, and then multiplied by the quantization parameter with the smallest loss to obtain the deployed quantization model.

[0015] In a second aspect, the present invention provides a large model quantization system based on convergent response, comprising: The data calibration module is used to obtain the data set that needs to be calibrated in the large language model to form a calibration data set; The weight determination module is used to use the full model to infer the calibration data set, obtain the response characteristics of each layer and channel in the full model, and determine the weight response convergence and weight importance of each layer and channel in the full model according to the response characteristics; The model iteration module is used to compress the weights of each layer and channel in the full model according to the weight response convergence and weight importance of each layer and channel in the full model, obtain the quantized full model according to the compressed weights of each layer and channel in the full model, and iterate the quantized full model to obtain the quantization parameters with the minimum loss; The model deployment module is used to deploy the quantization model according to the quantization parameters with the minimum loss.

[0016] A further improvement of the present invention is that the function of the data calibration module is realized by the following method: Select the documents that need to be calibrated according to business needs; Parse the document and obtain the content of the document; De-duplicate and desensitize the content of the document, delete low-quality data, and obtain cleaned data; According to business needs, the data that needs to be calibrated is selected from the cleaned data to form a calibration data set.

[0017] A further improvement of the present invention is that the function of the weight determination module is realized by the following method: The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model; According to the response characteristics of each layer and channel in the full model, the response characteristics of each layer and channel in the full model are spliced; According to the response characteristics after splicing, the convergence of each layer and channel is calculated; According to the convergence of each layer and channel, the average response amplitude of each layer and channel is inferred by weight, so as to obtain the weight response convergence and weight importance of each layer and channel in the full model.

[0018] A further improvement of the present invention is that the function of the model iteration module is implemented by the following method: Obtain the weight response convergence of each layer and channel in the full model, compare the weight response convergence with the preset threshold, select the weight that can be mapped according to the comparison result and the convergence multiple, compress the weight, and obtain the compressed full model; According to the weight importance of each layer and channel in the full model, the compressed full model is quantized to obtain a quantized full model; The quantized full-volume model is used to infer the loss parameters, and then the loss parameters are sent to the quantized full-volume model for iteration until the quantization parameters with the smallest loss are obtained while the compression ratio meets the requirements.

[0019] A further improvement of the present invention is that the function of the model deployment module is implemented by the following method: According to the weight importance of each layer and channel in the full model, the weight of each layer and channel in the full model is obtained; The weights of each layer and channel in the full model are divided by the quantization scaling factor, and then multiplied by the quantization parameter with the smallest loss to obtain the deployed quantization model.

[0020] In a third aspect, the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of a large model quantization method based on convergent response when executing the computer program.

[0021] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of a large model quantization method based on convergent response are implemented.

[0022] Compared with the prior art, the present invention has the following beneficial effects: The present invention can significantly reduce the storage requirements and computational burden of the model by compressing and quantizing the weights of each layer and channel of the large model. This helps to transform a model that originally requires a large amount of computing resources into a smaller model that adapts to resource-constrained devices (such as edge devices or mobile devices), reducing its dependence on hardware resources. The compressed model usually contains fewer parameters and operations, and the reasoning process becomes more efficient, thereby accelerating the response speed of the model. For tasks applied in real-time scenarios such as end and edge, this speed improvement is crucial. The parameter accuracy of the quantized model is usually reduced, which greatly reduces the storage space required for the model. In devices with limited resources, this can effectively save memory and storage space. By analyzing the convergence and importance of the weight response of each layer and channel, the present invention can finely compress the model, retain the most important features, and discard the redundant parts. This ensures that even after quantization and compression, the key performance of the model remains at a high level. Quantized models and compressed computing operations usually require less computing and storage bandwidth, which not only improves performance, but also helps to reduce the energy consumption of hardware devices and adapt to applications on edge devices with limited power consumption. The present invention can adjust the model to adapt to different hardware and application scenarios while maintaining minimum loss by iteratively optimizing quantization parameters. This adaptability enables the model to run effectively on a variety of edge devices, avoiding the need to retrain each device. The quantized model is easier to deploy on different hardware platforms and environments, especially in edge computing scenarios with limited resources. For example, mobile devices, IoT devices, etc. can effectively run large models through this quantization method, thereby improving the universality of the application. In summary, the present invention effectively reduces resource consumption and improves operating efficiency while ensuring model performance through a sophisticated compression and quantization mechanism, and can flexibly adapt to computing scenarios with limited resources such as ends and edges. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flow chart of the present invention; Figure 2 is a system diagram of the present invention; Figure 3 This is a system diagram of Example 5 of the present invention. DETAILED DESCRIPTION

[0024] In order to further understand the content of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and are not intended to limit it.

[0025] See also Figure 1 , a large model quantification method based on convergent response, comprising the following steps: S1, obtain the data set that needs to be calibrated in the large language model to form a calibration data set.

[0026] S2, use the full model to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model, and determine the weight response convergence and weight importance of each layer and channel in the full model based on the response characteristics.

[0027] S3, according to the weight response convergence and weight importance of each layer and channel in the full model, compress the weights of each layer and channel in the full model, obtain the quantized full model based on the compressed weights of each layer and channel in the full model, iterate the quantized full model, and obtain the quantization parameters with the smallest loss.

[0028] S4, deploy the quantization model according to the quantization parameter with the minimum loss.

[0029] See also Figure 2 , a large model quantization system based on convergent response, comprising: The data calibration module is used to obtain the data set that needs to be calibrated in the large language model to form a calibration data set; The weight determination module is used to use the full model to infer the calibration data set, obtain the response characteristics of each layer and channel in the full model, and determine the weight response convergence and weight importance of each layer and channel in the full model according to the response characteristics; The model iteration module is used to compress the weights of each layer and channel in the full model according to the weight response convergence and weight importance of each layer and channel in the full model, obtain the quantized full model according to the compressed weights of each layer and channel in the full model, and iterate the quantized full model to obtain the quantization parameters with the minimum loss; The model deployment module is used to deploy the quantization model according to the quantization parameters with the minimum loss.

[0030] Embodiment 1: In S1, the data set that needs to be calibrated in the large language model is obtained. The specific method of forming the calibration data set is as follows: S11, data selection: select the documents to be calibrated according to the model modality, application scenario and generalization requirements. For example, the directional vision model only needs to select the pictures in the application scenario. The large language model selects the scope of its data set according to its generalization scenario, such as language understanding, summary generation, knowledge reasoning, specific business data, etc. For common sense and general knowledge, open source data can be directly selected, while specific business data needs further processing; S12, data analysis: parse word, pdf, excel and other documents that need to be calibrated through libraries such as pandas, pandoc, python-docx, reportlab, etc. to obtain the document content; S13, data cleaning: The main work is to remove duplicate data (document level, paragraph level, sentence level, etc.), desensitize (confidential data, privacy data, etc.), remove low-quality data (incomplete content, meaningless data, etc.), and complete the cleaning through text quality recognition model and rule matching to obtain cleaned data; S14, data labeling: formulate relevant labeling specifications according to specific business needs, label the collected data, and form a calibrated data set after verification; Embodiment 2: In S2, the full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model. The specific method for determining the weight response convergence and weight importance of each layer and channel in the full model based on the response characteristics is as follows: S21, use the full model to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model; S22, storing the dimensions of each layer and channel in the full model, and splicing them according to the response characteristics of the dimensions of each layer and channel in the full model to different data; S23, according to the response characteristics after splicing, calculate the convergence of each layer and channel. The convergence is determined based on the KL divergence (Kullback-Leibler Divergence) of different channels after splicing. The specific method is as follows:

[0031]

[0032] in, For channels in the same layer and Channel The convergence of For Channel The response distribution of For Channel The response distribution of , so The threshold of is (0,1], For Channel The convergence of for The number of channels in the layer, The larger the value of , the greater the convergence.

[0033] S24, according to the convergence of each layer and channel, the average response amplitude of each layer and channel is inferred by weight, so as to obtain the weight response convergence and weight importance of each layer and channel in the full model. The specific method is as follows:

[0034] in, For Channel The average response amplitude, To calibrate the number of test set samples, For Channel In the sample The response amplitude on .

[0035] Embodiment 3: In S3, according to the weight response convergence and weight importance of each layer and channel in the full model, the weights of each layer and channel in the full model are compressed, and the quantized full model is obtained according to the compressed weights of each layer and channel in the full model. The quantized full model is iterated to obtain the quantization parameters with the minimum loss. The specific method is as follows: S31, obtaining the weight response convergence of each layer and channel in the full model, comparing the weight response convergence with a preset threshold, selecting the weights that can be mapped according to the comparison results and the convergence multiples, compressing the weights, and obtaining the compressed full model. The method for selecting the weights that can be mapped according to the comparison results and the convergence multiples is as follows:

[0036] in, For Channel i The weight after mapping, is the original weight, For Channel i The convergence multiple of i The number of channels whose convergence is greater than the threshold, K is the threshold of mapping convergence multiples, The original weight is mapped to the channel weight with the highest average convergence with other channels in the aggregated cluster weight. The mapped weight does not participate in subsequent activation and calculation, and the response is directly reused result.

[0037] S32, according to the weight importance of each layer and channel in the full model, the compressed full model is quantized to obtain a quantized full model. The specific method is as follows: For the full model after convergence compression, the full model is adaptively quantified according to its importance, and the weight of the top 1% or 0.1% with higher importance is obtained by high-precision quantization. , others use baseline quantization to obtain lower precision weights , the specific method is as follows:

[0038] in, is the quantized weight, is the baseline quantization method (Round-to-Nearest), is the weight of the full model after convergence compression, s The quantization scaling factor is obtained by calculating the quantization loss on the calibration test set.

[0039] S33, using the full model after response convergence compression and importance adaptive quantization to calculate the loss of full-precision inference, and continuously iterating the parameters of convergence compression and importance adaptive quantization according to the loss until the quantization parameters with the minimum quantization loss are obtained while the compression ratio meets the requirements.

[0040] Embodiment 4: In S4, according to the quantization parameter with the smallest loss, the specific method of deploying the quantization model is as follows: According to the weight importance of each layer and channel in the full model, the weight of each layer and channel in the full model is obtained; Divide the weights of each layer and channel in the full model by the quantization scaling factor s , and then multiply it with the quantization parameter with the smallest loss to get the deployment quantization model, which can avoid the compatibility issues caused by mixed precision. The specific method is as follows:

[0041] in, is the output of the quantization model, To quantify the importance of weights in the model.

[0042] The present invention compresses the redundant parameters of the model based on the response convergence, and directly uses the response mapping method for reasoning during reasoning, which can not only effectively reduce the video memory occupation of model reasoning, but also avoid the reasoning calculation of redundant parameters and improve the reasoning speed. The present invention adopts a convergence compression strategy, combines the importance of weights, adopts a high-precision quantization method for the top 1% important weights, and adopts a more efficient benchmark quantization method for low importance, which can effectively balance the model prediction effect and compression ratio, and highly restore the prediction effect of the full-precision model while achieving efficient compression. The present invention adopts a non-training PTQ (Post-Training Quantization) strategy, which will not change the weights of the original model, avoids the overfitting problem of the quantization model for a specific data set, and helps to enhance the generalization ability of the quantization model in different tasks and fields. The present invention is mainly applied to large models based on the transformer mechanism, but the proposed convergence response compression strategy and the quantization method combined with the importance of weights do not depend on the specific model structure, and can be used for networks with traditional architectures such as CNN and LSTM, and have a wide range of applicability. The model quantization technology based on convergence response has important practical application value. Through theoretical analysis and experimental verification, the present invention explores the technical principle of convergent response compression combined with a quantization method of weight importance, and conducts actual verification, and the method is effective and practical.

[0043] Embodiment 5: See also Figure 3 As shown, the present invention also provides an electronic device 100 for a large model quantization method based on convergent response; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0044] The memory 101 can be used to store the computer program 103, and the processor 102 implements the steps of a large model quantification method based on convergence response described in Example 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data (such as audio data) created according to the use of the electronic device 100, etc. In addition, the memory 101 can include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0045] The at least one processor 102 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and uses various interfaces and lines to connect various parts of the entire electronic device 100.

[0046] The memory 101 in the electronic device 100 stores a plurality of instructions to implement a large model quantization method based on convergence response, and the processor 102 can execute the plurality of instructions to implement: Obtain the data set that needs to be calibrated in the large language model to form a calibration data set; The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model. The weight response convergence and weight importance of each layer and channel in the full model are determined based on the response characteristics. According to the weight response convergence and weight importance of each layer and channel in the full model, the weights of each layer and channel in the full model are compressed, and a quantized full model is obtained according to the compressed weights of each layer and channel in the full model. The quantized full model is iterated to obtain the quantization parameters with the minimum loss. Deploy the quantization model based on the quantization parameters with the smallest loss.

[0047] Embodiment 6: If the module / unit integrated in the electronic device 100 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory).

[0048] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0049] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0050] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A large model quantification method based on convergent response, characterized in that: The following steps are involved: Obtain the data set that needs to be calibrated in the large language model to form a calibration data set; The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model. The weight response convergence and weight importance of each layer and channel in the full model are determined based on the response characteristics. According to the weight response convergence and weight importance of each layer and channel in the full model, the weights of each layer and channel in the full model are compressed, and a quantized full model is obtained according to the compressed weights of each layer and channel in the full model. The quantized full model is iterated to obtain the quantization parameters with the minimum loss. Deploy the quantization model based on the quantization parameters with the smallest loss.

2. A large model quantification method based on convergent response according to claim 1, characterized in that: The specific method for obtaining the data set that needs to be calibrated in the large language model is as follows: Select the documents that need to be calibrated according to business needs; Parse the document and obtain the content of the document; De-duplicate and desensitize the content of the document, delete low-quality data, and obtain cleaned data; According to business needs, the data that needs to be calibrated is selected from the cleaned data to form a calibration data set.

3. The large model quantification method based on convergent response according to claim 1, characterized in that: The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model. The specific method for determining the weight response convergence and weight importance of each layer and channel in the full model based on the response characteristics is as follows: The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model; According to the response characteristics of each layer and channel in the full model, the response characteristics of each layer and channel in the full model are spliced; According to the response characteristics after splicing, the convergence of each layer and channel is calculated; According to the convergence of each layer and channel, the average response amplitude of each layer and channel is inferred by weight, so as to obtain the weight response convergence and weight importance of each layer and channel in the full model.

4. The large model quantification method based on convergent response according to claim 3 is characterized in that: According to the response characteristics after splicing, the specific method for calculating the convergence of each layer and channel is as follows: in, For channels in the same layer and Channel The convergence of For Channel The response distribution of For Channel The response distribution of , so The threshold of is (0,1], For Channel The convergence of for The number of channels of the layer.

5. The large model quantification method based on convergent response according to claim 3 is characterized in that: According to the convergence of each layer and channel, the method of using weighted inference to calculate the average response amplitude of each layer and channel is as follows: in, For Channel The average response amplitude, To calibrate the number of test set samples, For Channel In the sample The response amplitude on .

6. The large model quantification method based on convergent response according to claim 1, characterized in that: According to the weight response convergence and weight importance of each layer and channel in the full model, the weights of each layer and channel in the full model are compressed, and the quantized full model is obtained according to the compressed weights of each layer and channel in the full model. The quantized full model is iterated to obtain the quantization parameters with the smallest loss. The specific method is as follows: Obtain the weight response convergence of each layer and channel in the full model, compare the weight response convergence with the preset threshold, select the weight that can be mapped according to the comparison result and the convergence multiple, compress the weight, and obtain the compressed full model; According to the weight importance of each layer and channel in the full model, the compressed full model is quantized to obtain a quantized full model; The quantized full-volume model is used to infer the loss parameters, and then the loss parameters are sent to the quantized full-volume model for iteration until the quantization parameters with the smallest loss are obtained while the compression ratio meets the requirements.

7. The large model quantification method based on convergent response according to claim 1, characterized in that: According to the quantization parameters with the smallest loss, the specific method of deploying the quantization model is as follows: According to the weight importance of each layer and channel in the full model, the weight of each layer and channel in the full model is obtained; The weights of each layer and channel in the full model are divided by the quantization scaling factor, and then multiplied by the quantization parameter with the smallest loss to obtain the deployed quantization model.

8. A large model quantization system based on convergent response, characterized in that: include: The data calibration module is used to obtain the data set that needs to be calibrated in the large language model to form a calibration data set; The weight determination module is used to use the full model to infer the calibration data set, obtain the response characteristics of each layer and channel in the full model, and determine the weight response convergence and weight importance of each layer and channel in the full model according to the response characteristics; The model iteration module is used to compress the weights of each layer and channel in the full model according to the weight response convergence and weight importance of each layer and channel in the full model, obtain the quantized full model according to the compressed weights of each layer and channel in the full model, and iterate the quantized full model to obtain the quantization parameters with the minimum loss; The model deployment module is used to deploy the quantization model according to the quantization parameters with the minimum loss.

9. A large model quantization system based on convergent response according to claim 8, characterized in that: The functions of the data calibration module are realized by the following methods: Select the documents that need to be calibrated according to business needs; Parse the document and obtain the content of the document; De-duplicate and desensitize the content of the document, delete low-quality data, and obtain cleaned data; According to business needs, the data that needs to be calibrated is selected from the cleaned data to form a calibration data set.

10. The large model quantization system based on convergent response according to claim 8, characterized in that: The function of the weight determination module is realized by the following methods: The full model is used to infer the calibration data set to obtain the response characteristics of each layer and channel in the full model; According to the response characteristics of each layer and channel in the full model, the response characteristics of each layer and channel in the full model are spliced; According to the response characteristics after splicing, the convergence of each layer and channel is calculated; According to the convergence of each layer and channel, the average response amplitude of each layer and channel is inferred by weight, so as to obtain the weight response convergence and weight importance of each layer and channel in the full model.

11. The large model quantization system based on convergent response according to claim 8, characterized in that: The functions of the model iteration module are implemented through the following methods: Obtain the weight response convergence of each layer and channel in the full model, compare the weight response convergence with the preset threshold, select the weight that can be mapped according to the comparison result and the convergence multiple, compress the weight, and obtain the compressed full model; According to the weight importance of each layer and channel in the full model, the compressed full model is quantized to obtain a quantized full model; The quantized full-volume model is used to infer the loss parameters, and then the loss parameters are sent to the quantized full-volume model for iteration until the quantization parameters with the smallest loss are obtained while the compression ratio meets the requirements.

12. The large model quantization system based on convergent response according to claim 8, characterized in that: The functionality of the model deployment module is implemented through the following methods: According to the weight importance of each layer and channel in the full model, the weight of each layer and channel in the full model is obtained; The weights of each layer and channel in the full model are divided by the quantization scaling factor, and then multiplied by the quantization parameter with the smallest loss to obtain the deployed quantization model.

13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a large model quantification method based on convergent response according to any one of claims 1 to 7 are implemented.

14. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a large model quantification method based on convergent response according to any one of claims 1 to 7 are implemented.