Quantitative quality evaluation method and device for large language model
By evaluating the accuracy, compression rate and memory usage of the quantitative large language model, the problem of performance and storage imbalance in the quantized model is solved, ensuring the stability and efficiency of the model in different environments.
Patent Information
- Application Number
- CN202510313597.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
In the process of quantifying large language models, the prior art cannot ensure that the quantized model achieves a balance in performance and storage, and cannot ensure that the model operates stably under different environments.
By evaluating the accuracy coefficient, compression coefficient and memory occupancy coefficient of the quantitative large language model, comprehensively considering the inference accuracy, model size and memory occupancy, a quantitative quality evaluation method for large language model is provided.
The quantitative model is balanced in performance and storage, and ensures that the model operates stably in different environments, providing a comprehensive evaluation of the quantitative quality of the model.
Smart Images

Figure CN120258134A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method and device for evaluating the quantization quality of large language models. Background Art
[0002] Model quantization is a method that reduces the computational resource requirements while maintaining the inference accuracy of the model. It is widely used in deep learning inference tasks, especially in edge computing, mobile devices, and large-scale distributed inference systems. The basic idea of quantization is to use lower numerical representations (such as int8, int4, or lower) to replace traditional FP32 or FP16 floating-point calculations, so as to reduce the storage occupancy and computational complexity, while minimizing the accuracy loss as much as possible.
[0003] In the process of model quantization of large models, although quantization, as a method to reduce computational complexity and storage requirements, can improve the usability of the model when hardware resources are limited, it will inevitably bring information loss. Therefore, how to evaluate whether the quantized model can still meet the expected task requirements and whether a reasonable balance has been achieved between computational efficiency and model accuracy has become a crucial issue.
[0004] Currently, in the process of quantization, usually only a single metric is considered, and the quality of model quantization varies greatly. It is impossible to ensure the balance between the performance and storage of the quantized model, nor to ensure the stable operation of the model in different environments. Summary of the Invention
[0005] Embodiments of this application provide a method and device for evaluating the quantization quality of large language models, which are used to solve the problem that the quality of model quantization varies greatly, and it is impossible to ensure the balance between the performance and storage of the quantized model and also impossible to ensure the stable operation of the model in different environments.
[0006] Embodiments of this application adopt the following technical solutions:
[0007] On the one hand, embodiments of this application provide a method for evaluating the quantization quality of large language models. The method includes: determining the accuracy coefficient of the quantized large language model according to the accuracy of the original large language model and the accuracy of the quantized large language model; determining the compression coefficient of the quantized large language model according to the file size of the original large language model and the file size of the quantized large language model; determining the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy size of the original large language model and the video memory occupancy size of the quantized large language model; and evaluating the quantized large language model according to the accuracy coefficient, the compression coefficient, and the video memory occupancy coefficient.
[0008] In one example, before determining the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy of the original large language model and the quantized large language model, the method further includes: monitoring the video memory occupancy information of the computing device during the execution of the accuracy evaluation every preset time interval; determining the video memory occupancy of the original large language model and the quantized large language model each time according to the video memory occupancy information during the operation of the original large language model and the quantized large language model, and adding the video memory occupancy of the original large language model and the quantized large language model each time to the original large language model video memory occupancy array and the quantized large language model video memory occupancy array respectively; at the end of the accuracy evaluation, determining the video memory occupancy of the original large language model and the quantized large language model respectively according to the original large language model video memory occupancy array and the quantized large language model video memory occupancy array.
[0009] In one example, the determining the video memory occupancy of the original large language model and the quantized large language model respectively according to the original large language model video memory occupancy array and the quantized large language model video memory occupancy array specifically includes: determining the video memory occupancy of the original large language model and the quantized large language model respectively according to the average value of the original large language model video memory occupancy array and the average value of the quantized large language model video memory occupancy array.
[0010] In one example, before monitoring the video memory occupancy information of the computing device during the execution of the accuracy evaluation every preset time interval, the method further includes: after installing the device management library, determining the monitored computing device number; creating a thread for monitoring the video memory occupancy of the computing device and setting the thread as a daemon thread to monitor the video memory occupancy of the computing device during the execution of the accuracy evaluation; passing the computing device number into the device management library to obtain the handle of the computing device, and passing the handle into the device management library to monitor the video memory occupancy information of the computing device through the thread.
[0011] In one example, the determining the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy of the original large language model and the quantized large language model specifically includes: obtaining the video memory occupancy difference between the video memory occupancy of the quantized large language model and the video memory occupancy of the original large language model; obtaining the ratio between the video memory occupancy difference and the video memory occupancy of the original large language model to determine the video memory occupancy coefficient of the quantized large language model.
[0012] In one example, before determining the accuracy coefficient of the quantized large language model based on the accuracy of the original large language model and the quantized large language model, the method further includes: loading the original large language model, the quantized large language model, and the tokenizer respectively through a natural language processing task library; creating a target class for adapting the natural language processing task library to the language model performance evaluation library according to the loaded original large language model, the loaded quantized large language model, the loaded tokenizer, and the type of computing device; passing the target class and the evaluation data set into the natural language processing task library to obtain the evaluation results of the original large language model and the evaluation results of the quantized large language model; parsing the evaluation results of the original large language model to obtain the accuracy of the original large language model, and parsing the evaluation results of the quantized large language model to obtain the accuracy of the quantized large language model.
[0013] In one example, determining the accuracy coefficient of the quantized large language model based on the accuracy of the original large language model and the quantized large language model specifically includes: obtaining the accuracy difference between the accuracy of the original large language model and the accuracy of the quantized large language model; obtaining the accuracy ratio between the accuracy difference and the accuracy of the original large language model; obtaining the difference between 1 and the accuracy ratio to determine the accuracy coefficient of the quantized large language model.
[0014] In one example, before determining the compression coefficient of the quantized large language model based on the file size of the original large language model and the file size of the quantized large language model, the method further includes: creating an original large language model folder path object according to the storage path of the original large language model to recursively traverse all files of the original large language model and generate the file size of the original large language model; creating a quantized large language model folder path object according to the storage path of the quantized large language model to recursively traverse all files of the quantized large language model and generate the file size of the quantized large language model.
[0015] In one example, determining the compression coefficient of the quantized large language model based on the file size of the original large language model and the file size of the quantized large language model specifically includes: obtaining the file size difference between the file size of the quantized large language model and the file size of the original large language model; obtaining the ratio between the file size difference and the file size of the original large language model to determine the compression coefficient of the quantized large language model.
[0016] On the other hand, an embodiment of the present application provides a large language model quantization quality evaluation device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a large language model quantization quality evaluation method described in any one of the above.
[0017] On the other hand, an embodiment of the present application provides a non-volatile computer storage medium for large language model quantization quality evaluation, storing computer-executable instructions that can execute a large language model quantization quality evaluation method described in any one of the above.
[0018] The above at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects:
[0019] By setting three indicators of inference accuracy, model size, and video memory occupancy, and fusing the three indicators, it is possible to evaluate from three key perspectives of inference accuracy, model size, and video memory occupancy, ensure the balance between the performance and storage of the quantized large language model, and also ensure the stable operation of the model in different environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the present application, some embodiments of the present application will be described in detail below with reference to the drawings. In the drawings:
[0021] Figure 1 is a schematic flowchart of a large language model quantization quality evaluation method provided by an embodiment of the present application;
[0022] Figure 2 is a schematic structural diagram of a large language model quantization quality evaluation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0024] Some embodiments of the present application will be described in detail below with reference to the drawings.
[0025] Figure 1It is a schematic flowchart of a method for evaluating the quantization quality of a large language model provided by an embodiment of this application. This method can be applied to different business fields, such as the Internet finance business field, the e-commerce business field, the instant messaging business field, the game business field, the official business field, etc. Some input parameters or intermediate results in this process allow manual intervention and adjustment to help improve accuracy.
[0026] The implementation of the analysis method involved in the embodiments of this application can be a terminal device or a server, and this application does not make special restrictions on this. For the convenience of understanding and description, the following embodiments will be described in detail using a server as an example.
[0027] It should be noted that this server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make specific limitations on this.
[0028] Figure 1 The process in
[0029] S101: Determine the accuracy coefficient of the quantized large language model according to the accuracy of the original large language model and the accuracy of the quantized large language model.
[0030] In some embodiments of this application, after saving the original large language model, perform a quantization operation on the original large language model and save the quantized large language model.
[0031] Based on this, the process of determining the accuracy of the original large language model and the accuracy of the quantized large language model is as follows:
[0032] First, load the original large language model, the quantized large language model, and the tokenizer respectively through the natural language processing task library. For example, the natural language processing task library can be the transformers library.
[0033] Then, create a target class for adapting the natural language processing task library to the language model performance evaluation library according to the loaded original large language model, the loaded quantized large language model, the loaded tokenizer, and the computing device type. For example, the language model performance evaluation library can be the lm-evaluation-harness library. After installing the lm-evaluation-harness library, create an HFLM object of the lm-evaluation-harness library according to the original large language model, the quantized large language model, the tokenizer, and the computing device type loaded by the transformers library. Its main function is to adapt the Transformer model to the unified interface of lm-eval for automated evaluation of tasks.
[0034] It should be noted that the computing device type can be GPU or CPU.
[0035] Then, the target class and the evaluation dataset are passed into the natural language processing task library to obtain the evaluation results of the original large language model and the evaluation results of the quantized large language model. For example, using the simple_evaluate method of the lm-evaluation-harness library, passing in the HFLM object and the dataset, conducting the evaluation on the specified evaluation dataset, exporting the evaluation results to json, only taking the result field, and storing the evaluation results of the original large language model and the quantized large language model in a database, or a storage medium such as a file.
[0036] Finally, parse the evaluation results of the original large language model to obtain the accuracy of the original large language model, and parse the evaluation results of the quantized large language model to obtain the accuracy of the quantized large language model.
[0037] Among them, parsing the evaluation results and obtaining the value of results.datasetname.acc is the accuracy of the large language model.
[0038] Based on this, the process of determining the accuracy coefficient of the quantized large language model is as follows:
[0039] First, calculate the accuracy difference between the accuracy of the original large language model and the accuracy of the quantized large language model.
[0040] Then, calculate the accuracy ratio between the accuracy difference and the accuracy of the original large language model.
[0041] Finally, calculate the difference between 1 and the accuracy ratio to determine the accuracy coefficient of the quantized large language model, and store the accuracy coefficient in a database, or a storage medium such as a file.
[0042] For example, the formula for the accuracy coefficient is as follows:
[0043]
[0044] Among them, k is the accuracy coefficient, S 原始 is the accuracy of the original large language model, S 量化 is the accuracy of the quantized large language model.
[0045] S102: Determine the compression coefficient of the quantized large language model according to the file size of the original large language model and the file size of the quantized large language model.
[0046] In some embodiments of the present application, the process of determining the file size of the original large language model and the file size of the quantized large language model is as follows:
[0047] On the one hand, create a folder path object for the original large language model according to the storage path of the original large language model to recursively traverse all files of the original large language model and generate the file size of the original large language model.
[0048] It should be noted that the recursively obtained file size needs to be converted from the byte unit to the MB unit to obtain the file size of the original large language model.
[0049] On the other hand, create a folder path object for the quantized large language model according to the storage path of the quantized large language model to recursively traverse all files of the quantized large language model and generate the file size of the quantized large language model.
[0050] It should be noted that the recursively obtained file size needs to be converted from the byte unit to the MB unit to obtain the file size of the quantized large language model.
[0051] For example, after finding the storage paths of the original large language model and the quantized large language model, use Path (path lib module) to create a folder path object, and use rglob('*') to recursively traverse all subdirectories and files of the folder, obtain the sizes of all files in bytes, traverse all files, accumulate their sizes, which is the total size, convert the total size from bytes to MB, and return the converted total size, so as to obtain the file sizes of the original large language model and the quantized large language model.
[0052] Based on this, the process of determining the compression coefficient of the quantized large language model according to the file size of the original large language model and the file size of the quantized large language model is as follows:
[0053] First, calculate the difference in file size between the file size of the quantized large language model and the file size of the original large language model.
[0054] Then, calculate the ratio between the file size difference and the file size of the original large language model to determine the compression coefficient of the quantized large language model, and store the compression coefficient and the model size in a database, or a storage medium such as a file.
[0055] For example, the formula for the compression coefficient is as follows:
[0056]
[0057] Among them, j is the compression coefficient, L 原始 is the file size of the original large language model, L 量化 is the file size of the quantized large language model.
[0058] S103: Determine the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy size of the original large language model and the video memory occupancy size of the quantized large language model.
[0059] In some embodiments of the present application, the process of determining the VRAM occupancy of the original large language model and the VRAM occupancy of the quantized large language model is as follows:
[0060] After installing the device management library, determine the monitored computing device number.
[0061] Create a thread for monitoring the VRAM occupancy of the computing device, and set the thread as a daemon thread to monitor the VRAM occupancy of the computing device during the execution of the accuracy evaluation.
[0062] It should be noted that the accuracy evaluation at this time can be the accuracy evaluation process of the above S101, or it can be the process of separately evaluating the accuracy of the dataset used for calculating the VRAM occupancy coefficient.
[0063] Pass the computing device number into the device management library to obtain the handle of the computing device, and pass the handle into the device management library to monitor the VRAM occupancy information of the computing device through the thread.
[0064] Monitor the VRAM occupancy information of the computing device during the execution of the accuracy evaluation every preset time interval.
[0065] According to the VRAM occupancy information of the original large language model and the quantized large language model during operation, determine the VRAM occupancy of the original large language model and the quantized large language model each time, and add the VRAM occupancy of the original large language model and the quantized large language model each time to the original large language model VRAM occupancy array and the quantized large language model VRAM occupancy array respectively.
[0066] When the accuracy evaluation is completed, determine the VRAM occupancy of the original large language model and the quantized large language model respectively according to the original large language model VRAM occupancy array and the quantized large language model VRAM occupancy array. Among them, determine the VRAM occupancy of the original large language model and the quantized large language model respectively according to the average values of the original large language model VRAM occupancy array and the quantized large language model VRAM occupancy array.
[0067] For example, first, install the pynvml library to monitor GPU VRAM usage. Then, set the GPU number to be monitored. Then, create a thread to monitor the VRAM occupancy once every second, set the thread as a daemon thread, and perform VRAM monitoring during the execution of the inference accuracy evaluation interval.
[0068] Then, use the nvmlDeviceGetHandleByIndex method of the pynvml library, pass in the specified GPU number to obtain the handle of the specified GPU, and subsequently, the status information of the GPU, such as VRAM usage, temperature, power consumption, etc., can be queried through this handle.
[0069] Finally, use the nvmlDeviceGetMemoryInfo method of the pynvml library, passing in the GPU handle to obtain the video memory information mem-info of the GPU device.
[0070] Obtain the video memory occupancy of the GPU based on mem-info, convert it to MB units, and add the video memory occupancy size to the video memory occupancy array.
[0071] Based on this, the process of determining the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy size of the original large language model and the quantized large language model is as follows:
[0072] Obtain the difference in video memory occupancy between the video memory occupancy size of the quantized large language model and the original large language model.
[0073] Obtain the ratio between the difference in video memory occupancy and the video memory occupancy size of the original large language model to determine the video memory occupancy coefficient of the quantized large language model.
[0074] For example, the formula for the video memory occupancy coefficient is as follows:
[0075]
[0076] Among them, m is the compression coefficient, X 原始 is the file size of the original large language model, X 量化 is the file size of the quantized large language model.
[0077] S104: Evaluate the quantized large language model according to the accuracy coefficient, the compression coefficient, and the video memory occupancy coefficient.
[0078] Among them, the full score values corresponding to the accuracy coefficient, the compression coefficient, and the video memory occupancy coefficient can be set, and then multiplied by the accuracy coefficient, the compression coefficient, and the video memory occupancy coefficient respectively.
[0079] For example, the comprehensive score = 100 * accuracy coefficient + 100 * compression coefficient + 100 * video memory occupancy coefficient. The range of the comprehensive score is between 0 - 300, and the higher the score, the higher the quality of the model quantization.
[0080] It should be noted that although the embodiments of the present application are introduced and described in sequence for steps S101 to S104 with reference to Figure 1 this does not mean that steps S101 to S104 must be executed in a strict order. The reason why the embodiments of the present application are in accordance with Figure 1The order shown in the figure introduces and explains steps S101 to S104 in sequence, which is to facilitate those skilled in the art to understand the technical solution of the embodiment of the present application. In other words, in the embodiment of the present application, the sequence order among steps S101 to S104 can be appropriately adjusted according to actual needs.
[0081] Through Figure 1 the method, it is possible to evaluate from three key aspects of inference accuracy, model size, and video memory occupancy, ensuring the balance between performance and storage of the quantized large language model, rather than only focusing on a single metric.
[0082] Furthermore, by installing the device management library, it is possible to continuously monitor the video memory occupancy of the computing device, and after the inference evaluation is completed, based on the generated video memory occupancy array, the video memory occupancy during the actual operation process can be obtained. In the prior art, usually, the comparison is based on the set video memory occupancy size of the large language model, and the video memory occupancy during the inference process cannot be obtained. Therefore, the present application can effectively measure the video memory usage of the quantized model on the GPU, provide a basis for large-scale inference deployment, and help select a suitable hardware environment.
[0083] Furthermore, by measuring the reduction rate of the quantized model size, it helps developers judge the impact of different quantization methods on the storage space, facilitating storage optimization decisions on different computing devices.
[0084] Furthermore, by installing the lm-evaluation-harness library, the Transformer model is adapted to the unified interface of lm-eval for automated evaluation of tasks, so as to obtain the accuracy coefficient of the quantized large language model.
[0085] Based on this, it is simple to use. The user only needs to specify the storage location of the large language model to automatically complete the model evaluation, and a comprehensive score is provided, which can intuitively judge the quality of model quantization.
[0086] Based on the same idea, some embodiments of the present application also provide a device and a non-volatile computer storage medium corresponding to the above method.
[0087] Figure 2 FIG. is a schematic structural diagram of a device for evaluating the quantization quality of a large language model provided by an embodiment of the present application, including:
[0088] At least one processor; and,
[0089] A memory communicatively connected to the at least one processor; wherein,
[0090] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a large language model quantization quality assessment method described in any one of the above.
[0091] A non-volatile computer storage medium provided by some embodiments of the present application stores computer-executable instructions that can execute a large language model quantization quality assessment method described in any one of the above.
[0092] The various embodiments in the present application are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0093] The devices and media provided by the embodiments of the present application correspond one by one to the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0094] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0095] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0098] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0099] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0100] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0101] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0102] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the technical principle of the present application shall fall within the protection scope of the present application.
Claims
1. A method for evaluating the quantization quality of a large language model, characterized in that, The method includes: Determine the accuracy coefficient of the quantized large language model according to the accuracy of the original large language model and the accuracy of the quantized large language model; Determine the compression coefficient of the quantized large language model according to the file size of the original large language model and the file size of the quantized large language model; Determine the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy of the original large language model and the video memory occupancy of the quantized large language model; Evaluate the quantized large language model according to the accuracy coefficient, the compression coefficient and the video memory occupancy coefficient.
2. The method according to claim 1, characterized in that Before determining the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy of the original large language model and the video memory occupancy of the quantized large language model, the method further includes: Monitor the video memory occupancy information of the computing device during the execution of the accuracy evaluation every preset time interval; Determine the video memory occupancy of the original large language model and the quantized large language model each time according to the video memory occupancy information of the original large language model and the quantized large language model during operation, and add the video memory occupancy of the original large language model and the quantized large language model each time to the original large language model video memory occupancy array and the quantized large language model video memory occupancy array respectively; When the accuracy evaluation is completed, determine the video memory occupancy of the original large language model and the quantized large language model respectively according to the original large language model video memory occupancy array and the quantized large language model video memory occupancy array.
3. The method according to claim 2, wherein The determining the video memory occupancy of the original large language model and the quantized large language model respectively according to the original large language model video memory occupancy array and the quantized large language model video memory occupancy array specifically includes: Determine the video memory occupancy of the original large language model and the quantized large language model respectively according to the average value of the original large language model video memory occupancy array and the average value of the quantized large language model video memory occupancy array.
4. The method according to claim 2, wherein Before monitoring the video memory occupancy information of the computing device during the execution of the accuracy evaluation every preset time interval, the method further includes: After installing the device management library, determine the monitored computing device number; Create a thread for monitoring the video memory occupancy of the computing device, and set the thread as a daemon thread to monitor the video memory occupancy of the computing device during the execution of the accuracy evaluation; Pass the computing device number into the device management library to obtain the handle of the computing device, and pass the handle into the device management library to monitor the video memory occupancy information of the computing device through the thread.
5. The method according to claim 1, wherein The determining the video memory occupancy coefficient of the quantized large language model according to the video memory occupancy of the original large language model and the video memory occupancy of the quantized large language model specifically includes: Obtain the video memory occupancy difference between the video memory occupancy of the quantized large language model and the video memory occupancy of the original large language model; Obtain the ratio between the video memory occupancy difference and the video memory occupancy of the original large language model to determine the video memory occupancy coefficient of the quantized large language model.
6. The method according to claim 1, characterized in that Before determining the accuracy coefficient of the quantized large language model based on the accuracy of the original large language model and the accuracy of the quantized large language model, the method further includes: Loading the original large language model, the quantized large language model, and the tokenizer respectively through a natural language processing task library; Creating a target class for adapting the natural language processing task library to the language model performance evaluation library according to the loaded original large language model, the loaded quantized large language model, the loaded tokenizer, and the type of computing device; Passing the target class and the evaluation data set into the natural language processing task library to obtain the evaluation results of the original large language model and the evaluation results of the quantized large language model; Parsing the evaluation results of the original large language model to obtain the accuracy of the original large language model, and parsing the evaluation results of the quantized large language model to obtain the accuracy of the quantized large language model.
7. The method according to claim 1, wherein The determining the accuracy coefficient of the quantized large language model based on the accuracy of the original large language model and the accuracy of the quantized large language model specifically includes: Calculating the accuracy difference between the accuracy of the original large language model and the accuracy of the quantized large language model; Calculating the accuracy ratio between the accuracy difference and the accuracy of the original large language model; Calculating the difference between 1 and the accuracy ratio to determine the accuracy coefficient of the quantized large language model.
8. The method according to claim 1, characterized in that, Before determining the compression coefficient of the quantized large language model based on the file size of the original large language model and the file size of the quantized large language model, the method further includes: Creating an original large language model folder path object according to the storage path of the original large language model to recursively traverse all files of the original large language model and generate the file size of the original large language model; Creating a quantized large language model folder path object according to the storage path of the quantized large language model to recursively traverse all files of the quantized large language model and generate the file size of the quantized large language model.
9. The method according to claim 8, wherein The determining the compression coefficient of the quantized large language model based on the file size of the original large language model and the file size of the quantized large language model specifically includes: Calculating the file size difference between the file size of the quantized large language model and the file size of the original large language model; Calculating the ratio between the file size difference and the file size of the original large language model to determine the compression coefficient of the quantized large language model.
10. A large language model quantization quality evaluation device, characterized in that Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a method for evaluating the quantization quality of a large language model according to any one of claims 1-8 above.