Model compression method, model deployment method, electronic device, and storage medium
By grouping and quantizing the graphics processor's video memory capacity, the problem of low efficiency in the quantization of large-model weights is solved, and the processing performance of the hardware system and the efficiency of inference tasks of large-models are improved.
Patent Information
- Application Number
- PCT/IB2024/063335
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-17
AI Technical Summary
Graphics processors are less efficient when quantifying large model weights, resulting in increased pressure on the hardware system and affecting the efficiency of model inference tasks.
By grouping multiple model layers based on the graphics processor's video memory capacity, and quantizing the weight parameters of each layer grouping in sequence, the iterative algorithm is used to generate the target matrix to adjust the unquantized parameters, reduce the pressure of the hardware system, and improve processing performance.
It effectively reduces the pressure on the hardware system, improves the efficiency of large models in inference tasks, and solves the problem of inefficiency of graphics processors when weight quantization.
Smart Images

Figure IB2024063335_17072025_PF_FP_ABST
Abstract
Description
[0001] TECHNICAL FIELD The present disclosure relates to large model technology and cloud computing, and more specifically, to a model compression method, a model deployment method, an electronic device, and a storage medium. Background: With technological advancements, large models are increasingly used. Conventional large model weights are typically represented using floating-point numbers, which consumes a significant amount of video memory. However, different graphics processors have varying video memory capacities, resulting in low quantization efficiency when using graphics processors to quantize the weight parameters of large models. Currently, no effective solution has been proposed to address these issues. SUMMARY: Embodiments of the present disclosure provide a model compression method, a model deployment method, an electronic device, and a storage medium to at least address the technical issue of low graphics processor efficiency when performing weight quantization. According to one aspect of an embodiment of the present disclosure, a model compression method is provided, comprising: obtaining an initial model, wherein the initial model is a pre-trained machine learning model and includes multiple model layers; grouping the multiple model layers based on the graphics memory capacity of a graphics processing unit to obtain at least one layer group; and sequentially quantizing, via the graphics processing unit, weight parameters of the model layers included in the at least one layer group to obtain a target model. According to another aspect of an embodiment of the present disclosure, a model deployment method is provided, comprising: obtaining an initial model, wherein the initial model is a pre-trained machine learning model and includes multiple model layers; grouping the multiple model layers based on the graphics memory capacity of a graphics processing unit to obtain at least one layer group; sequentially quantizing, via the graphics processing unit, weight parameters of the model layers included in the at least one layer group to obtain a target model; and deploying the target model. According to another aspect of an embodiment of the present disclosure, a model compression method is provided, comprising: in response to an input instruction on an operation interface, displaying an initial model on the operation interface, wherein the initial model is a pre-trained machine learning model and includes multiple model layers; and in response to a model compression instruction on the operation interface, displaying a target model on the operation interface, wherein the target model is obtained by sequentially quantizing weight parameters of model layers included in at least one layer group by a graphics processor, wherein the at least one layer group is obtained by grouping the multiple model layers based on the video memory capacity of the graphics processor.According to another aspect of an embodiment of the present disclosure, a model compression method is provided, comprising: obtaining an initial model by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the initial model, the initial model is a pre-trained machine learning model, and the initial model includes multiple model layers; grouping the multiple model layers based on the memory capacity of a graphics processor to obtain at least one layer group; sequentially quantizing, via the graphics processor, weight parameters of the model layers included in the at least one layer group to obtain a target model; and outputting the target model by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the target model. According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a memory storing an executable program; and a processor for executing the program, wherein when the program is executed, any one of the above methods is executed. According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device containing the computer-readable storage medium is controlled to execute any one of the above methods. In an embodiment of the present disclosure, an initial model is obtained, where the initial model is a pre-trained machine learning model comprising multiple model layers; the multiple model layers are grouped based on the graphics processor's memory capacity to obtain at least one layer group; and the weight parameters of the model layers within the at least one layer group are sequentially quantized by the graphics processor to obtain a target model. It is readily apparent that grouping the multiple model layers based on the graphics processor's memory capacity and, after obtaining the at least one layer group, sequentially quantizing the weight parameters of the model layers within the at least one layer group effectively reduces the pressure on the hardware system during quantization, thereby improving the hardware system's processing performance and further enhancing the efficiency of large models in performing inference tasks. This resolves the technical issue of low GPU efficiency during weight quantization. It is readily apparent that the general description above and the detailed description that follow are merely illustrative and illustrative of the present disclosure and do not constitute a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.In the accompanying drawings: Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure; Figure 2 is a flow chart of a model compression method according to embodiment 1 of the present disclosure; Figure 3 is a schematic diagram of a model layer grouping according to an embodiment of the present disclosure; Figure 4a is a schematic diagram of a symmetric quantization method according to an embodiment of the present disclosure; Figure 4b is a schematic diagram of an asymmetric quantization according to an embodiment of the present disclosure; Figure 5 is a schematic diagram of a model compression method according to an embodiment of the present disclosure; Figure 6 is a schematic diagram of a quantization granularity according to an embodiment of the present disclosure; Figure 7 is a flow chart of a model deployment method according to embodiment 2 of the present disclosure; Figure 8 is a flow chart of a model compression method according to embodiment 3 of the present disclosure; Figure 9 is a flow chart of a model compression method according to embodiment 4 of the present disclosure; Figure 10 is a schematic diagram of a model compression device according to embodiment 5 of the present disclosure; Figure 11 is a schematic diagram of a model deployment device according to embodiment 6 of the present disclosure; Figure 12 is a schematic diagram of a model compression device according to embodiment 7 of the present disclosure; Figure 13 is a schematic diagram of a model compression device according to embodiment 8 of the present disclosure; Figure 14 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure. DETAILED DESCRIPTION To help those skilled in the art better understand the present disclosure, the following will provide a clear and complete description of the technical solutions in the embodiments of the present disclosure, in conjunction with the accompanying drawings. It should be noted that the described embodiments represent only a portion of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort should fall within the scope of protection of the present disclosure. It should be noted that the terms "first," "second," and so on, in the specification and claims of the present disclosure, and in the accompanying drawings, are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, such that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to the steps or units expressly listed, but may include other steps or units not expressly listed or inherent to such process, method, product, or apparatus. The technical solutions provided in this disclosure are mainly implemented using large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can typically include hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters.Large models, also known as foundational models, are pre-trained on large amounts of unlabeled corpora, producing pre-trained models with over 100 million parameters. These models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities. Examples include large language models (LLMs) and multimodal pre-training models. It should be noted that in practical applications, large models can be fine-tuned using a small number of samples, allowing them to be applied to different tasks. For example, large models can be widely applied in fields such as natural language processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation. They can also be widely used in natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios of large models include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. In the embodiments of the present disclosure, data processing using a large model in a human-computer interaction scenario is used as an example for explanation. First, some nouns or terms that appear in the description of the embodiments of the present disclosure are applicable to the following explanations: Transformer: Transformer is a deep learning model commonly used to process sequence data. It is widely used in natural language processing tasks (such as machine translation and speech recognition). Among them, the most important component of Transformer is the self-attention mechanism (also known as Self Attention). The self-attention mechanism can simultaneously consider all positions in the input sequence, calculate the correlation of each position with all other positions, and perform a weighted average of each position based on these correlations. In addition to the self-attention mechanism, the Transformer model also includes two parts: an encoder and a decoder.Model quantization: Quantization is the process of approximating the continuous value of a signal to a finite number of discrete values. It can be understood as a method of information compression. It is generally represented by "low bits" on computer systems. Quantization can also be called "fixed-point quantization." Optionally, large models can be quantized using smoothing quantization (also known as Smooth Quant). Smooth Quant can quantize the weights and features of large models into integers, allowing integer calculations to be used when using large models for model inference.
[0002] FP16: Half Precision Floating-Point, a half-precision floating-point representation that requires 16 bits of storage in a computer. Hessian matrix: Also known as the Hessian matrix, a Hessian matrix is a square matrix composed of second-order partial derivatives, used to describe the second-order derivative information of a multivariate function. A tensor: A tensor is a multidimensional array that can be used to represent multidimensional data. In mathematics and physics, tensors are a general concept that can describe vectors, matrices, and higher-dimensional arrays. In machine learning and deep learning, tensors are a commonly used data structure for storing and processing large amounts of data. The square root method: Also known as the Cholesky decomposition, is a method for solving linear equations. It decomposes a positive definite symmetric matrix into the product of a lower triangular matrix and its transpose. It is applicable to symmetric positive definite matrices. The quasi-Newton method: An iterative method for solving unconstrained optimization problems. The basic idea is to transform the original problem into a system of linear equations by approximating the second-order derivative matrix of the objective function. Channel quantization, also known as per-channel quantization, is a quantization technique used to compress and represent the weights and activation values of a neural network. For each weight or feature matrix, each channel has its own set of quantization parameters. Example 1: According to an embodiment of the present disclosure, a model compression method is provided. It should be noted that the steps shown in the flowcharts of the accompanying figures can be executed in a computer system, such as a set of computer-executable instructions. Furthermore, although the flowcharts illustrate a logical order, in some cases, the steps shown or described may be executed in a different order. Considering the large number of model parameters in large models and the limited computing resources of mobile terminals, the model compression method provided in the embodiments of the present disclosure can be applied to the application scenario shown in Figure 1, but is not limited thereto. FIG1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure. In the application scenario shown in FIG1 , a large model is deployed on a server 10. Server 10 can be connected to one or more client devices 20 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Client devices 20 herein may include, but are not limited to, smartphones, tablet computers, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users via a graphical user interface (GUI) to invoke the large model and thereby implement the methods provided by the embodiments of the present disclosure. The GUI displays a dialog box in which users can enter model generation instructions.In an embodiment of the present disclosure, a system consisting of a client device and a server may perform the following steps: The client device executes: inputting a target model generation instruction on a graphical user interface; the server executes: obtaining an initial model based on the target model generation instruction, wherein the initial model is a pre-trained machine learning model and comprises multiple model layers; after obtaining the initial model, the multiple model layers are grouped based on the graphics processor's video memory capacity to obtain at least one layer group; further, the graphics processor sequentially quantizes the weight parameters of the model layers included in the at least one layer group to obtain a target model. It should be noted that if the client device's operating resources meet the requirements for deploying and operating a large model, the present disclosure can be performed on the client device. In this operating environment, the present disclosure provides a model compression method as shown in Figure 2. Figure 2 is a flowchart of the model compression method according to Example 1 of the present disclosure. As shown in Figure 2, the method may include the following steps: Step S202: Obtaining an initial model, wherein the initial model is a pre-trained machine learning model and comprises multiple model layers. The pre-trained machine learning models described above can be large models. These models are large in scale and complexity, typically containing numerous parameters and complex network structures. These models are capable of processing larger amounts of data and more complex tasks. Large models improve model accuracy and performance, but also require greater computing resources and training time. The model layer described above is a crucial component of machine learning, used to represent and process data. Within the model layer, data is processed and transformed using specific algorithms and models to obtain useful information and prediction results. Optionally, the model layer typically includes the following aspects: Feature selection and extraction: Within the model layer, features useful for the task need to be selected and extracted. This can be achieved through feature engineering methods, including selecting the most relevant features, performing feature transformations, and combining them. Data preprocessing: Within the model layer, raw data needs to be preprocessed to meet the model's input requirements. This includes data cleaning, missing value handling, standardization, and normalization. In an optional embodiment, the initial model can be obtained in the following ways: Open source model: Many researchers and developers will release their trained models publicly, and these models can be found in various open source libraries for machine learning and deep learning.Dataset Pre-Trained Models: Some model training starts with pre-trained models based on large-scale datasets. These models have undergone extensive training and fine-tuning on common tasks (such as image classification and object detection). You can use these pre-trained models as initial models and fine-tune them on your own dataset. Self-Trained Models: If you have sufficient data and computing resources, you can train a model from scratch. Optionally, you need to prepare a well-labeled dataset and use an appropriate machine learning algorithm or deep learning architecture for training. Transfer Learning: Transfer learning is a technique that uses a pre-trained model as the initial model. You can use a model trained on a similar task and transfer it to the task you want to create. This method generally speeds up training and improves model performance. Step S204: Group multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group. The aforementioned graphics processing unit (GPU) is a processor specialized for graphics rendering, image processing, and computing. It can perform a large number of computing tasks, particularly intensive tasks such as processing images, videos, and multi-dimensional graphics. Compared to traditional central processing units (CPUs), GPUs have more processing cores and higher parallel computing capabilities. They utilize the Single Instruction, Multiple Data (SIMD) architecture, which allows them to process multiple data simultaneously, thereby accelerating computing speed. This has led to GPUs being widely used in many fields, such as gaming, virtual reality, artificial intelligence, and scientific computing. GPUs are typically composed of multiple processing units, each of which can independently execute instructions and store data. They also have dedicated memory and cache for storing image data and intermediate computational results. Through parallel processing and intelligent memory access, GPUs can quickly process a large number of images and computing tasks. The video memory capacity of the aforementioned graphics processor can be the size of the memory space used by the graphics processor to store graphics data. Video memory capacity is typically expressed in bytes. The size of video memory directly impacts the GPU's performance and capabilities when processing graphics data. Larger video memory capacities can accommodate more graphics data, providing a larger drawing area and higher resolution. Furthermore, larger video memory capacities help improve graphics processing speed and efficiency, reducing data exchange and increasing GPU access speed.Optionally, model quantization requires the calculation of quantization parameters. The most typical method for calculating weight quantization parameters is to calculate the maximum and minimum values of the model weight parameters. However, the implementation of the commonly used Generative Pre-trained Transformer Question Answering (GPTQ) model requires the model to be completely stored in GPU memory, traversing the model to obtain the minimum values of the weight parameters, and then obtaining the maximum values of the features through inference. Obviously, this is not feasible for larger models. For example, the opt-175b model (a neural network language model) requires 324 gigabytes (GB) of video memory to store only the FP16 weight parameters. This makes it impossible to compress large models, thus failing to take advantage of the memory savings brought by quantization. In an optional embodiment, to address this issue, a group-by-group quantization parameter calculation method can be used. Model layers are first grouped based on the GPU's video memory capacity. For example, assuming the GPU's video memory capacity is 80GB, if the weight parameters of a layer in a large model are stored using FP16, they occupy 4GB. The quantized weight parameters occupy 2GB, and the cache required during the calculation process is 1GB. Therefore, the video memory required for each layer of weight parameters is 7GB. Therefore, the large model can be divided into 80 / 7, rounded up, or 11 layer groups. For a large Transformer-based model, multiple decoding layers (also called decoder layers) are grouped together. Optionally, the weight parameters in each layer group can be sequentially stored in the GPU memory in the order of grouping, allowing quantization to be performed using the GPU memory. Figure 3 is a schematic diagram of model layer grouping according to an embodiment of the present disclosure. As shown in Figure 3, multiple model layers can be grouped together, for example, two model layers can be grouped together, and the model layers can be quantized using the GPU memory on a group-by-group basis. Step S206: The graphics processor sequentially quantizes the weight parameters of the model layers included in at least one layer group to obtain a target model. The weight parameters of the model layers may be parameters used to adjust the output of each model layer in the initial model. These parameters are learned during training to enable the initial model to better fit the training data and demonstrate good generalization capabilities on test data. The target model may be a model obtained by quantizing the weight parameters of the model layers of the initial model. Optionally, the target model has a lower graphics memory capacity utilization than the initial model.In an optional embodiment, a graphics processor can use Cholesky decomposition to calculate the inverse of the Hessian matrix. The graphics processor can then sequentially quantize the weight parameters of the model layers included in at least one layer group, thereby obtaining the quantized model layer weight parameters. The quantized model layer weight parameters are then used to replace the weight parameters of the model layers of the initial model, thereby obtaining the target model. In another optional embodiment, the graphics processor can also use a quasi-Newton method to calculate the inverse of the Hessian matrix. The graphics processor can then sequentially quantize the weight parameters of the model layers included in at least one layer group, thereby obtaining the quantized model layer weight parameters. The quantized model layer weight parameters are then used to replace the weight parameters of the model layers of the initial model, thereby obtaining the target model. Optionally, by quantizing the weight parameters of the model layers included in at least one layer group of the initial model, large model weights can be quantized from floating-point numbers to low-bit integers, significantly reducing the amount of graphics memory occupied by the large model weights and alleviating graphics memory pressure. This reduces hardware costs and enables the use of larger batch sizes for large model inference tasks, thereby improving model throughput. In the disclosed embodiments, an initial model is obtained, where the initial model is a pre-trained machine learning model comprising multiple model layers; the multiple model layers are grouped based on the graphics processor's memory capacity to obtain at least one layer group; and the weight parameters of the model layers within the at least one layer group are sequentially quantized by the graphics processor to obtain a target model. It is readily apparent that grouping the multiple model layers based on the graphics processor's memory capacity and, after obtaining the at least one layer group, sequentially quantizing the weight parameters of the model layers within the at least one layer group effectively reduces the pressure on the hardware system during quantization of the model layer weight parameters, thereby improving the hardware system's processing performance and further enhancing the efficiency of large models in performing inference tasks. This resolves the technical issue of low GPU efficiency during weight quantization. In the above-described embodiment of the present disclosure, grouping multiple model layers based on the graphics memory capacity of a graphics processor to obtain at least one layer group includes: determining target storage capacities corresponding to weight parameters and quantized weight parameters of the multiple model layers, wherein the quantized weight parameters are parameters obtained by quantizing the weight parameters, and the target storage capacity is used to represent the capacity occupied by the graphics memory of the graphics processor for storing the weight parameters and quantized weight parameters; determining a target number of model layers to be included in any layer group in the at least one layer group based on the graphics memory capacity and the target storage capacity; and grouping the multiple model layers based on the target number to obtain the at least one layer group.In an optional embodiment, the grouping criteria vary depending on the GPU. A larger number of layers in a group is preferred. Specifically, a finer quantization granularity minimizes the impact of the quantization algorithm on model accuracy. The upper limit for the number of layers is the limit at which the group can complete the quantization calculations while fully occupying the video memory. Generally, the upper limit for the number of layers can be estimated. Specifically, the multiple model layers are grouped based on a target number of model layers in any one of the at least one layer groupings. Alternatively, the target storage capacity corresponding to the weight parameters and quantization weight parameters of the multiple model layers can be determined, and the target number of model layers in any one of the at least one layer groupings can be determined based on the video memory capacity and the target storage capacity. Thus, the number of layers in a group can be determined based on the target data when grouping the multiple model layers. In the above embodiment of the present disclosure, weight parameters of model layers included in at least one layer grouping are quantized in sequence by a graphics processor to obtain a target model, including: determining whether there is a target layer group in at least one layer grouping, wherein the weight parameters of the model layers included in the target layer group are not quantized; if there is a target layer group in at least one layer grouping, the weight parameters of the model layers included in the target layer group are loaded into a video memory of the graphics processor, the weight parameters stored in the video memory are quantized by the graphics processor to obtain quantized weight parameters of the model layers included in the target layer group, and the weight parameters stored in the video memory are deleted; if there is no target layer group in at least one layer grouping, the target model is obtained based on the quantized weight parameters of the model layers included in the at least one layer group. In an optional embodiment, when weight parameters of model layers included in at least one layer group are sequentially quantized by a graphics processor, first, the weight parameters of the model layers included in the at least one layer group are not quantized. The order of the at least one layer grouping can be determined based on the order of multiple model layers in the large model, and the weight parameters of the model layers included in the first layer grouping can be quantized. Optionally, the first layer grouping can be used as the target layer grouping, and the weight parameters of the model layers included in the target layer grouping can be loaded into the graphics memory of the graphics processor. The graphics processor can then be used to quantize the weight parameters stored in the graphics memory and obtain the quantized weight parameters, that is, obtain the quantized weight parameters of the model layers included in the target layer group. Furthermore, after obtaining the quantized weight parameters, the weight parameters stored in the graphics memory can be deleted, and the next layer group can be loaded, and the cycle is repeated until the weight parameters in all layer groups have been quantized.By loading the weight parameters into the graphics processor before quantizing them and releasing the graphics processor's memory space after quantization, it is determined that the graphics processor has sufficient video memory capacity when sequentially quantizing the weight parameters of the model layers included in the at least one layer group. In another optional embodiment, if no model layers in the at least one layer group have unquantized weight parameters, that is, if all model layers in the at least one layer group have been quantized, further quantization is not required. The target model can be directly derived based on the quantized weight parameters of the model layers included in the at least one layer group. In other words, the initial model in which all model layers have been quantized is determined as the target model. Optionally, quantization methods can be divided into symmetric quantization and asymmetric quantization. Symmetric quantization maps floating-point values of 0 to the integer 0. That is, when the smaller and larger floating-point values are not symmetrical, a large amount of integer space is wasted. Figure 4a is a schematic diagram of a symmetric quantization method according to an embodiment of the present disclosure. As shown in Figure 4a, for example, the smaller floating-point value is T28 and the larger floating-point value is 127. Therefore, it can be seen that symmetric quantization is not friendly to situations where the distribution of positive and negative numbers is uneven. Figure 4b is a schematic diagram of asymmetric quantization according to an embodiment of the present disclosure. As shown in Figure 4b, compared to symmetric quantization, asymmetric quantization does not need to follow the mapping rule of 0 unchanged. Therefore, it has a wider dynamic mapping range. The larger floating-point value 255 and the smaller floating-point value 0 exactly correspond to the maximum and minimum integer values, making full use of the integer space. Optionally, using asymmetric quantization can fully utilize the limited integer space (for example, 4-bit quantization only allows for 16 integers), thereby minimizing the impact of quantization on the model. Optionally, the present disclosure does not impose any specific restrictions on the quantization method, and asymmetric quantization is used as an example for illustration. In the above embodiment of the present disclosure, when there are multiple graphics processors, the method further includes: splitting the initial model based on the number of graphics processors to obtain multiple sub-models, where different sub-models include different dimension parameters in the weight parameters of multiple model layers; and grouping the multiple model layers included in the sub-models corresponding to the graphics processors based on the video memory capacity of the graphics processors to obtain at least one layer group. The different dimension parameters can be parameter values of different dimensions within the same weight parameter. Optionally, the present disclosure does not impose any specific restrictions on the number of dimensions; for example, it can be three dimensions.In an optional embodiment, the calculation of the Hessian matrix weights requires the use of the weights of all model layers. However, when there are multiple graphics processors, a tensor parallel approach can be used to process different tensors using different graphics processors, so that the weight parameters of a model layer are split onto different graphics processors. That is, multiple sub-models are obtained, and the different sub-models contain the same model layers, but the dimensions of the weight parameters of the model layers are different. Furthermore, tensor parallelism is detrimental to the calculation of the Hessian matrix. One solution is to communicate with multiple GPUs for each calculation. However, this not only reduces quantization efficiency but also places high demands on engineering implementation. Therefore, this solution hypothesizes that independently calculating the Hessian matrix using each GPU's own weights will not significantly affect the accuracy of the target model. Experimental verification confirms this assumption. Therefore, after implementing tensor parallelism, large models only need to independently calculate the corresponding Hessian matrices. Specifically, the initial model is split into multiple sub-models based on the number of GPUs. The multiple model layers within the sub-models corresponding to the GPUs are grouped based on the GPU's video memory capacity to obtain at least one layer group. This leverages the advantages of multiple GPUs to accelerate the quantization rate while ensuring the effectiveness of the quantized model, while also reducing engineering development efforts. In the above embodiment of the present disclosure, the method further includes: generating, by a graphics processor, a target matrix corresponding to the weight parameters of multiple model layers using an iterative algorithm, wherein the target matrix is used to determine quantization errors generated by quantizing the weight parameters of the multiple model layers; and, during the process of sequentially quantizing the weight parameters of the model layers included in at least one layer group by the graphics processor, quantizing a first parameter among the weight parameters of the model layers by the graphics processor, and adjusting an unquantized second parameter among the weight parameters of the model layers based on the target matrix. The iterative algorithm may be a quasi-Newton method. Optionally, the quasi-Newton method is an iterative algorithm for solving nonlinear optimization problems that updates the search direction by approximating the Hessian matrix. Specifically, the quasi-Newton method uses a series of positive definite matrices to approximate the inverse of the Hessian matrix, thereby avoiding the complexity of directly calculating the Hessian matrix. The target matrix may be the inverse of the Hessian matrix. The first parameter may be the weight parameter to be quantized. The second parameter may be the unquantized weight parameter among the weight parameters of the model layers.In an optional embodiment, since quantization of the weight parameters of multiple model layers may generate quantization errors, a graphics processor can use a quasi-Newton method to generate the inverse of the Hessian matrix corresponding to the weight parameters of the multiple model layers, that is, to generate a target matrix. This allows the graphics processor to adjust the unquantized second parameter of the model layer weight parameters based on the target matrix when quantizing the first parameter of the model layer weight parameters, thereby avoiding the complexity of directly calculating the Hessian matrix. Optionally, through continuous iteration, the BFGS algorithm gradually approximates the inverse of the Hessian matrix (i.e., B), thereby obtaining a better search direction. Furthermore, when the graphics processor sequentially quantizes the weight parameters of the model layers included in at least one layer group, the graphics processor can quantize the weight parameters of the model layers that need to be quantized, and adjust the unquantized weight parameters of the model layer weight parameters based on the target matrix. In the above embodiment of the present disclosure, when there are multiple graphics processors, the method further includes: obtaining target dimension parameters included in a sub-model corresponding to the graphics processor; generating, using an iterative algorithm, a sub-target matrix corresponding to the target dimension parameters by the graphics processor; and, while the graphics processor sequentially quantizes the target dimension parameters of the model layers included in at least one layer group, quantizing, by the graphics processor, a first sub-parameter in the target dimension parameters, and adjusting, based on the sub-target matrix, an unquantized second sub-parameter in the target dimension parameters. The sub-model may be a model corresponding to each of the multiple graphics processors. The target dimension parameters may be weight parameters corresponding to the sub-model. The sub-target matrix may be the inverse of the Hessian matrix corresponding to the sub-model. The first sub-parameter may be a weight parameter to be quantized in the sub-model. The second sub-parameter may be an unquantized weight parameter in the weight parameters of the model layers in the sub-model. In an optional embodiment, the model after tensor parallelization only needs to independently calculate the corresponding Hessian matrix. That is, when there are multiple graphics processors, it is necessary to obtain the target dimension parameters contained in the sub-model corresponding to the graphics processor, and use the quasi-Newton algorithm on the graphics processor to generate the sub-target matrix corresponding to the target dimension parameters. Therefore, the graphics processor can sequentially quantize the target dimension parameters of the model layers contained in at least one layer group. Optionally, during the process of sequentially quantizing the target dimension parameters of the model layers contained in at least one layer group, the graphics processor can quantize the first sub-parameter of the target dimension parameter and adjust the unquantized second sub-parameter of the target dimension parameter based on the sub-target matrix, thereby avoiding the complexity of directly calculating the Hessian matrix.In the above-described embodiment of the present disclosure, generating a target matrix corresponding to weight parameters of multiple model layers through an iterative algorithm includes: determining a current matrix and a current gradient vector in a current iteration process, wherein, if the current iteration process is the first iteration process, the current matrix is determined to be a unit matrix, and the current gradient vector is determined to be a gradient vector determined based on the weight parameters of multiple model layers; updating the current weight parameters in the current iteration process based on the current matrix and the current gradient vector to obtain updated weight parameters, wherein, if the current iteration process is the first iteration process, the current weight parameters do not include weight parameters of multiple model layers; determining an updated gradient vector based on the updated weight parameters; updating the current matrix based on the updated gradient vector and the current gradient vector to obtain an updated matrix; and determining the updated matrix as the target matrix if the current iteration process satisfies a stop condition. The stop condition may be that the number of iterations reaches a preset number or the difference between two iteration results is less than a threshold. The preset number and the threshold can be set by those skilled in the art as needed. In an optional embodiment, when calculating the inverse of the Hessian matrix, a quasi-Newton method, a classic method for solving nonlinear optimization problems, can be used. This algorithm approximates the inverse of the Hessian matrix by continuously updating a positive definite symmetric matrix. Optionally, if the current iteration is the first iteration, the identity matrix can be determined as the current matrix. Furthermore, a gradient vector can be determined based on weight parameters of multiple model layers. After determining the current matrix and the current gradient vector, the current weight parameters in the current iteration are updated based on the current matrix and the current gradient vector to obtain updated weight parameters. The updated weight parameters can be used to update the current matrix. Optionally, during the process of updating the current matrix, an updated gradient vector can be first determined based on the updated weight parameters, and the current matrix can be updated using the updated gradient vector and the current gradient vector to obtain an updated matrix. Optionally, after determining the updated matrix, if the number of iterations reaches a preset number or the difference between two iteration results is less than a threshold, the updated matrix can be determined as the target matrix. In the above-described embodiment of the present disclosure, updating the current weight parameter in the current iteration process based on the current matrix and the current gradient vector to obtain the updated weight parameter includes: determining the current search direction based on the current matrix and the current gradient vector; searching in the current search direction to determine the current search step size; and updating the current weight parameter based on the current search direction and the current search step size to obtain the updated weight parameter.In an optional embodiment, let the current matrix be matrix B. For each iteration, the current search direction d can be obtained by calculating the product of the current matrix and the current gradient vector g, or by calculating the inverse of the product of the current matrix and the current gradient vector. Furthermore, a search step size a can be determined by performing a one-dimensional search in the search direction, so that the current weight parameter can be updated based on the current search direction and the current search step size to obtain the updated weight parameter. In the above embodiment of the present disclosure, determining the current search direction based on the current matrix and the current gradient vector includes: obtaining the product of the current matrix and the current gradient vector to obtain a first vector; and obtaining the inverse of the first vector to obtain the current search direction. In an optional embodiment, when determining the current search direction, the current search direction d can be determined by obtaining the inverse of the product of the current matrix and the current gradient vector. Specifically, the following formula can be used, where B is the current matrix and g is the current gradient vector. In the above embodiment, updating the current weight parameter based on the current search direction and the current search step to obtain the updated weight parameter includes: obtaining the product of the current search step and the current search direction d to obtain a second vector; and obtaining the sum of the second vector and the current weight parameter to obtain the updated weight parameter. In an optional embodiment, the product of the current search step size a and the current search direction d may be first determined to obtain a second vector axd, and the sum of the second vector and the current weight parameter x may be calculated. Specifically, the following formula may be used to calculate x' = x + axd. In the above embodiment of the present disclosure, updating the current matrix based on the updated gradient vector and the current gradient vector to obtain the updated matrix includes: determining the current search direction based on the current matrix and the current gradient vector; searching in the current search direction to determine the current search step size; obtaining the difference between the updated gradient vector and the current gradient vector to obtain a third vector; obtaining the product of the current search step size and the current search direction to obtain a second vector; and determining the updated matrix based on the third vector, the transposed vector of the third vector, the second vector, the transposed vector of the second vector, the current matrix, and the transposed vector of the current matrix. In an optional embodiment, the difference between the updated gradient vector and the current gradient vector may be obtained to obtain a third vector y. Specifically, y = g' - g may be calculated using the following formula. Furthermore, the second vector s may be calculated, where s = axd. Further, the update matrix can be determined based on the third vector, the transposed vector of the third vector, the second vector, the transposed vector of the second vector, the current matrix, and the transposed vector of the current matrix. Specifically, the following formula can be used: Wherein, y' is the transposed vector of the third vector, s' is the transposed vector of the second vector, and & is the transposed vector of the current matrix. In the above embodiment of the present disclosure, if the current iteration process does not meet the stop iteration condition, the method further includes: using the updated matrix as the matrix for the next iteration process, using the updated gradient vector as the gradient vector for the next iteration process, and using the updated matrix as the matrix for the next iteration process; and repeating the next iteration process until the next iteration process meets the stop iteration condition. In an optional embodiment, if the current iteration process does not meet the stop iteration condition, that is, the number of iterations has not reached a preset number, or the difference between two iteration results is greater than or equal to a threshold, the updated matrix may be used as the current matrix for the next iteration process, the updated gradient vector may be used as the gradient vector for the next iteration process, and the iteration process may be repeated until the next iteration process meets the stop iteration condition. In the above embodiment of the present disclosure, the current iteration process meeting the stop iteration condition includes one of the following: the number of iterations corresponding to the current iteration process reaches a preset number; or the difference between the updated matrix and the current matrix is less than a threshold. The above-mentioned preset times and thresholds can be set by those skilled in the art based on actual needs. In an optional embodiment, if the current iteration number corresponding to the current iteration process reaches the preset number, or the difference between the updated matrix and the current matrix is less than the threshold, the iteration process can be considered to have met the termination condition, and therefore, the updated matrix can be determined as the target matrix. In the above-mentioned embodiment of the present disclosure, quantizing a first parameter among the weight parameters of the model layer by a graphics processor and adjusting an unquantized second parameter among the weight parameters of the model layer based on the target matrix includes: sequentially obtaining the first parameter of the weight parameters of the model layer in a preset direction; quantizing the first parameter by a graphics processor to obtain a quantization parameter corresponding to the first parameter; determining a quantization error corresponding to the first parameter based on the first parameter, the quantization parameter, and the target matrix; and adjusting the second parameter based on the quantization error to obtain an adjusted weight parameter. In an optional embodiment, when quantizing the first parameter among the weight parameters of the model layer by a graphics processor and adjusting the unquantized second parameter among the weight parameters of the model layer based on the target matrix, the following formula can be used:
[0003] E: ,j — i~ (W: ,j — Q: ,j) / [HT]jj,
[0004] W: ,j(i + B) ~ W: ,j(i + B) — E: ,j — HTj,j:(i + B), where W: ,j is the first parameter, Q: ,j is the quantization parameter, E: ,j — i is the error quantization, W: ,j(i + B) is the second parameter, and HT is the inverse of the Hessian matrix. When obtaining quantized weights based on the Hessian matrix and calculating the quasi-Newton algorithm for the Hessian matrix, one column of the weights can be first selected for quantization. Furthermore, the Hessian matrix is used to calculate the impact of the quantization error on the remaining weight parameters, and finally, the remaining weight parameters are updated to eliminate the impact. Figure 5 is a schematic diagram of a model compression method according to an embodiment of the present disclosure. As shown in Figure 5 , assuming two quantization processes, a portion of quantized weight parameters are obtained after the first quantization operation. Therefore, the remaining weight parameters need to be updated. The result obtained after the first quantization operation serves as the initial result for the second quantization operation. The first quantization process is then repeated, i.e., the quantization operation is performed and the remaining weight parameters are updated. In the above embodiment of the present disclosure, quantizing a first parameter by a graphics processor to obtain a quantization parameter corresponding to the first parameter includes: grouping different channels included in the first parameter to obtain multiple sub-channels; and quantizing the multiple sub-channels by the graphics processor based on the quantization methods corresponding to the sub-channels to obtain the quantization parameters. In an optional embodiment, quantization granularity can be achieved by grouping multiple model layers. This granularity goes a step further than conventional per-channel quantization. Grouping is performed within each channel 1, with the weights or features of each group sharing a set of quantization parameters. Figure 6 is a schematic diagram of a quantization granularity according to an embodiment of the present disclosure. As shown in Figure 6, the left side of Figure 6 shows multiple model layers before grouping, and the right side of Figure 6 shows the layer groups obtained after grouping, such as the first group and the second group. Optionally, asymmetric quantization can be used to quantize the model layers. Asymmetric quantization converts weights from floating-point numbers to integers, that is, Q:,j ~ quant(W:,j). The error generated during the quantization calculation can be determined by E:,j - i ~ (W:,j - Q:,j) / [HT]jj. Optionally, the different channels included in the first parameter can be grouped, and grouping can be performed within each channel 1, thereby obtaining multiple sub-channels.It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. It should be noted that for the sake of simplicity, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited by the order of the actions described, as certain steps can be performed in a different order or simultaneously according to this disclosure. Secondly, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this disclosure. Through the above description of the embodiments, those skilled in the art will clearly understand that the methods according to the above embodiments can be implemented using software and a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the prior art, can essentially be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk) and includes instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present disclosure. Embodiment 2: According to an embodiment of the present disclosure, a model deployment method is also provided. Figure 7 is a flowchart of the model deployment method according to Embodiment 2 of the present disclosure. As shown in Figure 7, the method includes the following steps: Step S702: Obtaining an initial model, where the initial model is a pre-trained machine learning model and includes multiple model layers. Step S704: Grouping the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group. Step S706: Using the graphics processor, sequentially quantizing the weight parameters of the model layers included in the at least one layer group to obtain a target model. Step S708: Deploying the target model.In an optional embodiment, after obtaining the initial model, the different model layers in the initial model can be grouped using the graphics processor's video memory capacity to obtain at least one layer group. Furthermore, the graphics processor sequentially quantizes the weight parameters of the model layers contained in the at least one layer group, using the layer group as a unit, to obtain an update matrix. When an iteration during the quantization process meets a stop condition, the update matrix is determined as the target matrix, thereby obtaining a target model, which can then be deployed. Embodiment 3 According to an embodiment of the present disclosure, a model compression method is also provided. Figure 8 is a flowchart of the model compression method according to Embodiment 3 of the present disclosure. As shown in Figure 8 , the method includes the following steps: Step S802: In response to an input instruction on the operation interface, displaying the initial model on the operation interface, wherein the initial model is a pre-trained machine learning model and comprises multiple model layers. Step S804: In response to the model compression instruction applied to the operation interface, a target model is displayed on the operation interface. The target model is obtained by sequentially quantizing the weight parameters of the model layers included in at least one layer group by the graphics processor. The at least one layer group is obtained by grouping multiple model layers based on the graphics processor's video memory capacity. The aforementioned input instruction can be used to instruct the display of the initial model on the operation interface. Optionally, the user can send the input instruction to the operation interface via any instruction sending method. The aforementioned model compression instruction can be used to instruct the compression of the initial model. Optionally, the user can send the model compression instruction to the operation interface via any instruction sending method. In an optional embodiment, the user can provide an input instruction on the operation interface of the client 20. After receiving the input instruction, the server 10 can retrieve the initial model and display the initial model on the operation interface. Furthermore, the user can provide a model compression instruction on the operation interface of the client 20. After receiving the user's model compression instruction, the server 10 can compress the initial model to obtain the target model and display the target model on the operation interface. Embodiment 4: According to an embodiment of the present disclosure, a model compression method is also provided. Figure 9 is a flowchart of the model compression method according to Embodiment 4 of the present disclosure. As shown in Figure 9, the method includes the following steps: Step S902: Acquire an initial model by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the initial model. The initial model is a pre-trained machine learning model, and the initial model includes multiple model layers. Step S904: Group the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group.Step S906: The graphics processor sequentially quantizes the weight parameters of the model layers included in at least one layer group to obtain a target model. Step S908: The target model is output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target model. In an optional embodiment, an initial model can be obtained through the first interface on the client 20. After obtaining the initial model, the graphics processor's video memory capacity can be accessed through the server 10 to group multiple model layers to obtain at least one layer group. The graphics processor then sequentially quantizes the weight parameters of the model layers included in the at least one layer group to obtain a target model. The target model can then be output through the second interface on the client 20. Example 5 According to an embodiment of the present disclosure, an apparatus for implementing the above-mentioned model compression method is also provided. Figure 10 is a schematic diagram of the model compression apparatus according to Example 5 of the present disclosure. As shown in Figure 10, the apparatus includes: an acquisition module 1002, a grouping module 1004, and a quantization module 1006. The acquisition module 1004 is configured to acquire an initial model, where the initial model is a pre-trained machine learning model and includes multiple model layers. The grouping module 1004 is configured to group the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group. The quantization module 1006 is configured to sequentially quantize the weight parameters of the model layers included in the at least one layer group using the graphics processor to obtain a target model. In the above embodiment of the present disclosure, the grouping module 1004 includes: a first determination unit configured to determine target storage capacities corresponding to the weight parameters and quantized weight parameters of the multiple model layers, where the quantized weight parameters are parameters obtained by quantizing the weight parameters, and the target storage capacity represents the capacity occupied by the graphics processor's video memory for storing the weight parameters and quantized weight parameters; a second determination unit configured to determine a target number of model layers to be included in any one of the at least one layer group based on the video memory capacity and the target storage capacity; and a grouping unit configured to group the multiple model layers based on the target number to obtain at least one layer group.In the above embodiment of the present disclosure, the quantization module 1006 includes: a third determination unit for determining whether a target layer group exists in at least one layer group, wherein the weight parameters of the model layers included in the target layer group are not quantized; a loading unit for, if the target layer group exists in at least one layer group, loading the weight parameters of the model layers included in the target layer group into the graphics memory of the graphics processor, quantizing the weight parameters stored in the graphics memory by the graphics processor to obtain quantized weight parameters of the model layers included in the target layer group, and deleting the weight parameters stored in the graphics memory; and a fourth determination unit for, if the target layer group does not exist in at least one layer group, obtaining a target model based on the quantized weight parameters of the model layers included in the at least one layer group. In the above embodiment of the present disclosure, if there are multiple graphics processors, the apparatus further includes: a splitting module for splitting the initial model based on the number of graphics processors to obtain multiple sub-models, wherein different sub-models include different dimension parameters in the weight parameters of the multiple model layers; and a second grouping module for grouping the multiple model layers included in the sub-model corresponding to the graphics processor based on the graphics memory capacity of the graphics processor to obtain at least one layer group. In the above embodiment of the present disclosure, the apparatus further includes: a first generation module for generating target matrices corresponding to weight parameters of multiple model layers using an iterative algorithm via a graphics processor, wherein the target matrix is used to determine quantization errors generated by quantizing the weight parameters of the multiple model layers; a second quantization module for quantizing a first parameter of the weight parameters of the model layers included in at least one layer group via the graphics processor, and adjusting an unquantized second parameter of the weight parameters of the model layers based on the target matrix. In the above embodiment of the present disclosure, if there are multiple graphics processors, the apparatus further includes: a second acquisition module for acquiring target dimension parameters included in a sub-model corresponding to the graphics processor; a second generation module for generating a sub-target matrix corresponding to the target dimension parameters via the graphics processor using an iterative algorithm; and a third quantization module for quantizing a first sub-parameter of the target dimension parameters via the graphics processor, and adjusting an unquantized second sub-parameter of the target dimension parameters based on the sub-target matrix, when the graphics processor sequentially quantizes the target dimension parameters of the model layers included in at least one layer group.In the above embodiment of the present disclosure, the first generation module includes: a fifth determination unit, configured to determine a current matrix and a current gradient vector in a current iteration process, wherein, when the current iteration process is the first iteration process, the current matrix is determined to be a unit matrix, and the current gradient vector is determined to be a gradient vector determined based on weight parameters of multiple model layers; a first update unit, configured to update the current weight parameters in the current iteration process based on the current matrix and the current gradient vector to obtain updated weight parameters, wherein, when the current iteration process is the first iteration process, the current weight parameters do not include weight parameters of multiple model layers; a sixth determination unit, configured to determine an updated gradient vector based on the updated weight parameters; a second update unit, configured to update the current matrix based on the updated gradient vector and the current gradient vector to obtain an updated matrix; and a seventh determination unit, configured to determine, when the current iteration process satisfies an iteration stop condition, that the updated matrix is a target matrix. In the above embodiment of the present disclosure, the first updating unit includes: a first determining subunit for determining a current search direction based on a current matrix and a current gradient vector; a second determining subunit for searching in the current search direction and determining a current search step size; and a first updating subunit for updating a current weight parameter based on the current search direction and the current search step size to obtain an updated weight parameter. In the above embodiment of the present disclosure, the first determining subunit is further configured to obtain the product of the current matrix and the current gradient vector to obtain a first vector; and obtain the inverse of the first vector to obtain the current search direction. In the above embodiment of the present disclosure, the first updating subunit is further configured to obtain the product of the current search step size and the current search direction to obtain a second vector; and obtain the sum of the second vector and the current weight parameter to obtain the updated weight parameter. In the above embodiment of the present disclosure, the second updating unit includes: a third determining subunit, configured to determine a current search direction based on a current matrix and a current gradient vector; a fourth determining subunit, configured to search in the current search direction and determine a current search step size; a first acquiring subunit, configured to acquire a difference between the updated gradient vector and the current gradient vector to obtain a third vector; a second acquiring subunit, configured to acquire a product of the current search step size and the current search direction to obtain a second vector; and a fifth determining subunit, configured to determine an update matrix based on the third vector, a transposed vector of the third vector, the second vector, the transposed vector of the second vector, the current matrix, and the transposed vector of the current matrix.In the above embodiment of the present disclosure, if the current iteration process does not meet the stop condition, the apparatus further includes: a determination module configured to use the updated matrix as the matrix for the next iteration process, the updated gradient vector as the gradient vector for the next iteration process, and the updated matrix as the matrix for the next iteration process; and an execution module configured to repeatedly execute the next iteration process until the next iteration process meets the stop condition. In the above embodiment of the present disclosure, the second quantization module includes: an acquisition unit configured to sequentially acquire first parameters of weight parameters of model layers in a preset direction; a quantization unit configured to quantize the first parameters using a graphics processor to obtain quantization parameters corresponding to the first parameters; an eighth determination unit configured to determine a quantization error corresponding to the first parameters based on the first parameters, the quantization parameters, and a target matrix; and an adjustment unit configured to adjust the second parameters based on the quantization error to obtain adjusted weight parameters. In the above embodiment of the present disclosure, the quantization unit includes: a grouping subunit configured to group different channels included in the first parameter to obtain multiple sub-channels; and a quantization subunit configured to quantize the multiple sub-channels using a graphics processor based on the quantization methods corresponding to the sub-channels to obtain quantization parameters. It should be noted that the acquisition module 1002, grouping module 1004, and quantification module 1006 described above correspond to steps S202 to S206 in Example 1. The examples and application scenarios implemented by the modules and corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the modules or units described above may be hardware or software components stored in a memory and processed by one or more processors. The modules described above may also be part of a device that can run on the server 10 provided in Example 1. It should be noted that the preferred implementation schemes involved in the above embodiments of the present disclosure are the same as the schemes, application scenarios, and implementation processes provided in Example 1, but are not limited to the schemes provided in Example 1. Example 6 According to an embodiment of the present disclosure, a device for implementing the above-described model deployment method is also provided. FIG11 is a schematic diagram of the model deployment device according to Example 6 of the present disclosure. As shown in FIG11 , the device includes an acquisition module 1102, a grouping module 1104, a quantification module 1106, and a deployment module 1108.The acquisition module 1102 is used to acquire an initial model, where the initial model is a pre-trained machine learning model and includes multiple model layers. The grouping module 1104 is used to group the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group. The quantization module 1106 is used to sequentially quantize the weight parameters of the model layers included in the at least one layer group using the graphics processor to obtain a target model. The deployment module 1108 is used to deploy the target model. It should be noted that the acquisition module 1102, grouping module 1104, quantization module 1106, and deployment module 1108 correspond to steps S702 to S708 in Example 2. The examples and application scenarios implemented by the modules and corresponding steps are the same, but are not limited to the contents disclosed in Example 2. It should be noted that the modules or units described above can be hardware components or software components stored in memory and processed by one or more processors. The modules can also be part of a device and run on the server 10 provided in Example 1. It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present disclosure are the same as the schemes, application scenarios, and implementation processes provided in Example 2, but are not limited to the schemes provided in Example 2. Example 7 According to an embodiment of the present disclosure, a device for implementing the above-mentioned model compression method is also provided. FIG12 is a schematic diagram of the model compression device according to Example 7 of the present disclosure. As shown in FIG12 , the device includes: a first display module 1202 and a second display module 1204. The first display module 1202 is configured to display an initial model on the operation interface in response to an input instruction on the operation interface. The initial model is a pre-trained machine learning model comprising multiple model layers. The second display module 1204 is configured to display a target model on the operation interface in response to a model compression instruction on the operation interface. The target model is obtained by sequentially quantizing weight parameters of model layers included in at least one layer group by a graphics processor. The at least one layer group is obtained by grouping multiple model layers based on the graphics processor's video memory capacity. It should be noted that the first display module 1202 and the second display module 1204 correspond to steps S802 to S804 in Example 3. The examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1.It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors. The above-mentioned modules may also be executed as part of a device in the server 10 provided in Example 1. It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present disclosure are the same as the schemes, application scenarios, and implementation processes provided in Example 3, but are not limited to the schemes provided in Example 3. Example 8 According to an embodiment of the present disclosure, an apparatus for implementing the above-mentioned model compression method is also provided. FIG13 is a schematic diagram of the model compression apparatus according to Example 8 of the present disclosure. As shown in FIG13 , the apparatus includes an acquisition module 1302, a grouping module 1304, a quantization module 1306, and an output module 1308. The acquisition module 1302 is configured to acquire an initial model by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the initial model, and the initial model is a pre-trained machine learning model comprising multiple model layers. The grouping module 1304 is configured to group the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group. The quantization module 1306 is configured to sequentially quantize the weight parameters of the model layers included in the at least one layer group using the graphics processor to obtain a target model. The output module 1308 is configured to output the target model by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the target model. It should be noted that the acquisition module 1302, grouping module 1304, quantization module 1306, and output module 1308 correspond to steps S902 to S908 in Example 4. The examples and application scenarios implemented by the modules and corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory and processed by one or more processors. The above-mentioned modules may also be part of a device and run in the server 10 provided in Example 1. It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present disclosure are the same as the schemes, application scenarios, and implementation processes provided in Example 4, but are not limited to the schemes provided in Example 4. Example 9 The embodiments of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above-mentioned computer terminal may be replaced with a terminal device such as a mobile terminal. Optionally, in this embodiment, the above-mentioned computer terminal may be located in at least one of multiple network devices in a computer network.In this embodiment, the computer terminal can execute program code for the following steps in the model compression method: obtaining an initial model, where the initial model is a pre-trained machine learning model comprising multiple model layers; grouping the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group; and sequentially quantizing the weight parameters of the model layers within the at least one layer group via the graphics processor to obtain a target model. Optionally, Figure 14 is a block diagram of a computer terminal according to an embodiment of the present disclosure. As shown in Figure 14, the computer terminal A may include: one or more (only one shown) processors 1402, a memory 1404, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display. The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the model compression method and apparatus in the embodiments of the present disclosure. The processor executes the software programs and modules stored in the memory to perform various functional applications and data processing, thereby implementing the aforementioned model compression method. The memory may include high-speed random access memory (RAM) and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory located remotely from the processor, which can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The processor may access information and applications stored in the memory via a transmission device to perform the following steps: obtaining an initial model, wherein the initial model is a pre-trained machine learning model comprising multiple model layers; grouping the multiple model layers based on the graphics processor's memory capacity to obtain at least one layer group; and sequentially quantizing weight parameters of the model layers contained in the at least one layer group via the graphics processor to obtain a target model. Optionally, the processor may further execute program code of the following steps: determining a target storage capacity corresponding to weight parameters and quantized weight parameters of multiple model layers, wherein the quantized weight parameters are parameters obtained by quantizing the weight parameters, and the target storage capacity is used to represent the capacity occupied by the graphics memory of the graphics processor for storing the weight parameters and quantized weight parameters; determining a target number of model layers included in any one layer group in at least one layer grouping based on the graphics memory capacity and the target storage capacity; and grouping the multiple model layers based on the target number to obtain at least one layer grouping.Optionally, the processor may further execute program code for the following steps: determining whether a target layer group exists in at least one layer group, wherein the weight parameters of the model layers included in the target layer group are not quantized; if the target layer group exists in at least one layer group, loading the weight parameters of the model layers included in the target layer group into the graphics memory of the graphics processor, quantizing the weight parameters stored in the graphics memory by the graphics processor to obtain quantized weight parameters of the model layers included in the target layer group, and deleting the weight parameters stored in the graphics memory; if the target layer group does not exist in at least one layer group, obtaining a target model based on the quantized weight parameters of the model layers included in the at least one layer group. Optionally, the processor may further execute program code for the following steps: splitting the initial model based on the number of graphics processors to obtain multiple sub-models, wherein different sub-models include different dimension parameters of the weight parameters of the multiple model layers; and grouping the multiple model layers included in the sub-model corresponding to the graphics processor based on the graphics memory capacity of the graphics processor to obtain at least one layer group. Optionally, the processor may further execute program code for the following steps: generating, by the graphics processor, a target matrix corresponding to weight parameters of multiple model layers using an iterative algorithm, wherein the target matrix is used to determine quantization errors generated by quantizing the weight parameters of the multiple model layers; quantizing, by the graphics processor, a first parameter among the weight parameters of the model layers included in at least one layer group, and adjusting, based on the target matrix, an unquantized second parameter among the weight parameters of the model layers. Optionally, the processor may further execute program code for the following steps: obtaining a target dimension parameter included in a sub-model corresponding to the graphics processor; generating, by the graphics processor, a sub-target matrix corresponding to the target dimension parameter using an iterative algorithm; and, by the graphics processor, quantizing, by the graphics processor, a first sub-parameter among the target dimension parameters, and adjusting, based on the sub-target matrix, an unquantized second sub-parameter among the target dimension parameters, in the process of sequentially quantizing, by the graphics processor, the target dimension parameters of the model layers included in at least one layer group.Optionally, the processor may further execute program code for the following steps: determining a current matrix and a current gradient vector in a current iteration process, wherein, if the current iteration process is the first iteration process, the current matrix is determined to be a unit matrix, and the current gradient vector is determined to be a gradient vector determined based on weight parameters of multiple model layers; updating the current weight parameters in the current iteration process based on the current matrix and the current gradient vector to obtain updated weight parameters, wherein, if the current iteration process is the first iteration process, the current weight parameters do not include weight parameters of multiple model layers; determining an updated gradient vector based on the updated weight parameters; updating the current matrix based on the updated gradient vector and the current gradient vector to obtain an updated matrix; and determining the updated matrix as a target matrix if the current iteration process satisfies a stop condition. Optionally, the processor may further execute program code for the following steps: determining a current search direction based on the current matrix and the current gradient vector; searching in the current search direction and determining a current search step size; and updating the current weight parameters based on the current search direction and the current search step size to obtain updated weight parameters. Optionally, the processor may further execute program code for the following steps: obtaining the product of the current matrix and the current gradient vector to obtain a first vector; obtaining the inverse value of the first vector to obtain a current search direction. Optionally, the processor may further execute program code for the following steps: obtaining the product of the current search step size and the current search direction to obtain a second vector; obtaining the sum of the second vector and the current weight parameter to obtain an updated weight parameter. Optionally, the processor may further execute program code for the following steps: determining the current search direction based on the current matrix and the current gradient vector; searching in the current search direction to determine the current search step size; obtaining the difference between the updated gradient vector and the current gradient vector to obtain a third vector; obtaining the product of the current search step size and the current search direction to obtain a second vector; and determining the update matrix based on the third vector, the transposed vector of the third vector, the second vector, the transposed vector of the second vector, the current matrix, and the transposed vector of the current matrix. Optionally, the processor may further execute program code of the following steps: using the updated matrix as the matrix in the next iterative process, using the updated gradient vector as the gradient vector in the next iterative process, and using the updated matrix as the matrix in the next iterative process; and repeatedly executing the next iterative process until the next iterative process satisfies a stop condition.Optionally, the processor may further execute program code for the following steps: sequentially obtaining first parameters of weight parameters of a model layer in a preset direction; quantizing the first parameters using a graphics processor to obtain a quantization parameter corresponding to the first parameters; determining a quantization error corresponding to the first parameters based on the first parameters, the quantization parameters, and a target matrix; and adjusting a second parameter based on the quantization error to obtain an adjusted weight parameter. Optionally, the processor may further execute program code for the following steps: grouping different channels included in the first parameters to obtain multiple sub-channels; quantizing the multiple sub-channels using a graphics processor based on quantization methods corresponding to the sub-channels to obtain quantization parameters. In an embodiment of the present disclosure, an initial model is obtained, where the initial model is a pre-trained machine learning model and includes multiple model layers; grouping the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group; and sequentially quantizing the weight parameters of the model layers included in the at least one layer group using the graphics processor to obtain a target model. It's easy to note that multiple model layers can be grouped based on the graphics processor's video memory capacity. After obtaining at least one layer group, the weight parameters of the model layers included in the at least one layer group are sequentially quantized. This effectively reduces the pressure on the hardware system when quantizing the model layer weight parameters, thereby improving the hardware system's processing performance and further enhancing the efficiency of large models when performing inference tasks. This, in turn, addresses the technical issue of low efficiency of the graphics processor when performing weight quantization. Those skilled in the art will appreciate that the structure shown in the figure is merely illustrative, and the computer terminal may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, or a mobile internet device (MID), PAD, or other terminal device. Figure 14 does not limit the structure of the aforementioned electronic device. For example, computer terminal A may include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 14, or may have a configuration different from that shown in Figure 14. Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the hardware of the terminal device. The program can be stored in a computer-readable storage medium, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Embodiment 10 The embodiments of the present disclosure further provide a storage medium.Optionally, in this embodiment, the storage medium can be used to store program code executed by the model compression method provided in the first embodiment. Optionally, in this embodiment, the storage medium can be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group. Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining an initial model, where the initial model is a pre-trained machine learning model comprising multiple model layers; grouping the multiple model layers based on the graphics processor's video memory capacity to obtain at least one layer group; and sequentially quantizing, via the graphics processor, the weight parameters of the model layers contained in the at least one layer group to obtain a target model. The serial numbers of the embodiments of the present disclosure are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above embodiments of the present disclosure, the descriptions of each embodiment are given with emphasis. For portions not detailed in one embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other divisions may be employed. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be through interfaces, or indirect couplings or communication connections between units or modules, and may be electrical or other. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of these units may be selected to achieve the objectives of the present embodiments as needed. Furthermore, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. These integrated units may be implemented in either hardware or software functional units. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for causing a computer device (such as a personal computer, server, or network device) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), a removable hard drive, a magnetic disk, or an optical disk. The above description is merely a preferred embodiment of the present disclosure. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present disclosure, and such improvements and modifications should also be considered within the scope of protection of the present disclosure.
Claims
Claims 1. A model compression method, comprising: Obtain an initial model, where the initial model is a machine learning model pre-trained, and the initial model includes multiple model layers; group the multiple model layers based on the video memory capacity of a graphics processor to obtain at least one layer group; sequentially quantize the weight parameters of the model layers included in the at least one layer group through the graphics processor to obtain a target model.
2. The method according to claim 1, wherein, Grouping the multiple model layers based on the video memory capacity of a graphics processor to obtain at least one layer group includes: determining the target storage capacity corresponding to the weight parameters and the quantized weight parameters of the multiple model layers, where the quantized weight parameters are parameters obtained by quantizing the weight parameters, and the target storage capacity is used to represent the capacity occupied by storing the weight parameters and the quantized weight parameters in the video memory of the graphics processor; based on the video memory capacity and the target storage capacity, determine the target number of model layers included in any one layer group in the at least one layer group; group the multiple model layers based on the target number to obtain the at least one layer group.
3. The method according to claim 1, wherein, Sequentially quantizing the weight parameters of the model layers included in the at least one layer group through the graphics processor to obtain a target model includes: determining whether there is a target layer group in the at least one layer group, where the weight parameters of the model layers included in the target layer group are not quantized; in the case where the target layer group exists in the at least one layer group, load the weight parameters of the model layers included in the target layer group into the video memory of the graphics processor, quantize the weight parameters stored in the video memory through the graphics processor to obtain the quantized weight parameters of the model layers included in the target layer group, and delete the weight parameters stored in the video memory; in the case where the target layer group does not exist in the at least one layer group, obtain the target model based on the quantized weight parameters of the model layers included in the at least one layer group.
4. The method according to claim 1, wherein In the case where there are multiple graphics processors, the method further includes: splitting the initial model based on the number of graphics processors to obtain multiple sub-models, where different sub-models include different dimensional parameters among the weight parameters of the multiple model layers; grouping the multiple model layers included in the sub-model corresponding to the graphics processor based on the video memory capacity of the graphics processor to obtain the at least one layer group. 25 5. The method according to claim 1, wherein The method further includes: using the graphics processor to generate a target matrix corresponding to the weight parameters of the multiple model layers by means of an iterative algorithm, where the target matrix is used to determine the quantization error generated by quantizing the weight parameters of the multiple model layers; in the process of successively quantizing the weight parameters of the model layers included in at least one layer group by the graphics processor, the graphics processor quantizes a first parameter in the weight parameters of the model layer, and adjusts a second parameter in the weight parameters of the model layer that has not been quantized based on the target matrix.
6. The method according to claim 5, wherein When there are multiple graphics processors, the method further includes: obtaining the target dimension parameters included in the sub-model corresponding to the graphics processor; using the iterative algorithm by the graphics processor to generate a sub-target matrix corresponding to the parameters of the target dimension; in the process of successively quantizing the target dimension parameters of the model layers included in at least one layer group by the graphics processor, the graphics processor quantizes a first sub-parameter in the target dimension parameters, and adjusts a second sub-parameter in the target dimension parameters that has not been quantized based on the sub-target matrix.
7. The method according to claim 5, wherein Generating a target matrix corresponding to the weight parameters of the multiple model layers by means of an iterative algorithm includes: determining a current matrix and a current gradient vector in the current iteration process, where when the current iteration process is the first iteration process, determining the current matrix as the identity matrix and the current gradient vector as the gradient vector determined based on the weight parameters of the multiple model layers; updating the current weight parameters in the current iteration process based on the current matrix and the current gradient vector to obtain updated weight parameters, where when the current iteration process is the first iteration process, the current weight parameters do not have the weight parameters of the multiple model layers; determining an updated gradient vector based on the updated weight parameters; updating the current matrix based on the updated gradient vector and the current gradient vector to obtain an updated matrix; when the current iteration process meets the stop iteration condition, determining the updated matrix as the target matrix.
8. The method according to claim 7, wherein Updating the current weight parameters in the current iteration process based on the current matrix and the current gradient vector to obtain updated weight parameters includes: determining a current search direction based on the current matrix and the current gradient vector; searching in the current search direction to determine a current search step size. Updating the current weight parameters based on the current search direction and the current search step size to obtain the updated weight parameters.
9. The method according to claim 8, wherein Determining a current search direction based on the current matrix and the current gradient vector includes: obtaining the product of the current matrix and the current gradient vector to obtain a first vector; obtaining the opposite value of the first vector to obtain the current search direction.
10. The method according to claim 8 or 9, wherein, Updating the current weight parameter based on the current search direction and the current search step size to obtain the updated weight parameter, including: obtaining a product of the current search step size and the current search direction to obtain a second vector; obtaining a sum of the second vector and the current weight parameter to obtain the updated weight parameter.
11. The method according to claim 7, wherein Updating the current matrix based on the updated gradient vector and the current gradient vector to obtain an updated matrix, including: determining a current search direction based on the current matrix and the current gradient vector; performing a search in the current search direction to determine a current search step size; obtaining a difference between the updated gradient vector and the current gradient vector to obtain a third vector; obtaining a product of the current search step size and the current search direction to obtain a second vector; determining the updated matrix based on the third vector, a transposed vector of the third vector, the second vector, a transposed vector of the second vector, the current matrix, and a transposed vector of the current matrix.
12. The method according to claim 7, wherein In a case where the current iteration process does not satisfy the stop iteration condition, the method further includes: using the updated matrix as the matrix in the next iteration process, using the updated gradient vector as the gradient vector in the next iteration process, and using the updated matrix as the matrix in the next iteration process; repeatedly performing the next iteration process until the next iteration process satisfies the stop iteration condition.
13. The method according to claim 7, wherein The current iteration process satisfying the stop iteration condition includes one of the following: the current iteration number corresponding to the current iteration process reaches a preset number; a difference between the updated matrix and the current matrix is less than a threshold.
14. The method according to claim 5, wherein Quantifying a first parameter in the weight parameter of the model layer by the graphics processing unit, and adjusting a second parameter in the weight parameter of the model layer that is not quantified based on the target matrix, including: sequentially obtaining the first parameter of the weight parameter of the model layer in a preset direction; quantifying the first parameter by the graphics processing unit to obtain a quantization parameter corresponding to the first parameter; determining a quantization error corresponding to the first parameter based on the first parameter, the quantization parameter, and the target matrix; adjusting the second parameter based on the quantization error to obtain an adjusted weight parameter. Quantifying the first parameter by the graphics processing unit to obtain a quantization parameter corresponding to the first parameter, including: grouping different channels included in the first parameter to obtain a plurality of sub-channels; quantifying the plurality of sub-channels by the graphics processing unit based on a quantization method corresponding to the sub-channels to obtain the quantization parameter.
15. The method according to claim 14, wherein Quantifying the first parameter by the graphics processing unit to obtain a quantization parameter corresponding to the first parameter, including: grouping different channels included in the first parameter to obtain a plurality of sub-channels; quantifying the plurality of sub-channels by the graphics processing unit based on a quantization method corresponding to the sub-channels to obtain the quantization parameter.
16. A model deployment method, comprising: Obtain an initial model, where the initial model is a machine learning model pre-trained, and the initial model includes multiple model layers; group the multiple model layers based on the video memory capacity of the graphics processing unit to obtain at least one layer group; sequentially quantize the weight parameters of the model layers included in the at least one layer group through the graphics processing unit to obtain a target model; deploy the target model.
17. A model compression method, comprising: In response to an input instruction acting on the operation interface, display the initial model on the operation interface, where the initial model is a machine learning model pre-trained, and the initial model includes multiple model layers; in response to a model compression instruction acting on the operation interface, display the target model on the operation interface, where the target model is obtained by sequentially quantizing the weight parameters of the model layers included in at least one layer group, and the at least one layer group is obtained by grouping the multiple model layers based on the video memory capacity of the graphics processing unit.
18. A model compression method, comprising: Obtain the initial model by calling a first interface, where the first interface includes a first parameter, and the parameter value of the first parameter includes the initial model, and the initial model is a machine learning model pre-trained, and the initial model includes multiple model layers; group the multiple model layers based on the video memory capacity of the graphics processing unit to obtain at least one layer group; sequentially quantize the weight parameters of the model layers included in the at least one layer group through the graphics processing unit to obtain a target model; output the target model by calling a second interface, where the second interface includes a second parameter, and the parameter value of the second parameter includes the target model. 28 19. An electronic device, comprising: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the method according to any one of claims 1 to 18.
20. A computer-readable storage medium, the computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 18. 29
Citation Information
Patent Citations
Model optimization method, grouping compression method, corresponding device and equipment
CN111738401A
Method for optimizing neural network model and method for providing graphical user interface
CN115220833A
Method and system for federated learning
CN115485700A
Virtual human resource role construction method and system based on large model
CN116821293A