Model quantification method and device, electronic equipment and computer readable storage medium

By quantizing and splitting the parameter set of large models and determining the quantization method based on numerical distribution information, the problem of high memory and computing resources of large models is solved, and more efficient model storage and calculation is achieved.

CN119918595APending Publication Date: 2025-05-02LYNXI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997166.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Due to the huge number of parameters and high accuracy, the large model has a high computing pressure on the processor or processing core, resulting in high memory space consumption and high computing resource consumption.

Method used

By obtaining the numerical distribution information of the model parameter set to be quantized, the model parameter set is quantized, split into multiple subsets to be quantized, and quantizing according to the target quantization method of each subset, reducing the storage amount and calculation pressure of model parameters.

Benefits of technology

It realizes the reduction of the memory space occupied by model parameters, reduces the computing pressure and resource consumption of processors or processing cores, and improves the accuracy of model parameters quantification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918595A_ABST
    Figure CN119918595A_ABST
Patent Text Reader

Abstract

The invention provides a model quantification method and device, electronic equipment and a computer readable storage medium, and the model quantification method comprises the steps: obtaining a to-be-quantized model parameter set, the model parameter set comprising a plurality of first parameters in a target model; according to numerical value distribution information of a model parameter set, quantizing a first parameter in the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set; and determining a quantized target model according to the plurality of quantized subsets of the model parameter set. According to the embodiment of the invention, the method can achieve the quantification of the model parameters of different numerical distribution information according to different target quantification modes, and improves the quantification accuracy of the first parameter in the model parameter set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a model quantization method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] In recent years, large models (LM) have achieved great success in various application fields. During the application of large models, the model structure has an important impact on the application effect of the model. In practical applications, large models usually have a huge number of parameters exceeding 100 billion, such as GPT-3 (Generative Pre-trained Transformer-3), GLM (General Language Model) and other large language models (LLM). Storing the parameters of large models requires a large amount of memory space. During the operation of large models, due to the large number of parameters and high parameter accuracy, the requirements for processors or processing cores are high. Therefore, it is crucial to reduce the memory space occupied by model parameters and reduce the computing pressure of processors or processing cores. Summary of the invention

[0003] The present disclosure provides a model quantization method and device, an electronic device, and a computer-readable storage medium.

[0004] In a first aspect, the present disclosure provides a model quantization method, comprising:

[0005] Acquire a model parameter set to be quantized, wherein the model parameter set includes a plurality of first parameters in a target model;

[0006] quantizing a first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of a numerical range, a numerical distribution law, and a numerical proportion of the plurality of first parameters, and the numerical proportion is used to characterize a ratio between the number of first parameters corresponding to any numerical value within the numerical range and the total number of first parameters;

[0007] A quantized target model is determined according to the multiple quantized subsets of the model parameter set.

[0008] In a second aspect, the present disclosure provides a model quantization device, comprising:

[0009] An acquisition module is configured to acquire a model parameter set to be quantized, wherein the model parameter set includes a plurality of first parameters in a target model;

[0010] a quantization module, configured to quantize a first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of a numerical range, a numerical distribution law, and a numerical proportion of the plurality of first parameters, and the numerical proportion is used to characterize a ratio between the number of first parameters corresponding to any numerical value in the numerical range and the total number of first parameters;

[0011] The determination module is configured to determine a quantized target model according to multiple quantized subsets of the model parameter set.

[0012] In a third aspect, the present disclosure provides an electronic device comprising: a plurality of processing cores; and an on-chip network configured to exchange data between the plurality of processing cores and external data; wherein one or more instructions are stored in one or more of the processing cores, and one or more of the instructions are executed by one or more of the processing cores, so that one or more of the processing cores can execute the above-mentioned model quantization method.

[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned model quantization method when executed by a processing core.

[0014] The embodiments provided by the present disclosure can perform model quantization on the model parameter set according to the numerical distribution information of the model parameter set to be quantized, and obtain multiple quantization subsets after the quantization of the model parameter set, so as to determine the quantized target model according to the multiple quantization subsets. By quantizing the target model, the storage amount of the model parameters is reduced, and the occupation of the memory space by the model parameters is reduced; model quantization can reduce the accuracy of the model parameters, thereby reducing the computing pressure of the model on the processor or processing core during operation. Furthermore, the model parameter set is quantized according to the numerical distribution information of the model parameter set, so as to improve the accuracy of quantizing the first parameter in the model parameter set.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings, in which:

[0017] Figure 1 It is a schematic diagram of an application scenario of a model quantization method provided by an embodiment of the present disclosure;

[0018] Figure 2 is a flow chart of a model quantization method provided by an embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram of numerical distribution of model weight parameters provided by an embodiment of the present disclosure;

[0020] Figure 4 is a schematic diagram of a method for determining target quantification provided by an embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of generating a quantization vector to be determined provided by an embodiment of the present disclosure;

[0022] Figure 6 is a block diagram of a model quantization device provided by an embodiment of the present disclosure;

[0023] Figure 7 It is a block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0025] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.

[0026] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0027] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "including" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. "Connected" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.

[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless explicitly defined as such herein.

[0029] In recent years, large models have achieved great success in various application fields. In the application process of large models, the model structure has an important impact on the application effect of the model. In practical applications, large models usually have a huge number of parameters exceeding 100 billion, which brings huge challenges to the application and deployment of large models. At the same time, with the growing demand for privacy protection, timeliness, etc., it has gradually become necessary to deploy large models on terminal devices such as mobile phones and laptops. Storing the parameters of large models requires a large amount of memory space. During the operation of large models, due to the large number of parameters and high parameter accuracy, the requirements for processors or processing cores are high. Therefore, it is crucial to reduce the memory space occupied by model parameters and reduce the computing pressure of processors or processing cores.

[0030] According to the model quantization method of the embodiment of the present disclosure, the model parameter set can be split into multiple subsets to be quantized according to the numerical distribution information of the model parameter set to be quantized, and the target quantization methods of the multiple subsets to be quantized can be determined respectively, so that each subset to be quantized can be quantized according to the target quantization method corresponding to each subset to be quantized, so as to obtain the quantized subsets corresponding to each subset to be quantized, and achieve the accuracy of quantizing the model parameter set. By quantizing the subsets to be quantized with different target quantization methods for each subset to be quantized, the storage amount of the model parameter set can be reduced, thereby reducing the occupation of memory space by the parameters of the quantized target model; by quantizing each subset to be quantized, the accuracy of the first parameter in each subset to be quantized can be reduced, thereby reducing the computing pressure and computing resources of the processor or processing core during the operation of the quantized target model.

[0031] The model quantization method according to the embodiment of the present disclosure can be executed by an electronic device such as a terminal device or a server, and the terminal device can be a vehicle-mounted device, a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor calling a computer-readable program instruction stored in a memory. Alternatively, the method can be executed by a server.

[0032] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram showing an application scenario of a model quantization method provided in an embodiment of the present disclosure. The application of the model quantization method provided in an embodiment of the present disclosure in an image processing scenario is used as an example for explanation. Figure 1 As shown, a large model is deployed on the user terminal, and an image processing task is performed on the user terminal based on the deployed large model. The image processing task can specifically be tasks such as detecting an object to be detected in an image and generating a corresponding image based on a text prompt word. The image processing task can be determined according to the actual application situation, and the present disclosure does not limit it here. In practical applications, since the memory space of the user terminal is limited, but the large model often has parameters of hundreds of billions, therefore, deploying a large model on the user terminal to perform image processing tasks will occupy a large memory space of the user terminal, and in the process of running the large model, it will bring huge computing pressure to the user terminal.

[0033] Based on this, the large model deployed by the user terminal can be quantized by the model quantization method provided by the embodiment of the present disclosure, so as to continue to perform the image processing task using the quantized large model. The model parameter set to be quantized in the large model and the numerical distribution information of the model parameter set are obtained, and the model parameter set is split according to the numerical distribution information to obtain multiple subsets to be quantized. After determining the target quantization method corresponding to each subset to be quantized, each subset to be quantized is quantized according to the target quantization method corresponding to each subset to be quantized, and the quantized subsets corresponding to each subset to be quantized can be obtained. The quantized model parameter set can be obtained according to the multiple quantized subsets, and the quantized large model can be determined according to the quantized model parameter set. The quantized large model can reduce the storage amount of the model parameter set and reduce the computing pressure and computing resources of the user terminal.

[0034] It should be noted that the above-mentioned image processing scenario is only used to schematically illustrate the application scenario of the model quantization method provided in the embodiment of the present disclosure. The model quantization method provided in the embodiment of the present disclosure can also be applied to other application scenarios, such as video processing scenarios, text processing scenarios, etc. The model quantization method provided in the embodiment of the present disclosure is described in detail below.

[0035] Figure 2 A flow chart of a model quantization method provided according to an embodiment of the present disclosure is shown. Figure 2 , the method specifically comprises the following steps:

[0036] Step 202: Obtain a model parameter set to be quantized, where the model parameter set includes a plurality of first parameters in the target model.

[0037] The model parameter set is specifically a parameter set that needs to be quantized in the target model, and the model parameter set includes multiple parameters that need to be quantized in the target model. The first parameter refers to the parameter that needs to be quantized in the target model, including but not limited to model weight parameters, bias terms, activation values ​​in neural networks, etc. The target model is the model that needs to be quantized. In practical applications, the target model can be a large model, a large language model, etc.

[0038] In practical applications, the target model is used to perform any of image processing tasks, speech processing tasks, text processing tasks, and video processing tasks.

[0039] Specifically, after determining the target model that needs to be quantized, a model parameter set of the target model is obtained, so that the model parameter set is quantized in a subsequent process to obtain a quantized parameter set.

[0040] Step 204: quantize the first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set.

[0041] In practical applications, a model often has a large number of model parameters, and different model parameters have different values. In order to improve the accuracy of quantifying model parameters, after determining the model parameters that need to be quantified, the determined large number of model parameters can be quantified according to the numerical distribution information of the model parameters.

[0042] The numerical distribution information includes at least one of the numerical range, numerical distribution law and numerical proportion of the multiple first parameters; the numerical distribution law may specifically include normal distribution, uniform distribution, Gaussian distribution, etc.; the numerical proportion is used to characterize the ratio between the number of first parameters corresponding to any numerical value within the numerical range and the total number of first parameters. The quantized subset refers to the multiple subsets corresponding to the quantized model parameter set, and the quantized subset includes multiple second parameters; the second parameter specifically refers to the model parameter obtained after the first parameter is quantized.

[0043] Take the model weight parameter as an example to illustrate, see Figure 3 , Figure 3 FIG. 2 shows a schematic diagram of a numerical distribution of a model weight parameter provided according to an embodiment of the present disclosure. Figure 3 As shown, the weight parameters of the target model are sorted from large to small according to the value, and the following is obtained: Figure 3 The numerical distribution diagram shown in FIG. 1 shows that the horizontal axis of the coordinate system is the number of weight parameter values, and the vertical axis of the coordinate system is the value of the weight parameter. Figure 3 It can be seen that the approximate numerical range of the weight parameters of the target model is -0.10 to 0.10, and there are some outliers in the weight parameters of the target model.

[0044] In practical applications, the model parameter set can be quantized according to the numerical distribution information of the model parameter set. To ensure the accuracy of the quantization of the model parameter set, the first parameter in the model parameter set can be quantized according to at least one of the numerical range, numerical distribution law, and numerical proportion of the first parameter in the model parameter set. Taking the numerical distribution information of the model parameter set including the numerical distribution law as an example, if the numerical distribution law of the model parameter set is uniform distribution, the model parameter set can be quantized directly based on the equal interval quantization method; if the numerical distribution of the model parameter set includes both data with uniform numerical distribution and data with uneven data distribution, the data with uniform numerical distribution in the model parameter set can be quantized at equal intervals, and the data with uneven numerical distribution can be quantized at non-equal intervals, etc.; if the numerical distribution law of the model parameter set is a normal distribution, the first parameter in the model parameter set with uniform numerical distribution can be quantized at equal intervals, and the first parameter in the model parameter set with an outlier value can be quantized at non-equal intervals, etc.

[0045] The following describes a specific implementation process of quantizing the first parameter in the model parameter set to obtain multiple quantization subsets according to the numerical distribution information of the model parameter set.

[0046] In a specific implementation provided by the present disclosure, according to the numerical distribution information of the model parameter set, the first parameter in the model parameter set is quantized to obtain multiple quantized subsets corresponding to the model parameter set, including:

[0047] According to the numerical distribution information of the model parameter set, the model parameter set is divided into a plurality of subsets to be quantized, and a target quantization mode of each subset to be quantized is determined respectively;

[0048] For any subset to be quantized, a first parameter in the subset to be quantized is quantized according to a target quantization mode of the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized.

[0049] Among them, the subset to be quantized refers to the parameter set obtained after splitting the model parameter set; the target quantization method refers to the actual quantization method of the subset to be quantized, such as equal-interval quantization, non-equal-interval quantization, clustering, outlier truncation or sparse storage.

[0050] Specifically, since the model parameters of a model are often large-scale, the numerical distribution information of the first parameter in the obtained model parameter set is also complex and different. Therefore, after obtaining the model parameter set to be quantized of the target model, the obtained model parameter set can be split into multiple subsets to be quantized according to the numerical distribution information of the model parameter set, and the target quantization method of each subset to be quantized can be determined respectively, so that in the subsequent process, the subsets to be quantized can be quantized according to the target quantization method corresponding to each subset to be quantized, and the quantized subsets corresponding to each subset to be quantized can be obtained. Thereby, the accuracy of quantizing the model parameter set can be provided, and the efficiency of quantizing the model parameter set can be improved. Further, the implementation process of splitting the model parameter set into multiple subsets to be quantized is explained below.

[0051] In a specific implementation provided by the present disclosure, the model parameter set is split into a plurality of subsets to be quantized according to the numerical distribution information of the model parameter set, including:

[0052] Determining multiple candidate splitting thresholds of subset splitting points of the model parameter set according to the numerical distribution information of the model parameter set;

[0053] Determine a split threshold of the subset split point among the multiple candidate split thresholds;

[0054] The model parameter set is split into a plurality of subsets to be quantized according to a splitting threshold of the subset splitting point.

[0055] Among them, the subset splitting point refers to the splitting point used to split the model parameter set into multiple subsets to be quantized. The number of subset splitting points can be one or more. The candidate splitting threshold refers to multiple selectable values ​​of the subset splitting point. For example, a splitting point of the model parameter set is determined to be point A, and the multiple selectable values ​​corresponding to point A include x1, x2...xn, then point A is the subset splitting point of the model parameter set, and x1, x2...xn are the candidate splitting thresholds of point A. The splitting threshold refers to one of multiple candidate splitting thresholds, which is used to split the model parameter set.

[0056] Specifically, in order to make the quantization error of the quantized target model as small as possible, multiple candidate split thresholds of the subset splitting point of the model parameter set can be determined according to the numerical distribution information of the model parameter set, and the splitting threshold for splitting the model parameter set is determined from the multiple candidate splitting thresholds according to the quantization error of the target model. According to the determined splitting threshold, the model parameter set is split into multiple subsets to be quantized.

[0057] Multiple candidate splitting thresholds for subset splitting points are determined through the numerical distribution information of the model parameter set, and a splitting threshold for splitting the model parameter set is determined from the multiple candidate splitting thresholds. The model parameter set is split into multiple subsets to be quantized according to the splitting threshold, so as to combine the numerical distribution information of the model parameter set and determine the splitting processing of the model parameter set.

[0058] In practical applications, in order to make the quantization error of the quantized target model as small as possible, the calibration data set can be combined to obtain calibration input information, and the split threshold of the subset split point can be determined from multiple candidate split thresholds in combination with the calibration input information. The specific implementation method is as follows:

[0059] In a specific implementation provided by the present disclosure, the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, and the target quantization method includes equal-interval quantization and non-equal-interval quantization;

[0060] Determining the splitting threshold of the subset splitting point from among the multiple candidate splitting thresholds includes:

[0061] Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model;

[0062] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0063] Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values;

[0064] Determining a plurality of calibration output results of the model operator according to the calibration input information of the model operator and the plurality of candidate parameter matrices;

[0065] According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

[0066] Among them, the calibration data set is specifically a subset of the training data set of the target model, which can be obtained by screening from the training data set of the target model. The sample data is the data in the calibration data set. In practical applications, the sample data is determined according to the actual application scenario of the target model. If the target model is used to perform image processing tasks, the sample data can be sample image data; if the target model is used to perform speech processing tasks, the sample data can be sample speech data; if the target model is used to perform text processing tasks, the sample data can be sample text; if the target model is used to perform video processing tasks, the sample data can be sample video, etc. The calibration input information refers to the input information of the model operator to be quantized; the calibration output result refers to the output result of the model operator to be quantized. The original output result refers to the output result of the model operator without model quantization.

[0067] Quantization point refers to the quantization point after the subset to be quantized is quantized at non-equal intervals. The number of quantization points can be one or more. When the number of quantization points is one, it means that all parameters in the subset to be quantized will be quantized to one value. Candidate quantization values ​​refer to multiple selectable values ​​of the quantization point. For example, it is determined that the quantization point after the subset to be quantized is quantized at non-equal intervals is B, and the multiple selectable values ​​corresponding to quantization point B include y1, y2...yN, then the candidate quantization values ​​of quantization point B include y1, y2...yN. The candidate parameter matrix refers to the parameter matrix reconstructed after the parameter matrix before quantization is quantized at equal intervals, quantized at non-equal intervals, and inverse quantized according to the candidate splitting threshold and the candidate quantization values.

[0068] Specifically, the obtained calibration data set is input into the target model, the calibration input information of the model operator to be quantized is determined, and according to the numerical distribution information of the model parameter set, multiple candidate quantization values ​​of the quantization point after the subset to be quantized is subjected to non-equal interval quantization are determined. According to multiple candidate splitting thresholds and multiple candidate quantization values, the parameter matrix before quantization is reconstructed after equal interval quantization, non-equal interval quantization and inverse quantization to obtain multiple candidate parameter matrices. According to the calibration input information of the model operator and the multiple candidate parameter matrices, the multiple calibration output results of the model operator are determined, and according to the multiple calibration output results and the original output results of the sample data, the quantization error after quantizing the target model is determined, and the candidate splitting threshold corresponding to the calibration output result with the smallest quantization error is determined as the splitting threshold of the subset splitting point.

[0069] Further, after constructing multiple candidate parameter matrices of the model operator, based on the calibration input information of the model operator and the candidate parameter matrix, the implementation process of determining the splitting threshold of the subset splitting point from multiple candidate splitting thresholds can be specifically referred to the following formula 1:

[0070]

[0071] Among them, x1 is the candidate split threshold, y1 is the candidate quantization value, and W is the parameter matrix before quantization. is the candidate parameter matrix, and X is the calibration input information. It should be noted that the above formula 1 corresponds to the case where the number of subset splitting points and quantization points are both one.

[0072] In practical applications, the above formula 1 can be solved for x1 and y1 using a non-gradient optimization method or a grid search, so as to determine the splitting threshold of the subset splitting point and the non-equally spaced quantization value of the quantization point. The non-equally spaced quantization value is specifically one or more of the multiple candidate quantization values, and the non-equally spaced quantization value is used to perform non-equally spaced quantization on the subset to be quantized.

[0073] Furthermore, the process of quantizing the parameter matrix before quantization can be referred to the following formula 2:

[0074]

[0075] Where s is the scaling factor, which can usually be the maximum absolute value of W, and m is the quantization bit width. The dequantization process can be seen in the following formula 3:

[0076]

[0077] Furthermore, when the number of subset splitting points is one and the number of quantization points is multiple, the implementation process of determining the splitting threshold of the subset splitting point from multiple candidate splitting thresholds based on the calibration input information of the model operator and the candidate parameter matrix can be specifically referred to the following formula 4:

[0078]

[0079] Based on the same solution method as the above formula 1, the above formula 4 is solved to obtain the splitting threshold of the subset splitting point and multiple non-equally spaced quantization values ​​of the quantization point.

[0080] The model quantization method provided by the embodiment of the present disclosure can combine the numerical distribution information of the calibration data set and the model parameter set, determine the splitting threshold of the subset splitting point, and split the model parameter set according to the splitting threshold, so as to obtain multiple subsets to be quantized, so as to quantize each subset to be quantized according to different quantization methods in the subsequent process. Different quantization methods are used according to the numerical distribution of different model parameters to improve the accuracy of quantizing the model parameter set.

[0081] After determining the splitting threshold of the subset splitting point and splitting the model parameter set to obtain multiple subsets to be quantized, the target quantization method corresponding to each subset to be quantized can be determined.

[0082] Based on this, in a specific implementation provided by the present disclosure, the model parameter set corresponds to a subset splitting point;

[0083] Determine the target quantization method for each subset to be quantized, including:

[0084] When the data absolute value of the first parameter in any subset to be quantized is less than or equal to the first splitting threshold of the subset splitting point, determining that the target quantization mode of the subset to be quantized is equal interval quantization;

[0085] When the data absolute value of the first parameter in any subset to be quantized is greater than the first split threshold, it is determined that the target quantization mode of the subset to be quantized is non-uniformly spaced quantization or sparse storage.

[0086] In practical applications, the first parameter in the model parameter set usually has positive and negative values. Therefore, after determining the splitting threshold of the subset splitting point, it is necessary to determine the target quantization method for each subset to be quantized based on the splitting threshold of the subset splitting point and the absolute value of the data of the first parameter.

[0087] The first splitting threshold refers to the splitting threshold of the subset splitting point determined when the number of subset splitting points is one. Specifically, for any one of the multiple subsets to be quantized, if the data absolute value of the first parameter in the subset to be quantized is less than or equal to the first splitting threshold, it can be determined that the target quantization method of the subset to be quantized is equal interval quantization; if the data absolute value of the first parameter in the subset to be quantized is greater than the first splitting threshold, it can be determined that the target quantization method of the subset to be quantized is non-equal interval quantization or sparse storage.

[0088] That is to say, in the embodiment of the present disclosure, when the number of subset splitting points is one, the model parameter set can be split into two subsets to be quantized according to the first splitting threshold, and the target quantization methods of the two subsets to be quantized are equal-interval quantization and non-equal-interval quantization, or equal-interval quantization and sparse storage.

[0089] Furthermore, taking the model weight parameter as an example, the target quantization method for determining the model parameter set is explained. Figure 4 , Figure 4 FIG. 2 shows a schematic diagram of a method for determining target quantization according to an embodiment of the present disclosure. Figure 4As shown, the numerical distribution law of the weight parameters in the model parameter set is a normal distribution, and the numerical range of the weight parameters in the model parameter set is approximately between -0.07 and 0.07. The data of the weight parameters whose absolute values ​​are less than 0.025 account for a relatively large proportion, and the data of the weight parameters whose absolute values ​​are greater than 0.025 account for a relatively low proportion. Based on this, after determining the first splitting threshold (for example, 0.025) of the subset splitting point, the model parameter set is split into subset 1 to be quantized and subset 2 to be quantized according to the first splitting threshold, and the target quantization method of subset 1 to be quantized is determined to be equal interval quantization, and the target quantization method of subset 2 to be quantized is determined to be non-equal interval quantization.

[0090] Furthermore, in practical applications, the model parameter set can also be split into three categories according to the numerical distribution information of the model parameter set, that is, the model parameter set can be split into three subsets to be quantized. At this time, the target quantization methods of the three subsets to be quantized can be determined as equal interval quantization, non-equal interval quantization and sparse storage respectively. In this case, the calibration input information can also be obtained in combination with the calibration data set, and the splitting threshold of the subset splitting point can be determined from multiple candidate splitting thresholds according to the calibration input information. Since the target quantization method of the subset to be quantized also includes sparse storage, the cost of sparse storage must also be considered in the process of determining the splitting threshold of the subset splitting point.

[0091] Therefore, in a specific implementation provided by the present disclosure, the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, and the target quantization method includes equal-interval quantization, non-equal-interval quantization and sparse storage;

[0092] Determining the splitting threshold of the subset splitting point from among the multiple candidate splitting thresholds includes:

[0093] Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model;

[0094] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0095] Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values;

[0096] Determining multiple calibration output results of the model operator according to the calibration input information of the model operator, the multiple candidate parameter matrices and the cost of the sparse storage;

[0097] According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

[0098] Specifically, the obtained calibration data set is input into the target model, the calibration input information of the model operator to be quantized is determined, and according to the numerical distribution information of the model parameter set, multiple candidate quantization values ​​of the quantization point after the subset to be quantized is quantized at non-equal intervals are determined. According to multiple candidate split thresholds and multiple candidate quantization values, the parameter matrix before quantization is reconstructed after equal interval quantization, non-equal interval quantization and inverse quantization to obtain multiple candidate parameter matrices. According to the calibration input information of the model operator, multiple candidate parameter matrices and the cost of sparse storage, multiple calibration output results of the model operator are determined, and according to the multiple calibration output results and the original output results of the sample data, the quantization error after quantizing the target model is determined, and the candidate split threshold corresponding to the calibration output result with the smallest quantization error is determined as the split threshold of the subset split point.

[0099] Different from the above method of determining the splitting threshold of the subset splitting point, in the process of determining the calibration output result of the model operator, it is necessary to combine the cost of sparse storage to determine the splitting threshold of the subset splitting point.

[0100] Formula 5:

[0101]

[0102] Among them, λ is a constant used to balance and C, where C is the cost of sparse storage. The method of solving Formula 5 is the same as the method of solving Formula 1 above, and the embodiments of the present disclosure will not be repeated here.

[0103] The model quantization method provided by the embodiment of the present disclosure can split the model parameter set according to the numerical distribution information of the model parameter set and the cost of sparse storage, and determine the corresponding target quantization method.

[0104] Based on this, after determining the splitting threshold of the subset splitting point, the target quantization mode of each subset to be quantized can be determined. In a specific embodiment provided by the present disclosure, the model parameter set corresponds to two subset splitting points;

[0105] Determine the target quantization method for each subset to be quantized, including:

[0106] When the data absolute value of the first parameter in any subset to be quantized is less than or equal to the second splitting threshold of the subset splitting point, determining that the target quantization mode of the subset to be quantized is equal interval quantization;

[0107] When the data absolute value of the first parameter in any subset to be quantized is greater than the second split threshold and less than the third split threshold, determining that the target quantization mode of the subset to be quantized is non-equal interval quantization;

[0108] When the data absolute value of the first parameter in any subset to be quantized is greater than or equal to the third splitting threshold, it is determined that the target quantization mode of the subset to be quantized is sparse storage.

[0109] Among them, the second splitting threshold and the third splitting threshold refer to the splitting thresholds of the subset splitting points determined when the number of subset splitting points is two, and the second splitting threshold is less than the third splitting threshold. Specifically, for any one of the multiple subsets to be quantized, if the data absolute value of the first parameter in the subset to be quantized is less than or equal to the second splitting threshold, it can be determined that the target quantization method of the subset to be quantized is equal interval quantization; if the data absolute value of the first parameter in the subset to be quantized is greater than the second splitting threshold and less than the third splitting threshold, it can be determined that the target quantization method of the subset to be quantized is non-equal interval quantization; if the data absolute value of the first parameter in the subset to be quantized is greater than or equal to the third splitting threshold, it can be determined that the target quantization method of the subset to be quantized is sparse storage.

[0110] That is to say, in the embodiment of the present disclosure, when the number of subset splitting points is two, the model parameter set can be split into three subsets to be quantized according to the second splitting threshold and the third splitting threshold, and the target quantization methods of the three subsets to be quantized are equal-interval quantization, non-equal-interval quantization and sparse storage, respectively.

[0111] The model quantization method provided by the embodiment of the present disclosure can split the model parameter set according to the numerical distribution information of the model parameter set to obtain multiple subsets to be quantized, and determine the target quantization method of each subset to be quantized. Different target quantization methods are used for different subsets to be quantized, thereby improving the accuracy of quantizing the model parameter set.

[0112] After determining the target quantization mode corresponding to each subset to be quantized, the subset to be quantized is quantized according to the target quantization mode of each subset to be quantized, so as to obtain the quantized subset.

[0113] The following is an explanation of how to implement equal-interval quantization.

[0114] In a specific implementation provided by the present disclosure, the target quantization method includes equal interval quantization;

[0115] According to the target quantization mode of the subset to be quantized, quantizing the first parameter in the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized, including:

[0116] Determining the interval width according to the numerical distribution information of the model parameter set;

[0117] According to the interval width, the subset to be quantized is divided into a plurality of intervals to be quantized;

[0118] For any interval to be quantized, determining a scaling factor of the interval to be quantized according to an interval boundary value of the interval to be quantized;

[0119] Determine the equally spaced quantized values ​​of the interval to be quantized according to the scaling factor;

[0120] Determine the first parameter of the interval to be quantized as the equally spaced quantized value to obtain a second parameter;

[0121] A quantized subset corresponding to the subset to be quantized is obtained according to the multiple second parameters.

[0122] The interval width is used to divide the subset to be quantized into multiple intervals to be quantized; the interval to be quantized refers to multiple parameter intervals to be quantized that are divided into the subset to be quantized based on the interval width, and the interval to be quantized includes multiple first parameters. The equally spaced quantized value refers to the value of the model parameter after the interval to be quantized is quantized.

[0123] Specifically, after determining that the target quantization mode of the subset to be quantized is equal-interval quantization, the interval width is determined according to the numerical distribution information of the model parameter set or the subset to be quantized. According to the determined interval width, the subset to be quantized is divided into a plurality of equally spaced intervals to be quantized. According to the interval boundary value of each interval to be quantized, the scaling factor of each interval to be quantized is determined, and further the equally spaced quantization value of each interval to be quantized is determined according to the scaling factor and the floating point number of the first parameter. The first parameters of the intervals to be quantized are all quantized into equally spaced quantization values ​​to obtain the second parameters, and the quantization subset can be obtained according to the second parameters of each interval to be quantized.

[0124] Further, in practical applications, the actual equal-interval quantization method can be determined according to the value distribution information of the interval to be quantized. For example, symmetric quantization can be used for the interval to be quantized with symmetric value distribution, and asymmetric quantization can be used for the interval to be quantized with asymmetric value distribution. The specific method can be determined according to the actual application situation, and the present disclosure does not limit it here.

[0125] It should be noted that for the subset to be quantized using equal interval quantization, low-precision floating point quantization may also be used, for example, fp8 quantization, fp4 quantization, etc. In order to improve the accuracy of quantizing the model parameter set and considering the hardware conditions of the terminal device, the embodiment of the present disclosure preferably uses equal interval quantization.

[0126] The following is an explanation of the implementation method of non-uniform interval quantization.

[0127] In a specific implementation provided by the present disclosure, the target quantization method includes non-equal interval quantization;

[0128] According to the target quantization mode of the subset to be quantized, quantizing the first parameter in the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized, including:

[0129] Determining, according to the plurality of calibration output results and the original output results, non-equally spaced quantization values ​​of the quantization points of the non-equally spaced quantization from among the plurality of candidate quantization values;

[0130] respectively determining the distances between the first parameter in the to-be-quantized subset and a plurality of non-equally spaced quantized values;

[0131] quantizing the first parameter in the subset to be quantized to a non-uniformly spaced quantized value closest to the first parameter to obtain a second parameter;

[0132] A quantized subset corresponding to the subset to be quantized is obtained according to the multiple second parameters.

[0133] Specifically, in the process of determining the splitting threshold of the subset splitting point, the non-equally spaced quantization values ​​of the quantization points of the non-equally spaced quantization can also be determined from multiple candidate quantization values ​​based on the calibration output results and the original output results of the model operator. The distances between each first parameter in the subset to be quantized and the determined multiple non-equally spaced quantization values ​​are determined respectively, and the first parameter is quantized to the non-equally spaced quantization value closest to the first parameter to obtain the second parameter. For example, the non-equally spaced quantization values ​​include 0.1 and 0.8, and the first parameter 0.3 is quantized to 0.1. Based on the multiple second parameters obtained, the quantization subset can be obtained.

[0134] It should be noted that if there is only one non-uniformly spaced quantization value, the first parameters in the subset to be quantized will all be quantized to the non-uniformly spaced quantization value.

[0135] In practical applications, when it is determined that the target quantization methods of the subsets to be quantized in the model parameter set are equal-interval quantization and non-equal-interval quantization, each subset to be quantized can be quantized based on the implementation method of equal-interval quantization and the implementation method of non-equal-interval quantization as described above to obtain a quantized subset.

[0136] The following explains how to implement sparse storage.

[0137] In a specific implementation provided by the present disclosure, the target quantization method includes sparse storage;

[0138] According to the target quantization mode of the subset to be quantized, quantizing the first parameter in the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized, including:

[0139] Determining a positive value and a negative value of the splitting threshold;

[0140] Determine the value of the first parameter in the to-be-quantized subset that is greater than the positive value as the positive value;

[0141] Determine the value of the first parameter in the to-be-quantized subset that is smaller than the negative value as the negative value;

[0142] The first parameter and the positive value and the negative value in the subset to be quantized are sparsely stored to obtain a quantized subset corresponding to the subset to be quantized.

[0143] Specifically, when the target quantization method is determined to be sparse storage, the positive value and negative value of the split threshold are determined, the value of the first parameter in the subset to be quantized that is greater than the positive value is determined as the positive value, the value of the first parameter in the subset to be quantized that is less than the negative value is determined as the negative value, and the first parameter and the positive value and the negative value in the subset to be quantized are sparsely stored to obtain the quantized subset. In practical applications, sparse storage methods such as list storage (LIL), dictionary storage (DOK), compressed sparse row storage (CSR), and compressed sparse row storage (CSR) can be used. The present disclosure does not limit the specific implementation method of sparse storage.

[0144] It should be noted that the above-mentioned sparse storage quantization method can be applied to the case where the target quantization methods of the subset to be quantized in the model parameter set are respectively equal-interval quantization and sparse storage. In this case, for the subset to be quantized whose target quantization method is sparse storage, the above-mentioned sparse storage method is executed to obtain the corresponding quantization subset; for the subset to be quantized whose target quantization method is equal-interval quantization, the above-mentioned equal-interval quantization method is executed to obtain the corresponding quantization subset.

[0145] When it is determined that the target quantization methods of the subset to be quantized in the model parameter set are equal-interval quantization, non-equal-interval quantization and sparse storage respectively, for the subset to be quantized whose target quantization method is equal-interval quantization and the subset to be quantized whose target quantization method is non-equal-interval quantization, the equal-interval quantization method and non-equal-interval quantization method as described above are respectively executed, and corresponding quantization subsets are respectively obtained; for the subset to be quantized whose target quantization method is sparse storage, the first parameter in the subset to be quantized can be directly stored sparsely such as list storage, dictionary storage, compressed sparse row storage, compressed sparse row storage, etc., and the stored parameter set is determined as the corresponding quantization subset.

[0146] The model quantization method provided by the embodiment of the present disclosure splits the model parameter set into multiple subsets to be quantized, and adopts different target quantization methods according to the different numerical distributions of each subset to be quantized, so as to realize the fusion of quantization methods such as equal-interval quantization, non-equal-interval quantization, and sparse storage, thereby improving the accuracy of quantization of the target model.

[0147] In practical applications, the model quantization method provided by the embodiment of the present disclosure can also be applied to the processing core of the many-core system, and its specific implementation method is as follows:

[0148] In a specific implementation provided by the present disclosure, the method is applied to a first core group of a many-core system, a target model is deployed in multiple processing cores of the many-core system, the first core group includes at least one processing core of the multiple processing cores, wherein the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, the model operator is deployed in the first core group, the model parameter set is stored in an internal storage space of the first core group, the target quantization method includes equal interval quantization and non-equal interval quantization, wherein, among the multiple candidate split thresholds, determining the split threshold of the subset split point includes:

[0149] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0150] Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values;

[0151] Upon receiving calibration input information of the model operator, determining multiple calibration output results of the model operator according to the calibration input information of the model operator and the multiple candidate parameter matrices, wherein the calibration input information is output by a previous operator of the model operator after sample data in a calibration data set is input into the target model;

[0152] According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

[0153] Among them, the first core group refers to any group of processing cores in the many-core system, and specifically may include one or more of the multiple processing cores of the many-core system. Specifically, the target model is deployed in the multiple processing cores of the many-core system, and the model operators of the target model are deployed in the core group, and each core group can process different model operators in parallel. The parameters of the model operator to be quantized in the target model can be determined as the model parameter set to be quantized, and correspondingly stored in the internal storage space of the first core group. After obtaining the model parameter set of the target model, multiple candidate quantization values ​​of quantization points for non-equal interval quantization are determined according to the numerical distribution information of the model parameter set, and multiple candidate parameter matrices of the model operator are respectively constructed according to the multiple candidate split thresholds and the multiple candidate quantization values. The first core group can receive the calibration output result of the previous level operator of the model operator as the calibration input information of the model operator. After receiving the calibration input information of the model operator, the first core group can determine multiple calibration output results of the model operator according to the calibration input information of the model operator and the constructed multiple candidate parameter matrices. The calibration output result of the model operator can be used as the calibration input information of the next level operator of the model operator, and saved to the processing core group corresponding to the next level operator. Furthermore, according to the multiple calibration output results and the original output results of the model operator, a splitting threshold of the subset splitting point is determined from multiple candidate splitting thresholds.

[0154] According to the splitting threshold of the determined subset splitting point, the model parameter set is split to obtain multiple subsets to be quantized, and the target quantization method of each subset to be quantized is determined respectively. Based on the target quantization method of each subset to be quantized, each subset to be quantized is quantized respectively to obtain the quantized subset corresponding to each subset to be quantized. The specific implementation process of quantizing each subset to be quantized can refer to the implementation process of the above-mentioned quantization methods such as equal interval quantization, non-equal interval quantization and sparse storage, and the present disclosure will not repeat it here.

[0155] Furthermore, according to the multiple calibration output results and the original output results, the splitting threshold of the subset splitting point can be determined from the multiple candidate splitting thresholds, and the non-equally spaced quantization values ​​of the quantization points of the non-equally spaced quantization can be determined from the multiple candidate quantization values. In order to reduce the computational pressure of the processing core, after determining the splitting threshold of the subset splitting point and the non-equally spaced quantization values ​​of the quantization point, the splitting threshold of the subset splitting point and the non-equally spaced quantization values ​​of the quantization point can be sent to other core groups in the multi-core system except the first core group, and the other core groups perform quantization processing on multiple subsets to be quantized of the model parameter set based on the splitting threshold of the subset splitting point and the non-equally spaced quantization values ​​of the quantization point, so as to obtain multiple quantized subsets.

[0156] In practical applications, the model parameter set to be quantized can be determined according to the actual application situation. The model parameter set can specifically be a weight parameter or a weight row in the target model, or a set of weight parameters, etc. Before quantizing the model parameter set, different model parameter sets can be stored in the internal storage space of different core groups respectively, so as to realize parallel quantization of multiple model parameter sets through the many-core system. The multiple core groups of the many-core system provided in the embodiment of the present disclosure can execute the above-mentioned model quantization method in parallel to realize the quantization of the model parameter set and obtain multiple quantized subsets of the model parameter set. Thereby, a large amount of data transmission is avoided, the resource consumption of the many-core system is reduced, and the quantization efficiency is improved.

[0157] Step 206: Determine a quantized target model according to the multiple quantized subsets of the model parameter set.

[0158] After obtaining multiple quantized subsets of the model parameter set, the quantized target model can be determined. Further, in order to ensure the accuracy of the determined quantized target model in the actual application process, after obtaining multiple quantized subsets of the model parameter set, the target model can be trained based on the second parameters in the multiple quantized subsets to obtain a trained target model.

[0159] In practical applications, during the process of equally spaced quantization of the model parameter set of the target model, the data storage efficiency is low. To improve the data storage efficiency during the quantization process, the model parameter set can also be directly quantized at non-equal intervals, and the quantized target model can be determined. The implementation method is as follows:

[0160] Based on this, in a specific implementation provided by the present disclosure, the method further includes:

[0161] Performing non-uniform interval quantization on the first parameter in the model parameter set to obtain a quantization parameter set of the model parameter set;

[0162] A quantized target model is determined according to the quantization parameter set.

[0163] Specifically, the first parameter in the model parameter set may be directly quantized at non-uniform intervals without splitting the model parameter set to obtain a quantization parameter set of the model parameter set, and the quantized target model may be determined according to the quantization parameter set.

[0164] Similar to the implementation method of the above-mentioned non-equal interval quantization, in order to improve the universality and effectiveness of non-equal interval quantization, the model parameter set can be quantized in an non-equal interval manner by using a calibration input.

[0165] In a specific implementation provided by the present disclosure, the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model;

[0166] Performing non-uniform interval quantization on the first parameter in the model parameter set to obtain a quantization parameter set of the model parameter set includes:

[0167] Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model;

[0168] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0169] Generating a plurality of to-be-determined quantization vectors and a plurality of candidate parameter matrices according to the plurality of candidate quantization values;

[0170] Determining a quantization vector from among the plurality of quantization vectors to be determined according to the calibration input information of the model operator and the plurality of candidate parameter matrices;

[0171] The first parameter in the model parameter set is quantized according to the quantization vector to obtain the quantization parameter set.

[0172] Similar to the above-mentioned implementation method of determining the non-equally spaced quantization values ​​of the quantization points, in the embodiment provided by the present disclosure, the obtained calibration data set is input into the target model, the calibration input information of the model operator to be quantized is determined, and according to the numerical distribution information of the model parameter set, multiple candidate quantization values ​​of the quantization points after the model parameter set is quantized non-equally spaced are determined. Based on the multiple candidate quantization values, multiple quantization vectors to be determined are generated, see Figure 5 , Figure 5 FIG. 2 shows a schematic diagram of generating a quantization vector to be determined according to an embodiment of the present disclosure. Figure 5 As shown, quantization points for non-uniformly spaced quantization and multiple candidate quantization values ​​of the quantization points can be determined according to the numerical distribution information of the model parameter set, wherein V1, V2, ..., V7 are all different quantization vectors to be determined.

[0173] When the candidate quantization value is a quantization vector to be determined, the parameter matrix before quantization is reconstructed after non-uniform quantization and corresponding inverse quantization to obtain multiple candidate parameter matrices. According to the calibration input information of the model operator and the multiple candidate parameter matrices, a quantization vector is determined from the multiple quantization vectors to be determined, and the first parameter in the model parameter set is quantized according to the quantization vector to obtain a quantization parameter set corresponding to the model parameter set. Specifically, the implementation process of determining the quantization vector can be referred to the following formula 6:

[0174]

[0175] Among them, V is the vector to be determined, is the candidate parameter matrix reconstructed when the candidate quantization value is the quantization vector to be determined. When solving the vector to be determined in the above formula 6, non-gradient optimization, gradient descent or proxy gradient descent method can be used. The process of iterative solution using the proxy gradient descent method can be referred to the following formula 7:

[0176]

[0177] Wherein, η is the learning rate in the iterative process, and the termination condition of the iteration is that the change of the vector to be determined is less than the preset threshold. The granularity of the candidate quantization value sharing corresponding to the vector to be determined can be the target model, the entire weight matrix, a weight row, or a part of a weight row, which is not limited in the present disclosure.

[0178] After obtaining the calibration input information of the model operator and a plurality of candidate parameter matrices, the calibration output result of the model operator may be further determined, and the quantization vector may be determined based on the calibration output result.

[0179] Therefore, in a specific implementation provided by the present disclosure, determining a quantization vector from the multiple quantization vectors to be determined according to the calibration input information of the model operator and the multiple candidate parameter matrices includes:

[0180] Determining a plurality of calibration output results of the model operator according to the calibration input information of the model operator and the plurality of candidate parameter matrices;

[0181] A quantization vector is determined from among the plurality of quantization vectors to be determined according to the plurality of calibration output results and the original output results of the sample data.

[0182] Specifically, multiple calibration output results of the model operator are determined based on the calibration input information of the model operator and multiple candidate parameter matrices, and the quantization error after quantizing the target model is determined based on the multiple calibration output results and the original output results of the sample data, and the to-be-determined quantization vector corresponding to the calibration output result with the smallest quantization error is determined as the quantization vector, thereby reducing the quantization error of the target model.

[0183] After determining the quantization vector, the first parameter in the model parameter set can be quantized according to the quantization vector to obtain the quantization parameter set. Specifically, the nearest rounding method can be used, that is, by determining the distance between the quantization value of each data element in the quantization vector and each first parameter, each first parameter is quantized to the quantization value of the data element closest to the first parameter. The first parameter in the model parameter set can also be quantized by using a random rounding method.

[0184] In addition to the above method, the embodiment of the present disclosure also provides the following method to quantize the first parameter in the model parameter set to obtain a quantization parameter set.

[0185] In a specific implementation provided by the present disclosure, quantizing the first parameter in the model parameter set according to the quantization vector to obtain the quantization parameter set includes:

[0186] Determining a reconstruction parameter matrix corresponding to the quantization vector from among the multiple candidate parameter matrices;

[0187] Mapping the matrix elements in the reconstruction parameter matrix to the quantization vector to obtain a mapped parameter matrix;

[0188] The quantization parameter set is determined according to the mapped parameter matrix.

[0189] The reconstruction parameter matrix is ​​specifically one of the candidate parameter matrices. Specifically, after determining the quantization vector, the reconstruction parameter matrix corresponding to the quantization vector can be determined from multiple candidate parameter matrices, and the matrix elements in the reconstruction parameter matrix are mapped to the quantization vector to obtain the mapped parameter matrix, and the quantization values ​​in the mapped parameter matrix are determined as the quantization parameter set.

[0190] The candidate quantization values ​​are solved according to the calibration data set to determine the quantization parameter set corresponding to the model parameter set, thereby improving the universality and effectiveness of non-uniform interval quantization.

[0191] The model quantization method provided by the present disclosure includes: obtaining a model parameter set to be quantized, wherein the model parameter set includes multiple first parameters in a target model; quantizing the first parameters in the model parameter set according to numerical distribution information of the model parameter set to obtain multiple quantization subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of a numerical range, a numerical distribution law, and a numerical proportion of the multiple first parameters, and the numerical proportion is used to characterize the ratio between the number of first parameters corresponding to any numerical value within the numerical range and the total number of first parameters; determining the quantized target model according to the multiple quantization subsets of the model parameter set.

[0192] According to the embodiments of the present disclosure, the model parameter set can be split into multiple subsets to be quantized according to the numerical distribution information of the model parameter set to be quantized, and the target quantization methods of the multiple subsets to be quantized can be determined respectively, so that each subset to be quantized can be quantized according to the target quantization method corresponding to each subset to be quantized, so as to obtain the quantized subsets corresponding to each subset to be quantized, and achieve the accuracy of quantizing the model parameter set. By quantizing the subsets to be quantized with different target quantization methods for each subset to be quantized, the storage amount of the model parameter set can be reduced, thereby reducing the occupation of memory space by the parameters of the quantized target model; by quantizing each subset to be quantized, the accuracy of the first parameter in each subset to be quantized can be reduced, thereby reducing the computing pressure and computing resources of the processor or processing core during the operation of the quantized target model.

[0193] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not repeat them. It can be understood by those skilled in the art that in the above-mentioned method of the specific implementation method, the specific execution order of each step should be determined according to its function and possible internal logic.

[0194] In addition, the present disclosure also provides a model quantization device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any model quantization method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to in the corresponding records of the method part and will not be repeated here.

[0195] Figure 6 It is a block diagram of a model quantization device provided by an embodiment of the present disclosure.

[0196] Reference Figure 6 , the embodiment of the present disclosure provides a model quantization device, the model quantization device comprising:

[0197] An acquisition module 602 is configured to acquire a model parameter set to be quantized, wherein the model parameter set includes a plurality of first parameters in a target model;

[0198] A quantization module 604 is configured to quantize the first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of a numerical range, a numerical distribution law, and a numerical proportion of the plurality of first parameters, and the numerical proportion is used to characterize a ratio between the number of first parameters corresponding to any numerical value in the numerical range and the total number of first parameters;

[0199] The determination module 606 is configured to determine a quantized target model according to the multiple quantized subsets of the model parameter set.

[0200] Optionally, the quantization module 604 is further configured to:

[0201] According to the numerical distribution information of the model parameter set, the model parameter set is divided into a plurality of subsets to be quantized, and a target quantization mode of each subset to be quantized is determined respectively;

[0202] For any subset to be quantized, a first parameter in the subset to be quantized is quantized according to a target quantization mode of the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized.

[0203] Optionally, the quantization module 604 is further configured to:

[0204] Determining multiple candidate splitting thresholds of subset splitting points of the model parameter set according to the numerical distribution information of the model parameter set;

[0205] Determine a split threshold of the subset split point among the multiple candidate split thresholds;

[0206] The model parameter set is split into a plurality of subsets to be quantized according to a splitting threshold of the subset splitting point.

[0207] Optionally, the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, and the target quantization method includes equal-interval quantization and non-equal-interval quantization;

[0208] The quantization module 604 is further configured to:

[0209] Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model;

[0210] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0211] Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values;

[0212] Determining a plurality of calibration output results of the model operator according to the calibration input information of the model operator and the plurality of candidate parameter matrices;

[0213] According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

[0214] Optionally, the model parameter set corresponds to a subset split point;

[0215] The quantization module 604 is further configured to:

[0216] When the data absolute value of the first parameter in any subset to be quantized is less than or equal to the first splitting threshold of the subset splitting point, determining that the target quantization mode of the subset to be quantized is equal interval quantization;

[0217] When the data absolute value of the first parameter in any subset to be quantized is greater than the first split threshold, it is determined that the target quantization mode of the subset to be quantized is non-uniformly spaced quantization or sparse storage.

[0218] Optionally, the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, and the target quantization method includes equal-interval quantization, non-equal-interval quantization and sparse storage;

[0219] The quantization module 604 is further configured to:

[0220] Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model;

[0221] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0222] Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values;

[0223] Determining multiple calibration output results of the model operator according to the calibration input information of the model operator, the multiple candidate parameter matrices and the cost of the sparse storage;

[0224] According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

[0225] Optionally, the model parameter set corresponds to two subset splitting points;

[0226] The quantization module 604 is further configured to:

[0227] When the data absolute value of the first parameter in any subset to be quantized is less than or equal to the second splitting threshold of the subset splitting point, determining that the target quantization mode of the subset to be quantized is equal interval quantization;

[0228] When the data absolute value of the first parameter in any subset to be quantized is greater than the second split threshold and less than the third split threshold, determining that the target quantization mode of the subset to be quantized is non-equal interval quantization;

[0229] When the data absolute value of the first parameter in any subset to be quantized is greater than or equal to the third splitting threshold, it is determined that the target quantization mode of the subset to be quantized is sparse storage.

[0230] Optionally, the target quantization method includes equal interval quantization;

[0231] The quantization module 604 is further configured to:

[0232] Determining the interval width according to the numerical distribution information of the model parameter set;

[0233] According to the interval width, the subset to be quantized is divided into a plurality of intervals to be quantized;

[0234] For any interval to be quantized, determining a scaling factor of the interval to be quantized according to an interval boundary value of the interval to be quantized;

[0235] Determine the equally spaced quantized values ​​of the interval to be quantized according to the scaling factor;

[0236] Determine the first parameter of the interval to be quantized as the equally spaced quantized value to obtain a second parameter;

[0237] A quantized subset corresponding to the subset to be quantized is obtained according to the multiple second parameters.

[0238] Optionally, the target quantization method includes non-equal interval quantization;

[0239] The quantization module 604 is further configured to:

[0240] Determining, according to the plurality of calibration output results and the original output results, non-equally spaced quantization values ​​of the quantization points of the non-equally spaced quantization from among the plurality of candidate quantization values;

[0241] respectively determining the distances between the first parameter in the to-be-quantized subset and a plurality of non-equally spaced quantized values;

[0242] quantizing the first parameter in the subset to be quantized to a non-uniformly spaced quantized value closest to the first parameter to obtain a second parameter;

[0243] A quantized subset corresponding to the subset to be quantized is obtained according to the multiple second parameters.

[0244] Optionally, the target quantization method includes sparse storage;

[0245] The quantization module 604 is further configured to:

[0246] Determining a positive value and a negative value of the splitting threshold;

[0247] Determine the value of the first parameter in the to-be-quantized subset that is greater than the positive value as the positive value;

[0248] Determine the value of the first parameter in the to-be-quantized subset that is smaller than the negative value as the negative value;

[0249] The first parameter and the positive value and the negative value in the subset to be quantized are sparsely stored to obtain a quantized subset corresponding to the subset to be quantized.

[0250] Optionally, the apparatus is applied to a first core group of a many-core system, a plurality of processing cores of the many-core system having a target model deployed therein, the first core group including at least one processing core of the plurality of processing cores,

[0251] The first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, the model operator is deployed in the first core group, the model parameter set is stored in an internal storage space of the first core group, and the target quantization method includes equal-interval quantization and non-equal-interval quantization.

[0252] Wherein, the quantization module 604 is further configured as follows:

[0253] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0254] Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values;

[0255] Upon receiving calibration input information of the model operator, determining multiple calibration output results of the model operator according to the calibration input information of the model operator and the multiple candidate parameter matrices, wherein the calibration input information is output by a previous operator of the model operator after sample data in a calibration data set is input into the target model;

[0256] According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

[0257] Optionally, the device further includes a non-equal interval quantization module, configured to:

[0258] Performing non-uniform interval quantization on the first parameter in the model parameter set to obtain a quantization parameter set of the model parameter set;

[0259] A quantized target model is determined according to the quantization parameter set.

[0260] Optionally, the first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model;

[0261] The non-equal interval quantization module is further configured as follows:

[0262] Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model;

[0263] Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization;

[0264] Generating a plurality of to-be-determined quantization vectors and a plurality of candidate parameter matrices according to the plurality of candidate quantization values;

[0265] Determining a quantization vector from among the plurality of quantization vectors to be determined according to the calibration input information of the model operator and the plurality of candidate parameter matrices;

[0266] The first parameter in the model parameter set is quantized according to the quantization vector to obtain the quantization parameter set.

[0267] Optionally, the non-equal interval quantization module is further configured to:

[0268] Determining a plurality of calibration output results of the model operator according to the calibration input information of the model operator and the plurality of candidate parameter matrices;

[0269] A quantization vector is determined from among the plurality of quantization vectors to be determined according to the plurality of calibration output results and the original output results of the sample data.

[0270] Optionally, the non-equal interval quantization module is further configured to:

[0271] Determining a reconstruction parameter matrix corresponding to the quantization vector from among the multiple candidate parameter matrices;

[0272] Mapping the matrix elements in the reconstruction parameter matrix to the quantization vector to obtain a mapped parameter matrix;

[0273] The quantization parameter set is determined according to the mapped parameter matrix.

[0274] Optionally, the target model is used to perform any one of an image processing task, a speech processing task, a text processing task, and a video processing task.

[0275] The model quantization device provided by the present disclosure includes: an acquisition module, configured to acquire a model parameter set to be quantized, wherein the model parameter set includes multiple first parameters in a target model; a quantization module, configured to quantize the first parameters in the model parameter set according to the numerical distribution information of the model parameter set, and obtain multiple quantization subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of the numerical range, numerical distribution law and numerical proportion of the multiple first parameters, and the numerical proportion is used to characterize the ratio between the number of first parameters corresponding to any numerical value within the numerical range and the total number of first parameters; a determination module, configured to determine the quantized target model according to the multiple quantization subsets of the model parameter set.

[0276] The embodiments of the present disclosure can split the model parameter set into multiple subsets to be quantized according to the numerical distribution information of the model parameter set to be quantized, and determine the target quantization methods of the multiple subsets to be quantized respectively, so that each subset to be quantized can be quantized according to the target quantization method corresponding to each subset to be quantized, so as to obtain the quantized subsets corresponding to each subset to be quantized, and achieve the accuracy of quantizing the model parameter set. By quantizing the subsets to be quantized with different target quantization methods for each subset to be quantized, the storage amount of the model parameter set can be reduced, thereby reducing the memory space occupied by the parameters of the quantized target model; by quantizing each subset to be quantized, the accuracy of the first parameter in each subset to be quantized can be reduced, thereby reducing the computing pressure and computing resources of the processor or processing core during the operation of the quantized target model.

[0277] Figure 7 It is a block diagram of an electronic device provided by an embodiment of the present disclosure.

[0278] Reference Figure 7 An embodiment of the present disclosure provides an electronic device, which includes multiple processing cores 701 and an on-chip network 702, wherein the multiple processing cores 701 are connected to the on-chip network 702, and the on-chip network 702 is used to exchange data between the multiple processing cores and external data.

[0279] One or more instructions are stored in one or more processing cores 701 , and the one or more instructions are executed by one or more processing cores 701 , so that the one or more processing cores 701 can execute the above-mentioned model quantization method.

[0280] In some embodiments, the electronic device may be a brain-like chip, which can use vectorized computing and needs to load the weight information and other parameters of the neural network model through an external memory such as a double data rate (DDR) synchronous dynamic random access memory. Therefore, the batch processing used in the embodiments of the present disclosure has a higher computing efficiency.

[0281] The present disclosure also provides a computer-readable storage medium on which a computer program is stored, wherein the computer program implements the above-mentioned model quantization method when executed by a processing core. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0282] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned model quantization method.

[0283] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium).

[0284] As is known to those of ordinary skill in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, it is known to those of ordinary skill in the art that communication media typically contain computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and may include any information delivery medium.

[0285] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0286] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0287] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK), etc.

[0288] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0289] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0290] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0291] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0292] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for limiting purposes. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly noted, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, those skilled in the art will appreciate that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A model quantization method, characterized in that: include: Acquire a model parameter set to be quantized, wherein the model parameter set includes a plurality of first parameters in a target model; quantizing a first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of a numerical range, a numerical distribution law, and a numerical proportion of the plurality of first parameters, and the numerical proportion is used to characterize a ratio between the number of first parameters corresponding to any numerical value within the numerical range and the total number of first parameters; A quantized target model is determined according to the multiple quantized subsets of the model parameter set.

2. The method according to claim 1, characterized in that quantizing a first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set, including: According to the numerical distribution information of the model parameter set, the model parameter set is divided into a plurality of subsets to be quantized, and a target quantization mode of each subset to be quantized is determined respectively; For any subset to be quantized, a first parameter in the subset to be quantized is quantized according to a target quantization mode of the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized.

3. The method according to claim 2, characterized in that According to the numerical distribution information of the model parameter set, the model parameter set is divided into a plurality of subsets to be quantized, including: Determining multiple candidate splitting thresholds of subset splitting points of the model parameter set according to the numerical distribution information of the model parameter set; Determine a split threshold of the subset split point among the multiple candidate split thresholds; The model parameter set is split into a plurality of subsets to be quantized according to a splitting threshold of the subset splitting point.

4. The method according to claim 3, characterized in that The first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, and the target quantization method includes equal-interval quantization and non-equal-interval quantization; Determining the splitting threshold of the subset splitting point from among the multiple candidate splitting thresholds includes: Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model; Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization; Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values; Determining a plurality of calibration output results of the model operator according to the calibration input information of the model operator and the plurality of candidate parameter matrices; According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

5. The method according to claim 3, characterized in that The model parameter set corresponds to a subset splitting point; Determine the target quantization method for each subset to be quantized, including: When the data absolute value of the first parameter in any subset to be quantized is less than or equal to the first splitting threshold of the subset splitting point, determining that the target quantization mode of the subset to be quantized is equal interval quantization; When the data absolute value of the first parameter in any subset to be quantized is greater than the first split threshold, it is determined that the target quantization mode of the subset to be quantized is non-uniformly spaced quantization or sparse storage.

6. The method according to claim 3, characterized in that The first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, and the target quantization method includes equal-interval quantization, non-equal-interval quantization and sparse storage; Determining the splitting threshold of the subset splitting point from among the multiple candidate splitting thresholds includes: Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model; Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization; Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values; Determining multiple calibration output results of the model operator according to the calibration input information of the model operator, the multiple candidate parameter matrices and the cost of the sparse storage; According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

7. The method according to claim 3, characterized in that The model parameter set corresponds to two subset splitting points; Determine the target quantization method for each subset to be quantized, including: When the data absolute value of the first parameter in any subset to be quantized is less than or equal to the second splitting threshold of the subset splitting point, determining that the target quantization mode of the subset to be quantized is equal interval quantization; When the data absolute value of the first parameter in any subset to be quantized is greater than the second split threshold and less than the third split threshold, determining that the target quantization mode of the subset to be quantized is non-equal interval quantization; When the data absolute value of the first parameter in any subset to be quantized is greater than or equal to the third splitting threshold, it is determined that the target quantization mode of the subset to be quantized is sparse storage.

8. The method according to claim 5 or 7, characterized in that The target quantization method includes equal interval quantization; According to the target quantization mode of the subset to be quantized, quantizing the first parameter in the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized, including: Determining the interval width according to the numerical distribution information of the model parameter set; According to the interval width, the subset to be quantized is divided into a plurality of intervals to be quantized; For any interval to be quantized, determining a scaling factor of the interval to be quantized according to an interval boundary value of the interval to be quantized; Determine the equally spaced quantized values ​​of the interval to be quantized according to the scaling factor; Determine the first parameter of the interval to be quantized as the equally spaced quantized value to obtain a second parameter; A quantized subset corresponding to the subset to be quantized is obtained according to the multiple second parameters.

9. The method according to claim 4 or 6, characterized in that The target quantization method includes non-equal interval quantization; According to the target quantization mode of the subset to be quantized, quantizing the first parameter in the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized, including: Determining, according to the plurality of calibration output results and the original output results, non-equally spaced quantization values ​​of the quantization points of the non-equally spaced quantization from among the plurality of candidate quantization values; respectively determining the distances between the first parameter in the to-be-quantized subset and a plurality of non-equally spaced quantized values; quantizing the first parameter in the subset to be quantized to a non-uniformly spaced quantized value closest to the first parameter to obtain a second parameter; A quantized subset corresponding to the subset to be quantized is obtained according to the multiple second parameters.

10. The method according to claim 3, characterized in that The target quantization method includes sparse storage; According to the target quantization mode of the subset to be quantized, quantizing the first parameter in the subset to be quantized to obtain a quantized subset corresponding to the subset to be quantized, including: Determining a positive value and a negative value of the splitting threshold; Determine the value of the first parameter in the to-be-quantized subset that is greater than the positive value as the positive value; Determine the value of the first parameter in the to-be-quantized subset that is smaller than the negative value as the negative value; The first parameter and the positive value and the negative value in the subset to be quantized are sparsely stored to obtain a quantized subset corresponding to the subset to be quantized.

11. The method according to claim 3, characterized in that The method is applied to a first core group of a many-core system, wherein a target model is deployed in a plurality of processing cores of the many-core system, and the first core group includes at least one processing core of the plurality of processing cores. The first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model, the model operator is deployed in the first core group, the model parameter set is stored in an internal storage space of the first core group, and the target quantization method includes equal-interval quantization and non-equal-interval quantization. Among the multiple candidate splitting thresholds, determining the splitting threshold of the subset splitting point includes: Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization; Constructing a plurality of candidate parameter matrices of the model operator respectively according to the plurality of candidate splitting thresholds and the plurality of candidate quantization values; Upon receiving calibration input information of the model operator, determining multiple calibration output results of the model operator according to the calibration input information of the model operator and the multiple candidate parameter matrices, wherein the calibration input information is output by a previous operator of the model operator after sample data in a calibration data set is input into the target model; According to the multiple calibration output results and the original output results of the sample data, a splitting threshold of a subset splitting point is determined from the multiple candidate splitting thresholds.

12. The method according to claim 1, characterized in that The method further comprises: Performing non-uniform interval quantization on the first parameter in the model parameter set to obtain a quantization parameter set of the model parameter set; A quantized target model is determined according to the quantization parameter set.

13. The method according to claim 12, characterized in that The first parameter in the model parameter set is a parameter of a model operator to be quantized in the target model; Performing non-uniform interval quantization on the first parameter in the model parameter set to obtain a quantization parameter set of the model parameter set includes: Inputting sample data in a calibration data set into the target model to determine calibration input information of the model operator, wherein the calibration data set is obtained by screening a training data set of the target model; Determining, according to the numerical distribution information of the model parameter set, a plurality of candidate quantization values ​​corresponding to the quantization points of the non-equally spaced quantization; Generating a plurality of to-be-determined quantization vectors and a plurality of candidate parameter matrices according to the plurality of candidate quantization values; Determining a quantization vector from among the plurality of quantization vectors to be determined according to the calibration input information of the model operator and the plurality of candidate parameter matrices; The first parameter in the model parameter set is quantized according to the quantization vector to obtain the quantization parameter set.

14. The method according to claim 13, characterized in that Determining a quantization vector from the plurality of quantization vectors to be determined according to the calibration input information of the model operator and the plurality of candidate parameter matrices, comprising: Determining a plurality of calibration output results of the model operator according to the calibration input information of the model operator and the plurality of candidate parameter matrices; A quantization vector is determined from among the plurality of quantization vectors to be determined according to the plurality of calibration output results and the original output results of the sample data.

15. The method according to claim 13, characterized in that Quantizing the first parameter in the model parameter set according to the quantization vector to obtain the quantization parameter set includes: Determining a reconstruction parameter matrix corresponding to the quantization vector from among the multiple candidate parameter matrices; Mapping the matrix elements in the reconstruction parameter matrix to the quantization vector to obtain a mapped parameter matrix; The quantization parameter set is determined according to the mapped parameter matrix.

16. The method according to claim 1, wherein: The target model is used to perform any one of image processing tasks, speech processing tasks, text processing tasks, and video processing tasks.

17. A model quantization device, characterized in that: include: An acquisition module is configured to acquire a model parameter set to be quantized, wherein the model parameter set includes a plurality of first parameters in a target model; a quantization module, configured to quantize a first parameter in the model parameter set according to the numerical distribution information of the model parameter set to obtain a plurality of quantized subsets corresponding to the model parameter set, wherein the numerical distribution information includes at least one of a numerical range, a numerical distribution law, and a numerical proportion of the plurality of first parameters, and the numerical proportion is used to characterize a ratio between the number of first parameters corresponding to any numerical value in the numerical range and the total number of first parameters; The determination module is configured to determine a quantized target model according to multiple quantized subsets of the model parameter set.

18. An electronic device, characterized in that: include: Multiple processing cores; as well as An on-chip network is configured to exchange data between the multiple processing cores and external data; wherein one or more instructions are stored in one or more of the processing cores, and one or more of the instructions are executed by one or more of the processing cores, so that one or more of the processing cores can execute the model quantization method as described in any one of claims 1-16.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processing core, it implements the model quantization method as described in any one of claims 1 to 16.