Method and device for quantizing data, equipment and medium

By extracting and mapping matrix vectors in the machine learning model and utilizing the objective function and mapping parameters, the problem of low data compression accuracy in the existing technology is solved, and high-precision data compression is achieved at a low bit number, thereby improving the accuracy and efficiency of the model.

CN120688653APending Publication Date: 2025-09-23BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410323179.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing quantization technology solutions make it difficult to achieve effective data compression while ensuring data accuracy when compressing machine learning models, especially when the compressed data accuracy is unsatisfactory at extremely low bit numbers.

Method used

By extracting multiple first vectors from the matrix of the machine learning model, creating an associated objective function, and mapping the second vector to the third vector using mapping parameters, ensuring that the difference between the third vector and the first vector meets a predetermined condition, data width compression is achieved.

Benefits of technology

Maintaining high data precision at lower data width improves the accuracy of machine learning models and the efficiency of quantization operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688653A_ABST
    Figure CN120688653A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device, equipment and a medium for quantizing data. In one method, a plurality of first vectors are extracted from a matrix to be quantized. Creating a plurality of objective functions respectively associated with the plurality of first vectors, the plurality of objective functions respectively comprising the plurality of first vectors, a plurality of second vectors respectively corresponding to the plurality of first vectors, and mapping parameters for respectively mapping the plurality of second vectors to a plurality of third vectors, a second data width corresponding to the plurality of second vectors is less than a first data width corresponding to the plurality of first vectors. A plurality of second vectors and a mapping parameter are determined based on the plurality of objective functions, for a first vector of the plurality of first vectors, the mapping parameter causes a difference between a third vector of the plurality of third vectors corresponding to the first vector and the first vector to meet a predetermined condition. The data may be represented with a lower data width, and the accuracy of the quantization operation may be improved, thereby improving the accuracy of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary implementations of the present disclosure relate generally to data compression, and more particularly to methods, apparatuses, devices, and computer-readable storage media for quantizing data. Background Art

[0002] Machine learning technology has been widely used in multiple application environments. Machine learning models involve a large number of parameters, which leads to a large consumption of resources during the inference phase. Currently, various quantization techniques have been proposed for compressing machine learning models. For example, the data in the machine learning model can be compressed from a higher number of bits to a lower number of bits while ensuring data accuracy. However, the accuracy of the compressed data obtained with existing quantization technology solutions is not satisfactory, and therefore it is desired to provide a more efficient data quantization method. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for quantizing data is provided. In the method, multiple first vectors are extracted from a matrix to be quantized. Multiple objective functions are created, each associated with the multiple first vectors. The multiple objective functions respectively include multiple first vectors, multiple second vectors corresponding to the multiple first vectors, and mapping parameters. The mapping parameters are used to map the multiple second vectors to multiple third vectors, respectively. The second data width corresponding to the multiple second vectors is smaller than the first data width corresponding to the multiple first vectors. Based on the multiple objective functions, multiple second vectors and mapping parameters are determined. For a first vector among the multiple first vectors, the mapping parameters are such that the difference between the third vector corresponding to the first vector among the multiple third vectors and the first vector satisfies a predetermined condition.

[0004] In a second aspect of the present disclosure, a device for quantizing data is provided. The device includes: an extraction module configured to extract multiple first vectors from a matrix to be quantized; a creation module configured to create multiple objective functions respectively associated with the multiple first vectors, the multiple objective functions respectively including multiple first vectors, multiple second vectors respectively corresponding to the multiple first vectors, and mapping parameters, the mapping parameters being used to respectively map the multiple second vectors to multiple third vectors, wherein the second data width corresponding to the multiple second vectors is smaller than the first data width corresponding to the multiple first vectors; and a determination module configured to determine the multiple second vectors and the mapping parameters based on the multiple objective functions, wherein, for a first vector among the multiple first vectors, the mapping parameters ensure that the difference between the third vector corresponding to the first vector among the multiple third vectors and the first vector satisfies a predetermined condition.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.

[0007] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0009] Figure 1 A block diagram illustrating an application environment according to an exemplary implementation of the present disclosure is shown;

[0010] Figure 2 A block diagram for quantizing data according to some implementations of the present disclosure is shown;

[0011] Figure 3 A block diagram illustrating a machine learning model according to some implementations of the present disclosure is shown;

[0012] Figure 4 A block diagram illustrating quantization and dequantization processes according to some implementations of the present disclosure is shown;

[0013] Figure 5 A flowchart illustrating a method for quantifying data according to some implementations of the present disclosure is shown;

[0014] Figure 6 A block diagram illustrating an apparatus for quantizing data according to some implementations of the present disclosure; and

[0015] Figure 7 A block diagram is shown of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION

[0016] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0017] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.

[0018] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0019] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0020] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0023] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.

[0024] Sample Environment

[0025] Machine learning technology has been widely used in many application environments. Machine learning models involve a large number of parameters, which consumes a lot of resources during the inference phase. Currently, various quantization techniques have been proposed for compressing machine learning models. For example, the data in machine learning models can be compressed from a higher number of bits to a lower number of bits while ensuring data accuracy. Figure 1 Describes an application environment according to an example implementation of the present disclosure. Figure 1 A block diagram 100 illustrating an application environment according to an exemplary implementation of the present disclosure is shown.

[0026] like Figure 1 As shown, data 110 may include multiple bits (e.g., a width of 112). To reduce the storage space occupied by data 110, a quantization process may be performed using parameters 120 to convert data 110 into data 130 having a smaller width 132. Furthermore, during the inverse quantization process, parameters 140 may be used to restore data 130 to data 150. Here, parameters 120 and 140 may ensure that the difference between data 110 and 150 is not too large (e.g., satisfying a predetermined condition). Thus, the quantization and inverse quantization processes may reduce resource consumption involved in data storage, transmission, and use while ensuring that data accuracy meets expectations.

[0027] With the increasing popularity of machine learning models, they have been applied to a variety of industrial fields. Machine learning models (especially large models) typically have a large number of parameters and involve a huge amount of computation, which will consume a huge amount of resources during deployment and inference.

[0028] In the field of compression of machine learning models, a technical solution called Post-Training Quantization (PTQ) has been proposed. PTQ does not require training of the model, but only requires a few samples for calibration. This makes quantization simple and feasible, speeds up the iteration cycle and reduces the complexity of downstream processing. Although certain successes have been achieved in the field of PTQ quantization, it is difficult to achieve the desired accuracy in the case of extremely low bits (for example, 2 bits, or 3 bits, etc.). The compression rate and / or the accuracy of the compressed data of existing quantization technical solutions are not satisfactory, and it is therefore desired to provide a more effective data quantization method.

[0029] Overview of Data Quantification

[0030] In order to at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for quantizing data is proposed. For ease of description, in the context of the present disclosure, a matrix will be used as a specific example of data to describe more details of performing data quantization. Figure 2 An overview of an exemplary implementation of the present disclosure is described below. Figure 2 A block diagram 200 for quantizing data is shown, according to some implementations of the present disclosure.

[0031] like Figure 2 As shown, the matrix 210 to be quantized may have a first dimension (e.g., dimension 212) and a second dimension (e.g., dimension 214). The first dimension may, for example, represent rows in the matrix, and the second dimension may represent columns in the matrix. Alternatively and / or additionally, the first dimension may, for example, represent columns in the matrix, and the second dimension may represent rows in the matrix.

[0032] Multiple first vectors can be extracted from the matrix 210 to be quantized. For ease of description, the first vector can represent each column in the matrix. Alternatively and / or additionally, when the rows and columns in the matrix are swapped, the first vector can represent each row in the matrix. Furthermore, multiple objective functions can be created, each associated with the multiple first vectors. Here, the multiple objective functions each include the multiple first vectors, the multiple second vectors corresponding to the multiple first vectors, and mapping parameters, and the mapping parameters (e.g., parameters 232) are used to map the multiple second vectors to the multiple third vectors.

[0033] Specifically, for the first vector 220, the objective function 222 may include: the first vector 220, the second vector 230 corresponding to the first vector 220, and the parameter 232. Here, the parameter 232 may map the second vector 230 to the third vector 234. It should be understood that due to Figure 2Involving a quantization process, the second data width corresponding to the plurality of second vectors is smaller than the first data width corresponding to the plurality of first vectors. For example, the first data width may be 128 bits, 64 bits (or other values), and the second data width may be 32 bits, 16 bits, 8 bits, 4 bits, or even 2 bits.

[0034] It should be understood that although Figure 2 Only objective function 222 corresponding to first vector 220 is shown. Each column vector in matrix 210 can have its own objective function, and multiple objective functions may exist. Furthermore, multiple second vectors and mapping parameters can be determined based on the multiple objective functions. Specifically, multiple second vectors and mapping parameters that meet the requirements can be found by solving the multiple objective functions. In this case, for a first vector among the multiple first vectors, the mapping parameters ensure that the difference between a third vector among the multiple third vectors corresponding to the first vector and the first vector satisfies a predetermined condition.

[0035] In other words, using Figure 2 The process shown utilizes mapping parameters to restore multiple second vectors to multiple third vectors with greater widths, where the differences between the third vectors and the corresponding first vectors satisfy predetermined conditions. This means that the restored third vectors still have high data precision and result in smaller errors. This allows data to be represented using a lower data width, improving the precision of quantization operations and, consequently, enhancing the accuracy of machine learning models.

[0036] Detailed process of data quantification

[0037] An overview of an example implementation according to the present disclosure has been described, and more details about data quantization will be described below. According to an example implementation of the present disclosure, the matrix described above can be a weight matrix of a network layer in a machine learning model, and the number of first dimensions of the matrix is ​​determined by the width of the input data of the network layer, and the number of second dimensions is determined by the width of the output data of the network layer. Using the example implementation of the present disclosure, the weights of the machine learning model can be compressed in a more efficient manner, thereby reducing the various resource overheads involved in the operation of the machine learning model.

[0038] Figure 3 A block diagram 300 of a machine learning model according to some implementations of the present disclosure is shown. Figure 3As shown, the machine learning model 310 may include multiple network layers 312, ..., 314, ..., and 316. Here, each network layer may have a corresponding weight matrix, and the weight matrix of each network layer may be processed using the quantization process described above. Specifically, the initial weight matrix may be represented as W0, and the dimension of the matrix may be based on the width d of the input data. in and the width d of the output data out To represent. At this time, the dimension of W0 is represented by d in ×d out ,and By using the exemplary implementation of the present disclosure, corresponding quantization methods can be determined based on the formats of input data and output data of different network layers. In this way, the quantization accuracy and efficiency of each network layer can be improved.

[0039] According to an example implementation of the present disclosure, multiple first vectors correspond to a floating-point data space, and multiple second vectors correspond to an integer data space. Using this example implementation, data originally represented as floating-point data can be mapped to an integer data space, thereby reducing various resource overheads of a machine learning model through a quantization process.

[0040] According to an exemplary implementation of the present disclosure, the mapping parameters may include: a zero value parameter and a scaling parameter for mapping from an integer data space to a floating point data space. Figure 4 Describe the overall process of quantization. Figure 4 A block diagram 400 is shown of the quantization and dequantization process according to some implementations of the present disclosure. Figure 4 As shown, the matrix 210 can be mapped to the quantized matrix 410 through a quantization operation. Specifically, the following formula can be used:

[0041]

[0042] In the above formula, Represents the quantized matrix, clip() represents the truncation operation, α and β represent the lower and upper thresholds of the truncation operation respectively, Denotes a rounding operation, W0 denotes the initial matrix, z denotes the zero point parameter (i.e., zero point) used to perform the mapping operation from the floating-point data space to the integer data space, and s denotes the scaling parameter (i.e., scale) used to perform the mapping operation. The above formula maps data in the floating-point data space to the integer data space, and α and β denote the minimum and maximum values ​​represented by the integer data space, respectively.

[0043] Furthermore, in the inverse quantization process, the quantized matrix 410 may be converted into an inverse quantized matrix 420. Specifically, the following formula may be used:

[0044]

[0045] In the above formula, represents the inverse quantized matrix, represents the quantized matrix, z represents the zero point parameter (i.e., zero point) for performing the mapping operation from the integer data space to the floating-point data space, and s represents the scaling parameter (i.e., scale) for performing the mapping operation. It should be understood that in the context of the present disclosure, the zero point parameters in Formula 1 and Formula 2 may be the same or different, and the scaling parameters in Formula 1 and Formula 2 may be the same or different.

[0046] It should be understood that in order to ensure that the inverse quantization can more accurately represent the original matrix, the following conditions should be met:

[0047]

[0048] In the above formula, X represents the data input to the network layer, W0 represents the weight matrix of the network layer, Denotes the inverse quantized weight matrix, and tr denotes the matrix trace. Corresponding zero point parameters and scaling parameters can be determined while satisfying Formula 3. In this way, the error of the inverse quantized weight matrix can be ensured to be within an acceptable range.

[0049] According to an example implementation of the present disclosure, after the quantization of the model has been completed, only the formula 2 needs to be delivered when delivering the model to the downstream processing process. (s, z), and downstream processing does not need to know (s, z) in Formula 1. In this way, the (s, z) in Formula 1 and Formula 2 can be decoupled. At this point, a corresponding objective function can be created for each column vector in the matrix, that is, Formula 3 can be transformed into the following:

[0050]

[0051] In formula 4, g(w;s,z) represents the objective function associated with each column vector, that is, it can be used for To create multiple objective functions. According to an exemplary implementation of the present disclosure, in the process of determining the objective function, the objective function can be generated based on the transpose of the function components, the Hessian matrix and the product of the function components. Specifically, the objective function can be determined using the following formula:

[0052]

[0053] In the above formula, b represents the column vector in W0 w represents a quantized vector (that is, the quantized vector corresponding to each column vector can be expressed as w i , and i=1,2,…,d in ). The mapping parameters s and z represent the scaling parameter and zero-point parameter, respectively, in the inverse quantization process. At this point, the function component of the objective function can be created using the first vector, a second vector corresponding to the first vector from among the plurality of second vectors, and the mapping parameters. In Formula 5, the function component can be expressed, for example, as (w*s+zb).

[0054] According to an example implementation of the present disclosure, the objective function can be determined using the Hessian matrix (e.g., represented as H) of the function components and the input data. Using the example implementation of the present disclosure, each column vector can be processed independently, thereby reducing the amount of computation required to find the optimal solution of Formula 4, and enabling the determined optimal solution to further reduce the difference between the inverse quantized matrix and the original matrix. Specifically, the Hessian matrix can be expressed as: H = X T X, where X represents the input data of the network layer. Using the exemplary implementation of the present disclosure, the process of solving the optimal w, s, z can be converted into a mathematical calculation process, thereby improving the quantization accuracy under a predetermined data width.

[0055] According to an example implementation of the present disclosure, Formula 5 can be substituted into Formula 4, and the optimal value that conforms to Formula 4 can be determined using a variety of solutions currently known and / or to be developed in the future. By using the example implementation of the present disclosure, it is not necessary to worry about the details of existing quantization technology solutions. In other words, a variety of detailed issues, such as how to deal with outliers and how to deal with sensitive channels, can be converted into the problem of determining the optimal solution that conforms to Formula 4. In this way, the accuracy of the quantization model can be greatly improved, that is, the theoretical upper limit of the accuracy of the quantization model, that is, Formula 3, can characterize the ultimate accuracy of the model on a specific task.

[0056] According to an example implementation of the present disclosure, the integer data space has a lower threshold (e.g., represented by α) and an upper threshold (e.g., represented by β), where the lower threshold is lower than zero and the upper threshold is higher than zero. In other words, the integer data space spans zero. According to an example implementation of the present disclosure, the integer data space can be determined in a manner that is as symmetrical as possible, for example, the sum of the lower threshold and the upper threshold can meet a predetermined threshold.

[0057] According to an exemplary implementation of the present disclosure, assuming that the second data width for representing the integer data space is k, the lower limit threshold can be expressed as -2, for example. k-1, and the upper threshold can be expressed as 2 k-1 -1, and the sum of the lower threshold and the upper threshold is "-1". Alternatively and / or additionally, the lower threshold can be expressed as -2, for example. k-1 -1, and the upper threshold can be expressed as 2, for example k-1 , and the sum of the lower threshold and the upper threshold is "-1".

[0058] According to an example implementation of the present disclosure, the proposed quantization technology solution can enable the machine learning model to have acceptable accuracy in extremely low-bit quantization operations. In particular, the second data width includes at least any one of the following: 2, 3, or 4. In other words, even if only 2 bits (3 bits or 4 bits) are used to represent the model weights, the error caused by the dequantized weight data is still within an acceptable range compared to using the original weight data.

[0059] According to an exemplary implementation of the present disclosure, another weight matrix corresponding to the network layer can be generated using multiple third vectors, and another weight matrix (for example, represented as ), process the data input to the network layer. Specifically, each vector w can be determined i Combined into a matrix Then, the determined mapping parameters (s, z) are used to determine the inverse quantization weight matrix based on Formula 2: By using the exemplary implementation of the present disclosure, a dequantized weight matrix can be obtained, and the error caused by the dequantized weight matrix obtained in this way still satisfies an acceptable range.

[0060] Furthermore, at each network layer of the machine learning model, the corresponding weight matrix can be used In this way, the precision of the quantization operation can be improved when the weight matrix of the machine learning model is represented by a finite width, thereby improving the accuracy of the machine learning model.

[0061] Example Process

[0062] Figure 5A flowchart of a method 500 for quantizing data according to some implementations of the present disclosure is shown. At block 510, a plurality of first vectors are extracted from a matrix to be quantized. At block 520, a plurality of objective functions are created, each associated with the plurality of first vectors. The plurality of objective functions respectively include the plurality of first vectors, a plurality of second vectors corresponding to the plurality of first vectors, and mapping parameters, the mapping parameters being used to map the plurality of second vectors to a plurality of third vectors, wherein the second data width corresponding to the plurality of second vectors is less than the first data width corresponding to the plurality of first vectors. At block 530, a plurality of second vectors and the mapping parameters are determined based on the plurality of objective functions. For a first vector in the plurality of first vectors, the mapping parameters are such that a difference between a third vector corresponding to the first vector in the plurality of third vectors and the first vector satisfies a predetermined condition.

[0063] According to an example implementation of the present disclosure, the matrix is ​​a weight matrix of a network layer in a machine learning model, the number of first dimensions of the matrix is ​​determined by the width of the input data of the network layer, and the number of second dimensions is determined by the width of the output data of the network layer.

[0064] According to an example implementation of the present disclosure, creating multiple objective functions includes creating an objective function associated with a first vector among the multiple objective functions based on: creating a function component of the objective function using the first vector, a second vector corresponding to the first vector among multiple second vectors, and a mapping parameter; and determining the objective function using the function component and the Hessian matrix of the input data.

[0065] According to an example implementation of the present disclosure, determining the objective function includes generating the objective function based on a transpose of a function component, a Hessian matrix, and a product of the function component.

[0066] According to an example implementation of the present disclosure, the plurality of first vectors correspond to a floating-point data space, the plurality of second vectors correspond to an integer data space, and the mapping parameters include: a zero value parameter and a scaling parameter for mapping from the integer data space to the floating-point data space.

[0067] According to an example implementation of the present disclosure, the integer data space has a lower threshold and an upper threshold, the lower threshold being lower than a zero value and the upper threshold being higher than a zero value.

[0068] According to an exemplary implementation of the present disclosure, the sum of the lower threshold and the upper threshold satisfies a predetermined threshold.

[0069] According to an example implementation of the present disclosure, the method further includes: generating another weight matrix corresponding to the network layer using the plurality of third vectors; and processing data input to the network layer using the another weight matrix.

[0070] According to an exemplary implementation of the present disclosure, the second data width includes at least any one of the following: 2, 3, or 4.

[0071] According to an example implementation of the present disclosure, the plurality of first vectors include a plurality of columns in a matrix.

[0072] Example devices and equipment

[0073] Figure 6 A block diagram of an apparatus 600 for quantizing data according to some implementations of the present disclosure is shown. The apparatus 600 includes: an extraction module 610 configured to extract a plurality of first vectors from a matrix to be quantized; a creation module 620 configured to create a plurality of objective functions respectively associated with the plurality of first vectors, the plurality of objective functions respectively including the plurality of first vectors, a plurality of second vectors respectively corresponding to the plurality of first vectors, and mapping parameters, the mapping parameters being used to respectively map the plurality of second vectors to a plurality of third vectors, wherein the second data width corresponding to the plurality of second vectors is smaller than the first data width corresponding to the plurality of first vectors; and a determination module 630 configured to determine the plurality of second vectors and the mapping parameters based on the plurality of objective functions, wherein for a first vector among the plurality of first vectors, the mapping parameters ensure that a difference between a third vector corresponding to the first vector among the plurality of third vectors and the first vector satisfies a predetermined condition.

[0074] According to an example implementation of the present disclosure, the matrix is ​​a weight matrix of a network layer in a machine learning model, the number of first dimensions of the matrix is ​​determined by the width of the input data of the network layer, and the number of second dimensions is determined by the width of the output data of the network layer.

[0075] According to an example implementation of the present disclosure, creating multiple objective functions includes creating an objective function associated with a first vector among the multiple objective functions based on: a component creation module configured to create a function component of the objective function using the first vector, a second vector corresponding to the first vector among multiple second vectors, and a mapping parameter; and an objective function determination module configured to determine the objective function using the function component and the Hessian matrix of the input data.

[0076] According to an exemplary implementation of the present disclosure, the objective function determination module includes: a generation module configured to generate an objective function based on a transpose of a function component, a Hessian matrix, and a product of the function component.

[0077] According to an example implementation of the present disclosure, the plurality of first vectors correspond to a floating-point data space, the plurality of second vectors correspond to an integer data space, and the mapping parameters include: a zero value parameter and a scaling parameter for mapping from the integer data space to the floating-point data space.

[0078] According to an example implementation of the present disclosure, the integer data space has a lower threshold and an upper threshold, the lower threshold being lower than a zero value and the upper threshold being higher than a zero value.

[0079] According to an exemplary implementation of the present disclosure, the sum of the lower threshold and the upper threshold satisfies a predetermined threshold.

[0080] According to an example implementation of the present disclosure, the module further includes: a matrix generation module configured to generate another weight matrix corresponding to the network layer using multiple third vectors; and a processing module configured to process data input to the network layer using the other weight matrix.

[0081] According to an exemplary implementation of the present disclosure, the second data width includes at least any one of the following: 2, 3, or 4.

[0082] According to an example implementation of the present disclosure, the plurality of first vectors include a plurality of columns in a matrix.

[0083] Figure 7 FIG. 7 is a block diagram of a device 700 capable of implementing various implementations of the present disclosure. Figure 7 The illustrated computing device 700 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. Figure 7 The illustrated computing device 700 may be used to implement the methods described above.

[0084] like Figure 7 As shown, computing device 700 is in the form of a general-purpose computing device. Components of computing device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 700.

[0085] The computing device 700 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 700.

[0086] The computing device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 7 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0087] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.

[0088] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 700 may also communicate with one or more external devices (not shown) via communication unit 740, as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 700, or with any device that allows computing device 700 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0089] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0090] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0091] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0092] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0093] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0094] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for quantifying data, comprising: extracting a plurality of first vectors from a matrix to be quantized; creating a plurality of objective functions respectively associated with the plurality of first vectors, the plurality of objective functions respectively including the plurality of first vectors, a plurality of second vectors respectively corresponding to the plurality of first vectors, and mapping parameters, the mapping parameters being used to respectively map the plurality of second vectors to a plurality of third vectors, wherein a second data width corresponding to the plurality of second vectors is smaller than a first data width corresponding to the plurality of first vectors; as well as The plurality of second vectors and the mapping parameters are determined based on the plurality of objective functions, and for a first vector among the plurality of first vectors, the mapping parameters make the difference between a third vector among the plurality of third vectors corresponding to the first vector and the first vector satisfy a predetermined condition.

2. The method of claim 1 , wherein the matrix is ​​a weight matrix of a network layer in a machine learning model, the number of first dimensions of the matrix is ​​determined by the width of input data of the network layer, and the number of second dimensions is determined by the width of output data of the network layer.

3. The method of claim 2 , wherein creating the plurality of objective functions comprises creating an objective function associated with the first vector among the plurality of objective functions based on: creating a function component of the objective function using the first vector, a second vector corresponding to the first vector among the plurality of second vectors, and the mapping parameter; and The objective function is determined using the function components and the Hessian matrix of the input data.

4. The method of claim 3, wherein determining the objective function comprises: The objective function is generated based on the transpose of the function components, the Hessian matrix, and the product of the function components.

5. The method of claim 1 , wherein the plurality of first vectors correspond to a floating-point data space, the plurality of second vectors correspond to an integer data space, and the mapping parameters comprise: A zero value parameter and a scaling parameter are used to map from the integer data space to the floating point data space. The method of claim 5 , wherein the integer data space has a lower threshold and an upper threshold, the lower threshold being below zero and the upper threshold being above zero. The method according to claim 6 , wherein the sum of the lower threshold and the upper threshold satisfies a predetermined threshold.

8. The method according to claim 2, further comprising: generating another weight matrix corresponding to the network layer using the plurality of third vectors; as well as Data input to the network layer is processed using the other weight matrix. 9 . The method according to claim 1 , wherein the second data width comprises at least any one of the following: 2, 3, or 4.

10. The method of claim 1, wherein the plurality of first vectors comprises a plurality of columns in the matrix.

11. A device for quantifying data, comprising: an extraction module configured to extract a plurality of first vectors from a matrix to be quantized; a creation module configured to create a plurality of objective functions respectively associated with the plurality of first vectors, the plurality of objective functions respectively including the plurality of first vectors, a plurality of second vectors respectively corresponding to the plurality of first vectors, and mapping parameters, the mapping parameters being used to respectively map the plurality of second vectors to a plurality of third vectors, wherein a second data width corresponding to the plurality of second vectors is smaller than a first data width corresponding to the plurality of first vectors; as well as A determination module is configured to determine the multiple second vectors and the mapping parameters based on the multiple objective functions, and for a first vector among the multiple first vectors, the mapping parameters make the difference between a third vector among the multiple third vectors corresponding to the first vector and the first vector meet a predetermined condition.

12. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 10.